EconBase
← Back to paper

Cluster-robust inference with a single treated cluster using the t-test

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

87,279 characters · 23 sections · 72 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Cluster-robust inference with a single treated cluster using the t-test

\fontfamily{ppl}

abstractThis paper considers inference when there is a single treated cluster and a fixed number of control clusters, a setting that is common in empirical work, especially in difference-in-differences designs. We use the $t$-statistic and develop suitable critical values to conduct valid inference under weak assumptions allowing for unknown dependence within clusters. In particular, our inference procedure does not involve variance estimation. It only requires specifying the relative heterogeneity between the variances from the treated cluster and some, but not necessarily all, control clusters. Our proposed test works for any significance level when there are at least two control clusters. When the variance of the treated cluster is bounded by those of all control clusters up to some prespecified scaling factor, the critical values for our $t$-statistic can be easily computed without any optimization for many conventional significance levels and numbers of clusters. In other cases, one-dimensional numerical optimization is needed and is often computationally efficient. We have also tabulated common critical values in the paper so researchers can use our test readily. We illustrate our method in simulations and empirical applications. Keywords: Cluster-robust inference, single treated cluster, $t$-test, simultaneous inference, difference-in-differences.

\onehalfspacing

Introduction

In difference-in-differences designs, it is common for researchers to conduct inference using cluster-robust methods to account for correlation within clusters. However, inference becomes challenging when there is only a single treated cluster. Intuitively, this is because researchers are faced with only one estimate from the treated cluster, making its uncertainty difficult to quantify. Existing methods either impose strong assumptions on having a large number of clusters, require the variances to be homogeneous or estimable, or only work for specific significance levels depending on the number of clusters. In this paper, we develop a $t$-test associated with a suitable critical value to conduct valid inference that relaxes these assumptions.

Having a single treated cluster and a finite number of control clusters is common in empirical work. For instance, this can happen when researchers use data in the United States to examine the impact of a policy that takes place in one particular state but not the other states. The nearby states or the remaining states are commonly used as the control group. In these scenarios, it is reasonable for researchers to believe they are not in the scenario with a “large” number of control clusters. Some recent examples in this context include wangburke2022jfe on the effect of payday loan regulations in Texas, harrislarsen2023jhr on the effect of Hurricane Katrina on student outcomes in New Orleans, dillenderetal2023jpube on the effect of change in health care reimbursement rates in Illinois, alpertetal2024aejep on the impact of Kentucky's prescription drug monitoring programs on opioid prescribing, and kumarliang2024aejep on the labor market effects of constitutional amendments in Texas. hagemann2024wp has documented earlier related examples in the context of a single treated cluster. In the next section of this paper, we demonstrate that our test is also applicable to other empirical designs, in addition to difference-in-differences.

This paper contributes to the literature on cluster-robust inference. We assume that the number of clusters is fixed, as in some related work in this literature such as besteretal2011joe, ibraimovmuller2010jbes, ibraimovmuller2016restat, canayetal2017ecta, hagemann2022wp, and lau2025wp. Although the tests in the aforementioned papers are valid when there is a finite number of clusters, they are not suitable for the problem with a single treated cluster. The first test requires certain homogeneity conditions, and the other tests cannot be applied when there is a single treated cluster. canaysantosshaikh2021restat has showed that Wild cluster bootstrap popularized by cameronetal2008restat can be valid under strong homogeneity conditions when there is a fixed number of clusters. See, for instance, cameronmiller2015jhr, conleyetal2018jar, mackinnonetal2023joe, and alvarezetal2025wp for some surveys on the literature of conducting inference with a fixed number of clusters.

Several tests have been developed to conduct inference when there is a single treated cluster, but with assumptions that can be strong in practice. conleytaber2011restat assume homogeneous errors. fermanpinto2019restat relax the homogeneity assumption and allow for known heteroskedasticity. alvarezferman2023wp relax the homogeneity assumption and allows for spatial correlation. All these tests assume that there is an infinite number of control clusters. Our $t$-test neither assumes the variances are known nor assumes an infinite number of control clusters. Recently, hagemann2024wp develops a novel rearrangement test that relaxes the assumptions mentioned earlier. hagemann2024wp is the most related paper in that his test relies on a relative heterogeneity condition between the treated and control clusters. Both our test and hagemann2024wp work with a fixed number of clusters, allow for arbitrary correlation within clusters, and do not assume that we can estimate or know the cluster variances. He shows his test is valid when the relative heterogeneity condition holds with all or all but one control cluster. However, the validity of his test depends on the number of clusters, the significance level, and the relative heterogeneity condition. For example, when the standard deviation of the treated cluster is bounded by $2$ times all but one of the standard deviations of the control clusters, hagemann2024wp requires at least 14 control clusters to conduct a one-sided test at the 2.5% level using his rearrangement test, and at least 17 control clusters are needed for a 1% level test\footnote{Equivalently, to conduct two-sided tests, there have to be at least 14 control clusters for a 5% level test.}. This potentially limits the applicability of his test. Among the cases where his test is valid, he computes weights for his rearrangement test using numerical optimization that lead to valid tests.

There are several important distinctions between our work and hagemann2024wp. First, we show that our test can be valid for any choice of significance levels and heterogeneity parameters when there are at least two control clusters. For many conventional combinations of significance levels and number of clusters, we derive a closed-form critical value for the $t$-test. No optimization is necessary for these cases, so researchers can apply the test readily. For other cases, we can derive critical values that lead to valid tests through one-dimensional optimizations. Therefore, our test can have broader applicability. Second, we relax the relative heterogeneity condition in hagemann2024wp. Intuitively, the condition restricts the relative variance between the treated and control clusters. hagemann2024wp requires such condition to hold between the treated and all (or all but one) control clusters to show the validity of his test. We weaken this condition by allowing the researcher to impose this restriction between the treated cluster and any number of control clusters. This allows researchers to bound relative heterogeneity on variances depending on their choice on how to restrict the amount of heterogeneity between the treated and control clusters. Third, we allow for simultaneous inference across all the relative heterogeneity assumptions that we impose. Specifically, under the assumption of no treatment effect, we can infer the minimum amount of relative heterogeneity required to explain away the observed association between treatment and outcomes. If the true relative heterogeneity is unlikely to exceed this amount, the treatment is likely to have a significant nonzero effect.

While the $t$-test has been used in Bakirov:2006aa and ibraimovmuller2010jbes, ibraimovmuller2016restat, our proof on the validity of the $t$-test in the current context requires different proof strategies. This is because we only have a single treated cluster and the relative heterogeneity assumption imposes additional structure on the parameter space.

We confirm the performance of our proposed test in simulations. We find that our test controls size under a wide range of significance levels and number of clusters. We also have favorable power performance when compared to other tests in the literature.

The rest of the paper is organized as follows. Section (ref) outlines some popular empirical designs that our test can be applied to. Section (ref) presents our inference procedure. Readers interested in applying the test can refer to Algorithm (ref). Section (ref) presents our main theory. Section (ref) explains how our test can be easily used for simultaneous inference. Section (ref) presents the simulation results. Section (ref) contains two empirical studies. Section (ref) concludes. All proofs can be found in the supplementary material.

Motivating examples

We start by describing several empirically relevant designs that are common in applied work and are related to the issue of having a single treated cluster. These examples show that our method applies more broadly apart from standard difference-in-differences designs. The first three examples have also been discussed in hagemann2024wp.

In the following, $i \in \cI \equiv \{1, \ldots, n\}$ indices units, $j \in \cJ \equiv \{1, \ldots, m + 1\}$ indices clusters, and $t \in \cT \equiv \{ 1, \ldots, T\}$ indices time. Only cluster $(m+1)$ is treated, and the remaining clusters are controls. Thus, $\cJ$ can be partitioned as $\cJ_0 \cup \cJ_1$, where $\cJ_0 \equiv \{1, \ldots, m\}$ and $\cJ_1 \equiv \{m + 1\}$ denote the control clusters and treated cluster respectively. Let $D_j \equiv \ind[j \in \cJ_1]$ be a cluster-level treatment indicator across $j \in \cJ$.

As to be described in Section (ref), we require researchers to be able to run regression cluster by cluster, and write the estimator for the target parameter $\Delta$ in terms of the estimator for parameters from cluster-level regressions, denoted by $\{\widehat\theta_j\}^{m+1}_{j=1}$. The following examples explain how to connect $\{\widehat\theta_j\}^{m+1}_{j=1}$ with the estimator for the target parameters in common empirical designs.

eg[Clustered regression] Consider the following model \[ Y_{ij} = \beta_0 + \Delta D_j + U_{ij}, \] where $\beta_0, \Delta, U_{ij} \in \bR$. Under the assumption that $\bE[U_{ij}| D_j] = 0$, $\Delta$ can be estimated by \[ \widehat\Delta = \widehat{\theta}_{m+1} - \frac{1}{m} \sum^m_{j=1} \widehat{\theta}_j, \] where $\widehat{\theta}_j$ is the sample mean of $Y_{ij}$ for each $j \in \cJ$.
eg[Difference-in-differences] Suppose $T = 2$ and let $\text{Post}_t \equiv \ind[t = 2]$ for all $t \in \cT$. Consider the following model \begin{equation} Y_{jt} = \alpha_j + \beta Post_{t} + \Delta D_j Post_{t} + U_{jt}, \end{equation} where $U_{jt} \in \bR$, and $\{\alpha_{j}\}_{j=1}^{m+1}$ are the cluster fixed effects. Under the assumption that $\bE[U_{jt} | D_j, \text{Post}_t] = 0$, $\Delta$ can be estimated by \[ \widehat\Delta = \widehat{\theta}_{m+1} - \frac{1}{m} \sum_{j=1}^m \widehat{\theta}_j, \] where $\{\widehat{\theta}_j\}^{m+1}_{j=1}$ are the differences in the outcomes before and after treatment. They can be obtained as the estimates of $\{\theta_j\}^{m+1}_{j=1}$ from the following cluster-level regressions \begin{align} Y_{jt} = \alpha_j + \theta_j Post_t + \epsilon_{jt} \end{align} for each $j \in \cJ$.
eg[Two-way fixed effects] Consider the following two-way fixed effects model: \begin{equation} Y_{jt} = \alpha_j + \beta_t + \Delta D_j Post_t + U_{jt}, \end{equation} where $T \geq 2$, $\{\alpha_{j}\}_{j=1}^{m+1}$ are the cluster fixed effects, $\{\beta_t\}_{t=1}^T$ are the time fixed effects, and $\text{Post}_t \equiv \ind[t \geq t_0]$ for some $1 < t_0 \leq T$. Under the assumption that $\bE[U_{jt} | D_j, \text{Post}_t]=0$, $\Delta$ can be estimated by \[ \widehat\Delta = \widehat{\theta}_{m+1} - \frac{1}{m} \sum^m_{j=1} \widehat{\theta}_j, \] using the same cluster-level regression as in (ref).

For more discussion about recent advances in difference-in-differences and two-way fixed effects, see, for instance, the surveys by dechaisemartindhaultfœuille2023ej, rothetal2023joe, and bakeretal2025wp.

eg[Triple differences] Let $C_{ij}, D_{ij} \in \{0, 1\}$ be binary indicators that depend on the individual $i \in \cI$ and cluster $j \in \cJ$. Set $D_{ij} = \ind[j = {m+1}]$ and define $\text{Post}_t\equiv \ind[t = 2]$ as in Example (ref) with two periods. Assume for each $j \in \cJ$, there exist units with $C_{ij} = 1$ and units with $C_{ij} = 0$. Consider the following triple differences/difference-in-difference-in-differences model: \begin{align} \begin{split} Y_{ijt} & = \beta_0 + \beta_1 C_{ij} + \beta_2 D_{ij} + \beta_3 Post_t + \beta_4 C_{ij} D_{ij} \\ &\quad \quad + \beta_5 C_{ij} Post_t + \beta_6 D_{ij} Post_t + \Delta C_{ij} D_{ij} Post_t + U_{ijt}. \end{split} \end{align} Under the assumption that $\bE[U_{ijt} | C_{ij}, D_{ij}, \text{Post}_t] = 0$, $\Delta$ can be estimated by \[ \widehat\Delta = \widehat{\theta}_{m+1} - \frac{1}{m} \sum^m_{j=1} \widehat{\theta}_j, \] where $\{\widehat{\theta}_j\}^{m+1}_{j=1}$ are the estimates of $\{\theta_j\}^{m+1}_{j=1}$ from the following cluster-level regressions \[ Y_{ijt} = \alpha_j + \gamma_{C,j} C_{ij} + \gamma_{\text{Post},j} \text{Post}_t + \theta_j C_{ij} \text{Post}_t + \epsilon_{ijt}, \] for each $j \in \cJ$. See oldenmoen2023ej for a recent survey on triple difference estimators.

Inference procedure

In this section, we present the main assumptions and our algorithm of conducting inference using the $t$-test with a single treated cluster. Readers who are interested in applying our test can directly apply Algorithm (ref). We also present some brief intuition for the validity of our test in Section (ref) to facilitate the theoretical discussion in Section (ref).

Assumptions

We follow the notation used in Section (ref). Let $\{\widehat{\theta}_j\}^{m+1}_{j=1}$ be the cluster-level estimators, $\cJ_0 = \{1, \ldots, m\}$ index the control clusters and $\cJ_1 =\{m + 1\}$ index the treated cluster. For simplicity, we make the dependence of $\{\widehat{\theta}_j\}^{m+1}_{j=1}$ on the cluster size implicit.

Our procedure requires two main assumptions. First, we assume that an appropriate central limit theorem applies to $\{\widehat{\theta}_j\}^{m+1}_{j=1}$ as the sample size within each cluster, denoted by $n$, goes to infinity. A similar assumption is also imposed in related papers on a fixed number of clusters, such as ibraimovmuller2010jbes, ibraimovmuller2016restat, canayetal2017ecta, hagemann2022wp, hagemann2024wp and lau2025wp. Intuitively, this holds when each cluster has a large number of units or consists of a panel with a long time periods. This high-level assumption is formalized as follows.

assuThe following holds as the sample size within each cluster $n \longrightarrow \infty$: \begin{align} \sqrt{n} \begin{pmatrix} \widehat{\theta}_1 - \mu_0 \\ \vdots \\ \widehat{\theta}_m - \mu_0 \\ \widehat{\theta}_{m+1} - \mu_1 \end{pmatrix} \overset{d}{\longrightarrow} \cN(0, \Sigma), \end{align} where $\Sigma \equiv \text{diag}(\sigma_1^2, \ldots, \sigma_m^2, \sigma_{m+1}^2)$ is an diagonal matrix.

Below we give a few remarks regarding Assumption (ref). First, it is not necessary for all clusters to have the same size. In cases where clusters have varying sizes, we can define $n$ as the size of the smallest cluster, and our inference essentially requires the sample sizes in all clusters to be large. Second, we assume that the cluster-level estimators are independent, at least asymptotically, so that the asymptotic covariance matrix $\Sigma$ in (ref) is diagonal. This is trivially satisfied if the units are independent across clusters. We emphasize that, while we require independence across clusters, we allow for unknown dependence structure within clusters. Third, we assume that the estimators from the control clusters are consistent for a common parameter $\mu_0$, which can often be justified by model assumptions or study designs as illustrated in Section (ref).

Second, we impose the following relative heterogeneity assumption, which generalizes the relative heterogeneity assumption in hagemann2024wp. Let $\sigma_{(1)} \leq \sigma_{(2)} \leq \cdots \leq \sigma_{(m)}$ be the ordered values of $\{\sigma_j\}^m_{j=1}$ in Assumption (ref).

assuFor a given $\rho \ge 0$ and a given $k \in \{1, \ldots, m\}$, $\sigma_{m+1} \leq \rho \sigma_{(k)}$.

The above assumption does not require the variances to be known. It only restricts the relative heterogeneity of the standard deviations between the treated and control clusters. In particular, for any $\rho \geq 0$ and $k \in \{1, \ldots, m\}$, it requires the standard deviation of the treated cluster to be less than or equal to $\rho$ times the standard deviations of at least $(m - k + 1)$ control clusters. For example, $k = 1$ requires that the standard deviation of the treated cluster is smaller than or equal to $\rho$ times the standard deviations of all control clusters; when $k=m$, this means that the standard deviation of the treated cluster is smaller than or equal to $\rho$ times the largest standard deviation from the control clusters.

Assumption (ref) is a relaxed version of hagemann2024wp's maximum relative heterogeneity assumption. His assumption is equivalent to Assumption (ref) with $k = 1$ or $k = 2$. Our test allows a general choice of $k \in \{1, \ldots, m\}$. More importantly, as demonstrated in Section (ref), the inference can be simultaneously valid for all $k \in \{1, \ldots, m\}$. In other words, with additional choices of $k$, our test can provide more evidences against the null hypothesis than hagemann2024wp's test, which focuses on $k=1$ or $2$.

Before introducing our algorithm, we end this subsection with some discussion on choosing $(\rho, k)$ in Assumption (ref). If the researcher believes the standard deviation of the treated cluster cannot be larger than the standard deviations of the control clusters, then the researcher can set $k = 1$ and $\rho = 1$.

On the other hand, if the researcher does not want to commit to a particular value of $(\rho, k)$, the researcher can perform simultaneous inference and report the largest value of $\rho$ such that the null can be rejected for each $k$. This approach can also be interpreted as finding the largest $\rho$ such that the conclusion changes, which is related to the idea of finding breakdown points/frontiers in econometrics horowitzmanski1995ecta, klinesantos2013qe, mastenpoirier2020qe. We demonstrate this through two applications in Section (ref).

Algorithm

The goal is to test the following hypothesis:

equation[equation omitted — 138 chars of source]

Our procedure of conducting inference with a single treated cluster is as follows.

algoInputs: \begin{itemize} • Significance level $\alpha \in (0, \frac{1}{2})$. • The parameters on relative heterogeneity from Assumption (ref), i.e., $(\rho, k)$. • The cluster-level estimators $\{\widehat{\theta}_j\}^{m+1}_{j=1}$. \end{itemize} Steps: \begin{description} • Compute the $t$-statistic: \begin{align} \widehat{T}_m \equiv \frac{\widehat\theta_{m+1} - \overline{\widehat{\theta}}_m}{\widehat S_m}, \end{align} where $\overline{\widehat{\theta}}_m \equiv \frac{1}{m} \sum^m_{j=1} \widehat{\theta}_j$ and $\widehat S_m^2 \equiv \frac{1}{m-1} \sum^m_{j=1} (\widehat{\theta}_j - \overline{\widehat{\theta}}_m)^2$ represent the sample average and sample variance of the estimators from the control clusters, respectively. • We have tabulated many empirically-relevant critical values in Table (ref). These critical values can be immediately used. More generally, the critical value $\mathrm{cv}_{m, \alpha, k, \rho}$ can be computed as follows. \begin{itemize}[leftmargin=*] • Case 1: If $k = 1$ and $\alpha$ is below the thresholds in Table (ref), use the closed-form expression in (ref). Compute the closed-form critical value as \[ \mathrm{cv}_{m, \alpha, k, \rho} \equiv \sqrt{\rho^2 + \frac{1}{m}} \ t_{m-1, \frac{\alpha}{2}}, \] where $t_{ m-1, \frac{\alpha}{2}}$ is the $(1-\frac{\alpha}{2})$-th quantile of the $t$-distribution with $(m-1$) degrees of freedom. • Case 2: Search $\mathrm{cv}_{m, \alpha, k, \rho}$ such that $p_m(\mathrm{cv}_{m, \alpha, k, \rho}; k, \rho)$ defined in Theorem (ref) ahead is at most $\alpha$. This requires solving one-dimensional optimization problems. \end{itemize} • Reject (ref) if $|\widehat{T}_m| > \mathrm{cv}_{m, \alpha, k, \rho}$. $\blacksquare$ \end{description}

Computing p-value and confidence interval

To compute the $p$-value, the researcher does not need to run the entire Algorithm (ref). In particular, the researcher only needs to compute the test statistic $\widehat T_m$ in Step 1 of Algorithm (ref) and substitute it into the $p_m$ function in Step 2 of Algorithm (ref) and return the $p$-value as $p_m(\widehat T_m; k, \rho)$.

The researcher can take the critical value $\mathrm{cv}_{m, \alpha, k, \rho}$ from Step 2 of Algorithm (ref) and return $(\widehat\theta_{m+1} - \overline{\widehat{\theta}}_m) \pm \mathrm{cv}_{m, \alpha, k, \rho} \widehat S_m$ as an $(1 - \alpha)$ confidence interval for $\mu_1 - \mu_0$.

Discussion

Algorithm (ref) is a computationally simple three-step procedure. The first step is to compute the usual $t$-statistic using the cluster-level estimators.

The second step finds the critical value that depends on $(m, \alpha, k, \rho)$. This step aims to find the critical value $\mathrm{cv}_{m, \alpha, k, \rho}$ such that the procedure controls size for any configurations of $\{\sigma_j\}^{m+1}_{j=1}$ that satisfy Assumption (ref). In the first case, a closed-form critical value is available when $\alpha$ is below a certain threshold in Table (ref). If $\alpha$ does not satisfy the condition, we can compute the critical value $\mathrm{cv}_{m, \alpha, k, \rho}$ in the second step through one-dimensional optimization.

The reason for having a closed-form expression for the critical value in case 1 of Step 2 of Algorithm (ref) is as follows. For $k=1$ and any given $(m, \rho)$, for a sufficiently small significance level $\alpha$ (including many conventional choices), we find that the maximum rejection probability over all configurations of $\{\sigma_j\}^{m+1}_{j=1}$ is achieved when $\sigma_{m+1}=\rho \sigma_j$ for all $j = 1, \ldots, m$. The proof is nontrivial and we explain the technical details in Sections (ref). Knowing when the maximum rejection probability is achieved, we are able to derive a closed-form expression for the critical value using properties of the $t$-distribution.

Case 2 of Step 2 is in fact a general version that holds for any $(m, \alpha, k, \rho)$. For a general value of $k \in \{1, \ldots, m\}$, although we cannot find the exact configuration of $\{\sigma_j\}^{m+1}_{j=1}$ that achieves the desired level of rejection probability, we are able to substantially reduce the set of possible values. In particular, there are at most $m^2$ possible cases for the worst-case configuration of $\{\sigma_j\}^{m+1}_{j=1}$, each of which involves at most one unknown parameter. Therefore, we can find the maximum rejection probability through one-dimensional optimization. In addition, we have a good initial guess for the critical value. To facilitate the use of the test in practice, we have tabulated many critical values for researchers' use in Table (ref).

The last step rejects if the absolute value of the test statistic is above the critical value.

Theory

In this section, we develop the theory that justifies the validity of Algorithm (ref). In Section (ref), we show that the large-sample behavior of the test can be studied via a fixed number of normal random variables. In Section (ref), we show the general theory on computing the maximum rejection probability of the test statistic. This result is used to search for a critical value such that the test is valid. We show in Section (ref) that we can get a closed-form expression for the critical value as in Step 2 of Algorithm (ref) when $k = 1$ and the significance level is not “too large.” In Section (ref), we discuss the power of our $t$-test.

General framework with normal means

In this subsection, we show that under Assumption (ref), studying the behavior of the $t$-statistic in equation (ref) of Algorithm (ref) when $n$ is large is equivalent to studying the $t$-statistic of suitably defined $(m+1)$ normal random variables. Formally, consider the following assumption on the normal random variables.

assuLet $\{\psi_j\}^{m+1}_{j=1}$ be $m+1$ independent random variables, where $\psi_j \sim \cN(0, \sigma_j^2)$ for $1\le j \le m$ and $\psi_{m+1} \sim \cN(\delta, \sigma_{m+1}^2)$.

Compared to the notation in Section (ref), $\{\psi_j\}^m_{j=1}$ above correspond to the $m$ control clusters and $\psi_{m+1}$ corresponds to the single treated cluster. As before, the above assumption does not require us to know the variances $\{\sigma_j\}^{m+1}_{j=1}$. The variances can be arbitrarily heterogeneous as long as they satisfy the relative heterogeneity assumption stated in Assumption (ref). Analogous to (ref), we define the following $t$-statistic based on these normal random variables:

equation[equation omitted — 94 chars of source]

where $\overline\psi_m \equiv \frac{1}{m} \sum^m_{j=1} \psi_j$ and $S_m^2 \equiv \frac{1}{m-1} \sum^m_{j=1} (\psi_j - \overline\psi_m)^2$ denote the sample mean and variance for the control clusters, respectively.

The theorem below shows that, to achieve the desired size asymptotically, it suffices to consider a stylized setting with normally distributed random variables. Note that $\widehat{T}_m$ in (ref) depends on the sample size within each cluster $n$ implicitly. In addition, we allow $\mu_1$ and $\mu_0$ to vary with the sample size $n$ as well, which can facilitate power investigation under local alternatives.

thmLet $m\ge 2$. Suppose that Assumption (ref) holds, $\sqrt{n}(\mu_1 - \mu_0) \longrightarrow \delta$ as $n\longrightarrow \infty$ for some $\delta \in \mathbb{R}$, and the variances $\{\sigma_j^2\}^{m+1}_{j=1}$ are not all zero. Then, for any $c > 0$ and $c\ne m^{-1/2}$,\footnote{We impose $c\ne m^{-1/2}$ to avoid the case where $|T_m|$ has a positive point mass at $c$. This technical requirement excluding a single value of $c$ generally does not affect the practical use of our test, since the desired critical value at a usual significance level is greater than $m^{-1/2}$. Specifically, as discussed later in Remark (ref), the rejection probability $\bP[|T_m| > c]$ for any $c<m^{-1/2}$ can be as large as $1$, and the rejection probability $\bP[|T_m| > c]$ for $c$ close to but larger than $m^{-1/2}$ can at least be about $0.5$. } \begin{align*} \lim_{n\to\infty} \bP[|\widehat{T}_m| > c] = \bP[|T_m| > c], \end{align*} where $T_m$ is the $t$-statistic defined in (ref).

Note that in the asymptotics of Theorem (ref), the limit is taken as $n$ goes to infinity, while the number of clusters $(m+1)$ is fixed. In addition, we allow a general $\delta$ for the difference between $\mu_1$ and $\mu_0$ to facilitate the later discussion on the power of the test. To derive valid tests, it suffices to consider the case where $\mu_1 = \mu_0$ and $\delta = 0$. Specifically, to derive large-sample valid $t$-test with a single treated cluster and a finite number of control clusters, it suffices to test the following null hypothesis in the stylized setting with exactly normal observations:

equation[equation omitted — 125 chars of source]

In the remainder of this section, we will study valid $t$-test for (ref) in this stylized setting.

reIf we are interested in an one-sided alternative, such as $\mu_1 > \mu_0$, we can reject the null hypothesis if $\widehat{T}_m > c$ for some $c > 0$. Moreover, under the same condition as in Theorem (ref), $\lim_{n\to\infty} \bP[\widehat{T}_m > c] = \bP[T_m > c] = \frac{1}{2}\bP[|T_m| > c]$. Therefore, the critical value for a level-$\alpha$ one-sided test can be derived from the critical value for a level-($2\alpha$) two-sided test, for any $\alpha\in [0, \frac{1}{2}]$.
reThe setup in this paper also works when there is a single control cluster and at least two treated clusters. All the results would follow by labeling the first $m$ clusters as the treated clusters and the $(m+1)$th cluster as the control cluster.
reAlthough we focus on a fixed $m$, we discuss in Appendix (ref) on the behavior of the test when $m$ is large, where we derive a closed-form valid test and approximate its power.

Valid $t$-test

The key to constructing a valid $t$-test is to find an appropriate critical value $\mathrm{cv}_{m, \alpha, k, \rho}$ for Step 2 of Algorithm (ref). This critical value $\mathrm{cv}_{m, \alpha, k, \rho}$ has to be chosen such that for any $\{\sigma_j\}^{m+1}_{j=1}$ that satisfy Assumption (ref), $\bP[|T_m|>\mathrm{cv}_{m, \alpha, k, \rho}]$ under the null hypothesis (ref) is less than or equal to a given significance level $\alpha$. In the following, we will consider the maximum rejection probability at any critical value $c$.

Note that when the treated standard deviation is much larger than the control standard deviations, i.e., $\sigma_{m+1} \gg \max_{1\le j \le m}\sigma_j$, the rejection probability $\bP[|T_m|>c]$ will be close to $1$ for any $c>0$, under which we cannot derive a meaningful $t$-test. Therefore, Assumption (ref) on the relative heterogeneity of $\{\sigma_j\}^{m+1}_{j=1}$ is in some sense necessary.

For descriptive convenience, we introduce $\mathcal{S}_m(k, \rho)$ to denote all possible standard deviations that satisfy Assumption (ref) for a given $m$, $\rho \geq 0$ and $k \in \{1, \ldots, m\}$\footnote{With a slight abuse of notation, we also use $\{\sigma_j\}_{j=1}^{m+1}$ to denote generic values of the standard deviations. Note that $\{\sigma_j\}_{j=1}^{m+1}$ in Assumptions (ref) and (ref) represent the true (asymptotic) standard deviations for all the clusters.}:

align[align omitted — 180 chars of source]

Using the above notation, our goal of finding the maximum rejection probability as described above can be formalized as follows

align[align omitted — 157 chars of source]

where the $\bP_0$ notation with subscript $0$ indicates that the null hypothesis (ref) holds and $\delta=0$. To conduct a valid test at any given significance level $\alpha\in(0,1)$, it suffices to find the critical value $c$ such that $p_m(c; k, \rho) \le \alpha$ and ideally $p_m(c; k, \rho) = \alpha$. On the other hand, if the goal is to compute a $p$-value, then one can directly compute $p_m(|T_m|; k, \rho)$ without searching for the required critical value as discussed in Section (ref). That is, when we set $c$ to be the observed absolute value of the $t$-statistic, $p_m(c; k, \rho)$ gives a valid $p$-value for testing the null hypothesis in (ref).

In the following, we consider two cases depending on the standard deviation for the treated cluster $\sigma_{m+1}$. Section (ref) considers $\sigma_{m+1} = 0$ and shows that the corresponding maximum rejection probability has a closed-form solution that can be easily computed. Section (ref) considers $\sigma_{m+1} > 0$ and shows that we have an integral representation for the rejection probability. This facilitates its numerical calculation at any given values of $\{\sigma_j\}_{j=1}^{m+1}$. However, directly evaluating the optimization problem in (ref) generally results in solving an $m$-dimensional optimization problem. We discuss how this can be reduced to solving multiple one-dimensional optimization problems in Section (ref).

Case 1: $\sigma_{m+1} = 0$

When $\sigma_{m+1} = 0$, regardless of the choices of the relative heterogeneity parameters $\rho$ and $k$, the control standard deviations $\{\sigma_j\}^m_{j=1}$ can take arbitrary values in $\mathbb{R}^m_{\ge 0}$. Moreover, in this case, except for a multiplicative constant scaling factor of $\sqrt{m}$ in the $t$-statistic, our $t$-statistic essentially reduces to a one-sample $t$-statistic, and our $t$-test is equivalent to testing whether the mean of the control clusters is equal to zero. From Bakirov:2006aa, we know that the maximum rejection probability with a zero $\sigma_{m+1}$ has the following form:

align[align omitted — 280 chars of source]

where $R_m(c) \equiv \frac{m^2c^2}{mc^2+m-1}$, and $t_{j-1}$ denotes a random variable following the $t$-distribution with degrees of freedom $(j-1)$. In (ref), the maximum rejection probability is obtained when the control standard deviations $\{\sigma_j\}_{j=1}^m$ are either zero or take some common positive value. It can be efficiently computed by calculating the tail probabilities of at most $(m-1)$ $t$-distributions with various degrees of freedom.

It follows that when $\rho = 0$, we must have $\sigma_{m+1}=0$, and the maximum rejection probability of our $t$-test becomes the same as (ref), i.e., $p_m(c; k, 0) = p_{m,0}(c)$ for any $c>0$ and $k \in \{1, \ldots, m\}$. Thus, in the remainder of this section, we focus on $\rho > 0$.

reFrom (ref) and as commented in Bakirov:2006aa, $p_{m,0}(c) = 1$ for $0<c<m^{-1/2}$ and $p_{m,0}(c) = 0.5$ for $c = m^{-1/2}$. Furthermore, by the right-continuous property of distribution functions, we can verify that $p_{m,0}(c) \longrightarrow 0.5$ as $c$ approaches $m^{-1/2}$ from the right. These observations mean that, regardless of the values of $\rho$ and $k$, the maximum rejection probability of our $t$-test in (ref) is $1$ when $0<c<m^{-1/2}$, at least $0.5$ when $c = m^{-1/2}$, and at least about $0.5$ when $c$ is greater than but close to $m^{-1/2}$. Thus, it is generally innocuous to assume $c\ne m^{-1/2}$ (or even $c > m^{-1/2}$) as in Theorems (ref) and later in (ref) for most conventional significance levels.

Case 2: $\sigma_{m+1} > 0$

We now consider the case where $\sigma_{m+1} > 0$. As shown in the following lemma, the rejection probability $\bP_0[|T_m|>c]$ can be written as an integral. This not only facilitates numerical computation, but is also crucial for our later theoretical investigation.

lemSuppose that $\sigma_{m+1} > 0$, and define $\gamma_i \equiv \frac{\sigma_i}{\sigma_{m+1}}$ for $i = 1, 2, \ldots, m$. The rejection probability $\bP_0[|T_m|>c]$ can be written as: \begin{align} \bP_0[|T_m|>c] = \overline{p}_m(c; \gamma_1, \ldots, \gamma_m) \equiv \frac{1}{\pi} \int^{|\theta_{m+1}|}_0 \frac{s^{\frac{m-1}{2}}}{ [- g_c(-s)]^{\frac{1}{2}} } \ \mathrm{d}s. \end{align} where $\theta_{m+1} \in [-m - \max_{1\le i \le m}\gamma_i^2, \ -m]$ is the unique negative root of $g_c(\theta)$ defined below, \begin{align} g_c(\theta) \equiv -(m+ \theta) \prod_{i=1}^m ( \kappa\gamma_i^2 - \theta) + \left( \kappa + \frac{\kappa+1}{m} \theta\right) \cdot \sum_{i=1}^m \left[ \gamma_i^2 \prod_{j\ne i} ( \kappa\gamma_j^2 - \theta ) \right], \end{align} and $ \kappa \equiv \frac{m c^2}{m-1}$.

From Lemma (ref), when $\sigma_{m+1}>0$, the rejection probability in (ref) depends only on the ratios between the control standard deviations and the treated standard deviation. The relative heterogeneity assumption in Assumption (ref) equivalently assumes that at least $(m-k+1)$ of $\{\gamma_j\}^m_{j=1}$ is greater than or equal to $\rho^{-1}$. To find the maximum rejection probability, we need to solve an $m$-dimensional optimization over $(\gamma_1, \ldots, \gamma_m) \in [0, \infty)^{k-1} \times [\rho^{-1}, \infty)^{m-k+1}$.\footnote{Note that the rejection probability is invariant to any permutations of the $\{\gamma_j\}^m_{j=1}$. We can, for example, assume that the last $(m-k+1)$ of them is no less than $\rho^{-1}$ without loss of generality.} Optimizing this integral directly can be computationally challenging even for a moderate $m$. As demonstrated in the next subsection, for any $m\ge 2$, we can substantially simplify the $m$-dimensional optimization problem into multiple one-dimensional optimization problems. The reformulated problem can often be efficiently solved numerically.

Reformulating the problem of computing the maximum rejection probability

In this subsection, we discuss how to reduce the optimization for the maximum rejection probability in (ref) into multiple one-dimensional optimization problems. We summarize the key ideas here and defer the technical lemmas to Appendix (ref) of this main text.

First, Lemma (ref) shows that the maximum rejection probability in (ref) must be achieved at some finite values of $(\sigma_1, \ldots, \sigma_m, \sigma_{m+1}) \in \mathcal{S}_m(k, \rho)$. This means we do not need to worry about the boundary case where some of $\{\sigma_j\}_{j=1}^{m+1}$ approach infinity. Next, Lemma (ref) shows the necessary conditions for any $(\sigma_1, \ldots, \sigma_m, \sigma_{m+1}) \in \mathcal{S}_m(k, \rho)$ to be a maximizer. We summarize the implications of Lemma (ref) in the following theorem.

thmFor any given $m\ge 2$, $k \in \{1, \ldots, m\}$, $\rho \ge 0$, $c>0$ and $c\ne m^{-1/2}$, the maximum rejection probability in (ref) under Assumption (ref) must be attained at some $\{\sigma_j\}_{j=1}^{m+1} \in \mathcal{S}_m(k, \rho)$ that satisfy one of the following forms:\footnote{The result in (i) follows from Bakirov:2006aa and implies the maximum rejection probability when $\sigma_{m+1}=0$ as shown in (ref).} \begin{itemize} • $\sigma_{m+1} = 0$, and for $1\le j \le m$, $\sigma_j$ is either $0$ or some common value; • $\sigma_{m+1} = \rho$, and for $1\le j \le m$, $\sigma_j$ is either $0$, or $1$, or some common value. \end{itemize}

Note that $0$ is always a boundary value of $\sigma_j$ for $1\le j \le m$, and $1$ is also a boundary value for some $\sigma_j$ when $\sigma_{m+1} = \rho$ and Assumption (ref) holds. Therefore, Theorem (ref) essentially indicates that, at the maximizer of the rejection probability in (ref), each $\sigma_j$ must either lie on the boundary or share a common value for all $1 \le j \le m$. Importantly, Theorem (ref) explains why we can simplify the optimization for the rejection probability into multiple one-dimensional optimization problems. This is because for all possible maximizers shown in Theorem (ref), there is only one unknown value, which is the common value of the $\{\sigma_j\}^m_{j=1}$ that are not on the boundary. Note that the maximum rejection probability in case (i) with $\sigma_{m+1} = 0$ can be efficiently computed as shown in (ref). In the following, we will therefore focus on the optimization for the rejection probability in case (ii) with $\sigma_{m+1} > 0$.

Specifically, under Assumption (ref) with any $k \in \{1, \ldots, m\}$ and $\rho > 0$, if $(\sigma_1, \ldots, \sigma_m,$ $\sigma_{m+1}) \in \mathcal{S}_m(k, \rho)$ is the maximizer for the rejection probability and $\sigma_{m+1}>0$, then, for all $1\le j \le m$, $\gamma_j \equiv \frac{\sigma_j}{\sigma_{m+1}}$ is either on the boundary (equals $0$ or $\rho^{-1}$ ), or takes some common value $\gamma \ge 0$. Motivated by this, with a slight abuse of notation, we introduce $\overline{p}_m(c; \rho, \gamma; m_1, m_0)$ to denote the value of $\overline{p}_m(c; \gamma_1, \ldots, \gamma_m)$ in (ref) when $m_1$ of $\{\gamma_j\}_{j=1}^m$ equal $\rho^{-1}$, $m_0$ of them equal zero, and the remaining equal a common value $\gamma$. That is,

align[align omitted — 195 chars of source]

where $0 \leq m_0, m_1 \leq m$ and $m_0 + m_1 \leq m$.

Next, define the following as the supremum of $\overline p_m(c; \rho, \gamma; m_1, m_0)$ over $\gamma$, with the support of $\gamma$ depending on $m_1$ and $k$:

align[align omitted — 178 chars of source]

where $\underline\rho = 0$ if $m_1 \ge m-k+1$ and $\underline\rho = \rho^{-1}$ if $m_1 < m-k+1$.

With the above notations, we now state our main theorem of how to evaluate the maximum rejection probability for a given critical value $c$, heterogeneity parameters $(k, \rho)$ and the number of clusters $m$.

thmFor any given $m\ge 2$, $k \in \{1, \ldots, m\}$, $c>0$ and $c\ne m^{-1/2}$, the maximum rejection probability in (ref) under Assumption (ref) can be written as\footnote{ From Remark (ref), $p_m(c; k, \rho) = 1$ when $0 < c < m^{-1/2}$. For convenience, we also define $p_m(c; k, \rho) = 1$ at $c = 0$ and $c = m^{-1/2}$, so that $p_m(c; k, \rho)$ upper bounds the rejection probability for all $c \ge 0$. } \begin{align*} p_m(c; k, \rho) & \equiv \begin{cases} \max\Big\{ p_{m,0}(c), \max_{0\le m_0 \le k-1, 0\le m_1 \le m-m_0} \widetilde{p}_m(c; k, \rho; m_1, m_0) \Big\} & if $\rho > 0$ \\ p_{m,0}(c) & if $\rho = 0$ \end{cases}, \end{align*} where $p_{m,0}(c)$ is defined in (ref) and $\widetilde{p}_m(c; k, \rho; m_1, m_0)$ is defined in (ref).

From Theorem (ref), the key to obtaining the maximum rejection probability is to solve the optimization problem in (ref) for all combinations of $(m_1, m_0)$. Importantly, for each given $(m_1, m_0)$, the optimization in (ref) is an one-dimensional optimization problem, which is computationally much simpler than the original optimization in (ref), and we can solve it using numerical optimization. In particular, for any given $k \in \{1, \ldots, k\}$, we need to solve at most $\frac{1}{2}k(2m+1-k)$ one-dimensional optimization problems\footnote{When $m_1+m_0 = m$, (ref) no longer depends on $\gamma$, and thus no optimization over $\gamma$ is needed.}, which is $m$ when $k=1$, $2m-1$ when $k=2$, \ldots, and $\frac{1}{2}m(m+1)$ when $k=m$. Moreover, at a given $m$, $\rho \geq 0$ and $c > 0$, to find the maximum rejection probability over all $k \in \{1, \ldots, m\}$, we need to solve at most $m(m+1)$ one-dimensional optimization problems of the form (ref), because the optimizations required for different values of $k$ overlap with each other.

We illustrate the above theorem through the following examples. Example (ref) is about $k = 1$ and Example (ref) is about $k = 2$.

egLet $m=6$, $\rho = 2$, $\alpha = 0.05$, $k = 1$, and $c = t_{m-1, 1-\frac{\alpha}{2}} \sqrt{\rho^2 + \frac{1}{m}}$ where $t_{m-1, 1 - \frac{\alpha}{2}}$ is the $(1- \frac{\alpha}{2})$-th quantile of the $t$-distribution with $(m-1)$ degrees of freedom. We show the various functions from Theorem (ref) in Figure (ref). \begin{figure}[!ht] \caption{This above shows $\overline{p}_m(c; \rho, \gamma; m_1, m_0)$ against $\gamma$ for various values of $m_1$ with $m = 6$, $k = 1$ and $m_0 = 0$. The vertical dashed line represents $\gamma = \rho^{-1}$. The horizontal dotted line represents $p_{m,0}(c)$. See Theorem (ref) for the definitions of the functions.} \end{figure} A few observations from Figure (ref) are as follows. First, the curves show $\overline{p}_m(c; \rho, \gamma; m_1, m_0)$ as defined in (ref) against $\gamma$, for $m_1= 0, \ldots, 5$. Recall this function means that among $\{\gamma_j\}^6_{j=1}$, $m_1$ of them equals $\rho^{-1}$ and the remaining of them equals $\gamma$. We set $\gamma = e^x$ in order to display the behavior of large $\gamma$. Each colored curve corresponds to a specific value of $m_1$. It can be seen that $\overline{p}_m(c; \rho, \gamma; m_1, m_0)$ decreases as $\gamma$ increases for each $m_1$. The black dotted horizontal line plots $p_{m,0}(c)$ defined in (ref). It shows the maximum rejection probability when $\sigma_{m+1}=0$. Here, it is much smaller than 0.05. As $\gamma \longrightarrow \infty$, $\overline{p}_m(c; \rho, \gamma; m_1, m_0)$ converges to values less than or equal to $p_{m,0}(c)$. This is not surprising, as in such scenarios, $\sigma_{m+1}$ is much smaller than some control standard deviations, essentially approximating the case where $\sigma_{m+1} = 0$.
egConsider the same setup as in Example (ref) but with $k = 2$ and set the value of $c$ such that the maximum rejection probability equals $\alpha = 0.05$. Figure (ref) shows the rejection probability under different combinations of $\gamma$, $m_0$ and $m_1$, where $m_0$ here can only take the values 0 or 1. \begin{figure}[!ht] \caption{This above shows $\overline{p}_m(c; \rho, \gamma; m_1, m_0)$ against $\gamma$ for various $m_0$ and $m_1$ with $m = 6$ and $k = 1$. See the caption for Figure (ref) for more details.} \end{figure} As can be seen from the figure, the maximum rejection probability is achieved when one of the $\{\gamma_j\}^m_{j=1}$ equals 0 and the remaining $(m-1)$ terms from $\{\gamma_j\}^m_{j=1}$ equals $\rho^{-1}$. For instance, in the $m_0 = 0$ panel, the purple line represents $m_1 = 5$ terms of $\{\gamma_j\}^m_{j=1}$ equals $\rho^{-1}$. For this line, the maximum rejection probability is achieved when the remaining $m - m_1 = 1$ term equals $\gamma = 0$. Similarly, in the $m_0 = 1$ panel, exactly one of $\{\gamma_j\}^m_{j=1}$ is equal to 0. Each of the lines shows that the maximum rejection probability is achieved when all the remaining terms equal $\rho^{-1}$.
reAs discussed in Appendix (ref), when $m\longrightarrow \infty$, under certain regularity conditions, the maximum rejection probability $p_m(c; k, \rho)$ for any $c, k$ and $\rho>0$ is achieved when $\sigma_{m+1} = 1$, $(k-1)$ of $\{\sigma_j\}_{j=1}^m$ are $0$, and the remaining $(m-k+1)$ of $\{\sigma_j\}_{j=1}^m$ are equal to $\rho^{-1}$. This is indeed the case for both Examples (ref) and (ref). Based on this intuition, we can first use this configuration of $\{\sigma_j\}_{j=1}^{m+1}$ to get a candidate threshold $c$ such that $\bP_0 [|T_m| > c]=\alpha$, where $\alpha$ is the significance level of interest. We can then verify whether this candidate threshold $c$ achieves the desired type-I error control by solving the optimization in Theorem (ref).

Finally, we report the critical values for different numbers of control clusters $m$ and heterogeneity parameters $\rho$ for $\alpha = 0.01$ and $\alpha = 0.05$ in Table (ref) when $k = 1$. We report the critical values for $k = 2$ in the supplementary material.

table[table omitted — 3,213 chars of source]

Closed-form valid $t$-test when $k = 1$

In this subsection, we consider the case where $k = 1$ in Assumption (ref). This means that $\sigma_{m+1} \leq \rho \sigma_j$ for $j = 1, \ldots, m$, i.e., the treated standard deviation is smaller than or equal to $\rho$ times each of the control standard deviations.

In this case, we can obtain closed-form solutions for the maximum rejection probability when the threshold $c$ is large enough. Equivalently, this corresponds to testing at a significance level that is not “too large.” The theorem is stated as follows:

thmFor any given $m \ge 4$, $\rho > 0$, $c > \sqrt{\frac{3(m-1)}{m(m-3)}}$, define: \begin{align*} \overline H_m(c, \rho) & \equiv \max\left\{ \frac{3(m\rho^2 + 1)}{ m\rho^2 + \kappa + 1}, \frac{2\kappa + 3}{\kappa + 1} \right\} \notag \\ & \qquad + \frac{1-\tau}{1-\tau +\min\{ (1 - 2\tau)\kappa \underline Z - \frac{1}{2} , 0\}} -\frac{m\kappa}{m\rho^2 + \kappa+1} - 1, \end{align*} where $\kappa \equiv \frac{m c^2}{m-1}$ and $\tau \equiv \frac{\kappa+1}{m\kappa}$ are determined by $(m, c)$, and $\underline{Z} \equiv \frac{1}{2 \cdot \max\{m\rho^2 + 1, \kappa + 2\}}$ is determined by $(m, c, \rho)$. Under the above conditions and notations, the following statements hold. \begin{itemize} • $\overline H_m(c, \rho)$ is decreasing in $c$, and $\lim_{c\to \infty} \overline H_m(c, \rho) < 0$. • Let $\underline{c}_{m, \rho} \equiv \inf\{c > \sqrt{\frac{3(m-1)}{m(m-3)}}: \overline H_m(c, \rho) \le 0 \}$, which must be finite. Then, for any $c\ge \underline{c}_{m, \rho}$, the maximum rejection probability $p_m(c; 1, \rho)$ under Assumption (ref) with $k=1$ and the given $(\rho, m, c)$ has the following equivalent form: \begin{align*} p_m(c; 1, \rho) = \bP \left[ |t_{m-1}|\sqrt{\rho^2+\frac{1}{m}} > c \right]. \end{align*} \end{itemize}

Theorem (ref) gives a closed-form solution for the maximum rejection probability under the relative heterogeneity assumption with $k=1$ and any given $\rho > 0$. The maximum rejection probability is achieved when $\sigma_1 = \sigma_2 = \ldots = \sigma_m = \rho^{-1}$ and $\sigma_{m+1} = 1$. Importantly, for any given $m$ and $\rho>0$, the cutoff $\underline{c}_{m, \rho}$ can be easily computed numerically, due to the monotonicity of the function $\overline H_m(c, \rho)$ in $c$. Accordingly, Theorem (ref) gives a closed-form critical value of our $t$-test for any significance level less than or equal to

align[align omitted — 160 chars of source]

That is, for any significance level $\alpha \in (0, \underline{\alpha}_{m, \rho}]$, a valid critical value for our two-sided $t$-test is $\sqrt{\rho^2 + m^{-1}} \ t_{m-1,1- \frac{\alpha}{2}}$, where $t_{ m-1,1- \frac{\alpha}{2}}$ is the $(1-\frac{\alpha}{2})$ quantile of the $t$-distribution with degree of freedom $m-1$.

Table (ref) reports the largest significance level $\underline{\alpha}_{m, \rho}$ such that our two-sided $t$-test has a simple closed-form expression for the critical value, under Assumption (ref) with $k=1$ and various values of $(m, \rho)$. Note that our $t$-test can handle all values of $m\ge 2$, $k \in \{1, \ldots, m\}$ and $\rho \ge 0$, but it generally involves one-dimensional optimization as described in Section (ref). We can conduct a similar theoretical investigation to simplify the optimization of the rejection probability under Assumption (ref) with $k \geq 2$. However, in this case, the potential maximizers cannot be reduced to a single point, so one-dimensional numerical optimization is still required. We relegate the detailed discussion under Assumption (ref) with $k\geq 2$ to the supplementary material.

table[table omitted — 781 chars of source]

Power of the t-test

We now investigate the power of the proposed $t$-test under the alternative hypothesis $\overline{H}_1$ in (ref) with $\delta \ne 0$. Without loss of generality, we assume that $\delta>0$. The theorem below gives a lower bound for the power of the $t$-test.

thmSuppose that Assumption (ref) holds for some $\delta>0$, and define the $t$-statistic $T_m$ as in (ref). Then, for any finite $c>0$, \begin{align*} \bP[|T_m|>c] \ge \bP[T_m>c] \ge 1 - \frac{1}{\delta^2} \left[ \sigma^2_{m+1} + \frac{2(c^2 + m^{-1})}{m} \sum_{j=1}^m \sigma_j^2 \right]. \end{align*}

The lower bound of power in Theorem (ref) increases with $\delta$, where $\delta$ corresponds to $\sqrt{n}(\mu_1 - \mu_0)$ in our large-sample inference for a finite number of clusters as shown in Theorem (ref). If the gap between the means of the treated and control clusters is bounded away from zero, then, as the sample size $n\longrightarrow \infty$, the power of our $t$-test with any finite critical value will converge to $1$. In addition, Theorem (ref) also provides a rough power estimate when we have some information about the treatment effect size and the variances for the treated and control clusters.

Simultaneous inference

Recall that Assumption (ref) involves two parameters $\rho$ and $k$. They represent the allowable degree of relative heterogeneity between treated and control clusters. Specifically, larger values of $\rho$ and $k$ correspond to a greater degree of allowable heterogeneity. In practice, specifying the values of $(\rho, k)$ might be challenging, and it is often desirable to perform multiple analyses for a wide range of values for $(\rho, k)$. This raises the question of how to interpret the results of the analysis under different values of $(\rho, k)$. In this section, we show that the analyses over all possible values of $(\rho, k)$ can be simultaneously valid, without the need of any adjustment due to multiple analyses. Moreover, the analysis results can be easily visualized and interpreted. Similar simultaneous inference has been used for sensitivity analysis of observational studies cuili2025wp,wuli2025jasa.

We first introduce several notations to denote the true relative heterogeneity of the standard deviations between treated and control clusters. For any $k \in \{1, \ldots, m\}$, define

align[align omitted — 396 chars of source]

where $\sigma_{m+1}$ and $\sigma_{(k)}$ in (ref) denote the true standard deviation of the treated cluster and that of the control cluster at rank $k$. Consequently, the values of $\rho_k^\star$ reflect the true relative heterogeneity between the treated and control clusters. In particular, larger values of $\rho_k^\star$ indicate greater heterogeneity between the treated and control clusters.

We now explain the key idea for our simultaneous inference procedure. In Section (ref), we test the null hypothesis in (ref) of no mean difference between the treated and control clusters under Assumption (ref) for some $k \in \{1, \ldots, m\}$ and $\rho \ge 0$, which is equivalent to that $\rho_k^\star \le \rho$ with $\rho_k^\star$ defined as in (ref). Importantly, we can reinterpret the $t$-test in Algorithm (ref) as a valid test for the null hypothesis of $\rho_k^\star \le \rho$ about the true relative heterogeneity in (ref), under the assumption of no mean difference between treated and control clusters (i.e., $\delta = 0$). By standard test inversion, we can construct confidence intervals for $\rho_k^\star$, which must be one-sided confidence intervals with unbounded right endpoints and thus provide essentially lower bounds on the true relative heterogeneity, under the assumption that $\delta=0$. Moreover, these confidence intervals will be simultaneously valid across all $k \in \{1, \ldots, m\}$. We summarize the results in the following theorem, followed by a discussion of its practical implications and interpretation. For any $z\in \mathbb{R}$, we write $(z, \infty] \equiv (z, \infty)\cup \{\infty\}$ and analogously $[z, \infty] \equiv [z, \infty)\cup \{\infty\}$.

thmLet $\alpha \in (0, 1)$. Suppose Assumption (ref) holds with $\delta = 0$. Let $T_m$ be the test statistic as defined in (ref). \begin{itemize} • For any $k \in \{1, \ldots, m\}$, an $(1-\alpha)$-confidence set for $\rho^\star_k$ in (ref) is \begin{align} \mathcal{I}_{m, \alpha, k} \equiv \{ \rho \ge 0: p_m(|T_m|; k, \rho) > \alpha \} \cup \{\infty\} = (\hat{\rho}_{m, \alpha, k}, \infty ] or [\hat{\rho}_{m, \alpha, k}, \infty ], \end{align} which must be an one-sided confidence interval with $\hat{\rho}_{m, \alpha, k} \equiv \inf \mathcal{I}_{m, \alpha, k}$. • The confidence intervals in (i) are simultaneously valid across all $k \in \{1, \ldots, m\}$, in the sense that \begin{align*} \bP\left[ \rho^\star_k \in \mathcal{I}_{m, \alpha, k} for all k \in \{1, \ldots, m\} \right] \ge 1 - \alpha. \end{align*} \end{itemize}

The interpretation of Theorem (ref) is as follows. If there is no mean difference between treated and control clusters (i.e., $\delta=0$), then we need to believe that, with $(1-\alpha)$ confidence level, the relative heterogeneity $\rho_k^\star$ in (ref) must be greater than (or equal to) $\hat{\rho}_{m, \alpha, k}$, for all $k \in \{1, \ldots, m\}$. That is, with $(1-\alpha)$ confidence level, the treated standard deviation $\sigma_{m+1}$ must be at least $\hat{\rho}_{m, \alpha, k}$ times the control standard deviation $\sigma_{(k)}$ at rank $k$, for all $k \in \{1, \ldots, m\}$. If we question any of these statements on the true relative heterogeneity, then the assumption of $\delta=0$ is likely to fail, or equivalently there is likely significant mean difference between the treated and control clusters. Obviously, the larger the values of $\{\hat{\rho}_{m, \alpha, k}\}^m_{k=1}$, the stronger the evidence for a nonzero treatment effect. In practice, we can easily visualize these confidence intervals by, say, plotting $k$ against $\hat{\rho}_{m, \alpha, k}$; this is illustrated in the two empirical applications in Section (ref) ahead.

We now discuss the computation for the confidence intervals in Theorem (ref). By the definition of the $p$-value $p_m(c; k, \rho)$ in (ref) and the fact that the set $\mathcal{S}_m(k, \rho)$ in (ref) increases as $k$ or $\rho$ increase, $p_m(c; k, \rho)$ must be nondecreasing in both $k$ and $\rho$. The monotonicity in $\rho$ then explains the equivalent one-sided form of the set in (ref). Thus, we can use the bisection method to find the thresholds $\{\hat{\rho}_{m, \alpha, k}\}^m_{k=1}$. In addition, the monotonicity in $k$ implies the threshold $\hat{\rho}_{m, \alpha, k}$ is nonincreasing with $k$. Hence, we can first find $\hat{c}_{m, \alpha, 1}$, and then sequentially use $\hat{c}_{m, \alpha, k-1}$ as an upper bound for $\hat{\rho}_{m, \alpha, k}$ in the bisection method, for $2\le k \le m$.

rehagemann2024wp also reported thresholds like $\{\hat{\rho}_{m, \alpha, k}\}$ for his test. However, his test only works when $k=1$ or 2, while our test works for any $k \in \{1, \ldots, m\}$. In addition, as shown in Theorem (ref), these thresholds can be interpreted as lower confidence bounds for the true relative heterogeneity, and they are simultaneously valid, indicating that the confidence bounds for all $k$ are indeed “free lunch” added to those for $k=1$ or $2$.

Simulations

In this section, we consider two sets of numerical exercises to compare the performance of the $t$-test against other methods for conducting inference with a single treated cluster.

Simulation design 1: normal means

The first simulation design generates data using normal distributions. The data generating process (DGP) is as follows. We generate normal random variables as in Assumption (ref). In particular, $\{\psi_j\}^m_{j=1}$ are normally distributed with mean 0, $\psi_{m+1}$ has mean $\delta$, $\sigma_{m+1}^2 = \rho^2$, and $\{\sigma_j^2\}^m_{j=1}$ are specified below.

We consider two different DGPs to generate the random variables for the control clusters.

description$\sigma_j^2 = 1$ for $j = 1, \ldots, m$. • $\sigma_j^2 = 1 + \frac{j - 1}{m - 1}$ for $j = 1, \ldots, m$.

DGP 1 is the homogeneous design in which all control random variables have the same varaiance. We introduce heterogeneity in DGP 2 and allow the variance to vary between 1 and 2. The goal is to test the two-sided hypothesis (ref), under various heterogeneity parameter $\rho$, cluster size $m$, and significance level $\alpha$. We compare the performance of our $t$-test with hagemann2024wp in terms of size and power. This is because both of our tests work with a single treated cluster and a finite number of control clusters, and are valid under a certain relative heterogeneity assumption.

In the simulations, we consider $\rho$ from 0.1 to 2 with step size 0.1, $m \in \{5, 10, 25, 50\}$, and $\alpha \in \{0.01, 0.05, 0.1\}$. We focus on two-sided tests and choose $\rho$ in our $t$-test and hagemann2024wp so that it matches the one used in the DGP, i.e., the relative heterogeneity assumption is always correctly specified. The results are based on 500,000 Monte Carlo replications. Figure (ref) shows the result for $\alpha = 0.05$ under DGP 1 when $k = 1$. The figure plots the rejection rate against various values of $\rho$. Each facet represents a specific combination of $m$ and $\delta$. $\delta = 0$ corresponds to the results under the null. $\delta > 0$ corresponds to the results under the alternative. Note that hagemann2024wp's result does not appear or only partially appear in some of the facets. This is because hagemann2024wp may not necessarily be able to find a weight for his rearrangement test such that his test can be shown to be valid for some combinations of $m$, $\rho$ and $k$.

Figure (ref) shows that our test performs favorably when compared to hagemann2024wp. Both of our tests control size. The $t$-test is more powerful, especially when $\rho$ is small. As $\rho$ increases above 1, the power difference decreases, but we are still more powerful.

figure[figure omitted — 335 chars of source]

Next, Figure (ref) reports the results for DGP 2 that includes more heterogeneity among the control random variables for $\alpha = 0.05$ and $k=1$. As predicted by the theory, our $t$-test becomes more conservative when more heterogeneity are included. Our $t$-test continues to control size and has power against the alternative. Moreover, our test outperforms hagemann2024wp's test in most cases. In the supplementary material, we report the results for $k=1$ and $k=2$ for both DGP at various significance levels.

figure[figure omitted — 340 chars of source]

Simulation design 2: two-way fixed effects

Next, we conduct a simulation exercise with two-way fixed effects as in Example (ref). We consider a design that is based on the ones used in conleytaber2011restat and hagemann2024wp. As before, let $\cJ_0 \equiv \{1, \ldots, m\}$ be the control clusters and $\cJ_1 \equiv \{m+1\}$ be the treated cluster. Let $\cT \equiv \{1, \ldots, 10\}$ be the total number of time periods. Let $t_0 = 6$ be the intervention period and $D_{jt} \equiv \ind[j \in \cJ_1 \text{ and } t > t_0]$ for $j \in \cJ$ and $t \in \cT$. Each simulated data is generated from the following two-way fixed effects model

align[align omitted — 139 chars of source]

where $\alpha_t = 1$ for all $t \in \cT$ and $\gamma_j = 2 \ind[j \leq \frac{m}{2}] - 1$ for all $j \in \cJ$. We test the null of $\theta = 0$ and consider $\theta \in \{0, 1, 2, 3\}$ when generating the data. We consider the following DGPs based on model (ref). The first four DGPs are the same as the ones considered in hagemann2024wp. The remaining two DGPs are based on the error distributions considered in conleytaber2011restat.

description• Generate $U_{jt} = \eta U_{j, t-1} + \sigma^{\ind[j = m + 1]} V_{jt}$ where $\eta = 0.5$ and $V_{jt}$ are independently distributed standard normal random variables. • Same as DGP 1, but use $\eta = 0.1$. • Same as DGP 1, but use $\eta = 0.9$. • Same as DGP 1, but $V_{jt}$ follows a normalized $\chi_2^2$ distribution with mean 0 and variance 1. • Same as DGP 1, but use $V_{jt} \sim \text{Uniform}[-\sqrt{3}, \sqrt{3}]$.

For each DGP, we are interested in testing the two-sided hypothesis of \[ H_0: \theta = 0 \quad \text{ vs } \quad H_1 : \theta \neq 0. \] In addition, we consider simulations with $m \in \{5, 10, 25, 50\}$, $ \sigma \in \{0.1, 0.5, 1, 2\}$, and $\alpha = 0.05$. We report the rejection rate curves based on 5,000 Monte Carlo simulations. For each DGP, we consider conducting inference using the following four methods:

description• The $t$-test proposed in our paper. • The rearrangement test by hagemann2024wp. • The procedure of conleytaber2011restat that assumes homogeneity. • The bootstrap procedure of fermanpinto2019restat.

Figures (ref) and (ref) show the results for DGPs 1 and 2 under $\alpha = 0.05$ across different numbers of clusters $m$ and true relative heterogeneity $\sigma$. We set the relative heterogeneity parameter $\rho$ for both our and hagemann2024wp's method in all the DGPs to be the corresponding $\sigma$, so that the relative heterogeneity assumption is correctly specified. We also incorporate the correct value of $\sigma$ in applying FP to specify the heterogeneity.

figure[figure omitted — 226 chars of source]
figure[figure omitted — 227 chars of source]

It can be seen that our $t$-test performs well and compares favorably to other methods. In particular, the CT test exhibits inflated Type I error rates when the number of control clusters $m$ is small and $\sigma$ is large. This inflation arises because the validity of CT relies on the assumption of an infinite number of clusters and homogeneous variances between the treated and control clusters. The FP test also shows an inflated Type I error when $m$ is small, with the inflation becoming more pronounced as $\sigma$ decreases. This again reflects its reliance on the asymptotic validity under an infinite number of clusters. Notably, we assume that FP has access to the true heterogeneity between the treated cluster and each control cluster, which is a stronger assumption than those required by our $t$-test or hagemann2024wp's rearrangement test. Both our test and that of hagemann2024wp successfully control Type I error. In contrast, our test is applicable regardless of the number of control clusters and demonstrates higher power across different levels of heterogeneity. The results for DGPs 3 to 5 are contained in the supplementary material.

Empirical applications

In this section, we illustrate the $t$-test proposed in this paper with two recent empirical applications. For each of the empirical applications, we report two sets of results. First, we try to find the largest $\rho$ such that the null hypothesis is rejected when $k = 1$ among various regression specifications of the empirical examples. Among those specifications where the null can be rejected at some $\rho$, we conduct simultaneous inference as discussed in Section (ref).

Empirical application 1: depewswensen2022ej

depewswensen2022ej examines the impact of the 1911 New York State Sullivan Act on mortality rate. The Sullivan Act required New York citizens to obtain a permit and license to carry concealable weapon. We consider the following two-way fixed effects model:

equation[equation omitted — 142 chars of source]

where $\text{Treated}_{jt}$ equals 1 if state $j$ is New York and $t$ is the post-treatment period, $\alpha_t$ is the year fixed effects, $\gamma_j$ is the state fixed effects, and $\epsilon_{jt}$ is the residual term. They consider the following four outcome variables: homicide rate, suicide rate, gun suicide rate, and non-gun suicide rate.

They report cluster-robust standard errors and $p$-value from the Wild cluster bootstrap. There is only nine control clusters in this empirical application. Hence, hagemann2024wp's test may not be applicable because there can be no weights such that his rearrangement test is valid for some relative heterogeneity parameters $\rho$.

We are interested in testing

equation[equation omitted — 127 chars of source]

Table (ref) summarizes the estimation and inference results of the baseline model described in (ref) for the four outcomes described above. The point estimates and the Wild cluster bootstrap $p$-values are taken from Table 2 of depewswensen2022ej.\footnote{The point estimates and $p$-values that depewswensen2022ej report are based on weighted least squares.}

table[table omitted — 970 chars of source]

The remaining rows of Table (ref) report the largest $\rho$ such that the null can be rejected for significance levels $\alpha = 0.01$, $0.05$, and $0.1$. For entries with “NA” in the table, it refers to the situation where there does not exist a $\rho \geq 0$ such that the null is rejected. If we focus on the gun suicide rate under $\alpha = 0.1$, it means that the null can be rejected when the variance in New York is at most $1.3^2 \approx 1.69$ larger than the smallest variance in the control clusters. The true relative level of heterogeneity such that the null can be rejected may be related to the population size of the various states. In particular, New York state has a larger population than the other control states in 1911 fred2025website: the population of New York state was 9.249 million, and the largest control state was Massachusetts with 3.383 million and the smallest control state was Vermont with 0.358 million.

Next, we conduct simultaneous inference for our $t$-test as in Section (ref). We focus on the outcome “gun suicide rate” as we can find $\rho$ such that the null is rejected when $k=1$. Figure (ref) reports the results for $\alpha \in \{0.01, 0.05, 0.1\}$ on the range of $\rho$ for various $k$ such that the null can be rejected. The shaded area represents the simultaneous confidence region for true relative heterogeneity when there is no treatment effect.

figure[figure omitted — 235 chars of source]

Empirical application 2: hiraiwaetal2024wp

hiraiwaetal2024wp examine whether firms enforce noncompete agreements (NCAs). NCAs are restrictions that prevent workers from joining or starting competing firms. In 2020, Washington state passed a law that prohibits NCAs for workers earning below a certain threshold. The threshold was \$100,000 per year in 2020 and \$101,390 per year in 2021. They study the impact of being above or below the threshold on employment for different earnings bins. They consider the following two-way fixed effects model in their panel B of Table 1: \[ \log \text{Emp}_{b,t} = \beta \text{Treated}_{b,t} + \alpha_t + \gamma_b + \epsilon_{b,t}, \] where $\text{Emp}_{b,t}$ is the employment count of bin $b$ at year $t$, $\text{Treated}_{b,t}$ equals 1 for the focal bin in the focal year, $\alpha_t$ and $\gamma_b$ are the year and bin fixed effects, and $\epsilon_{b,t}$ is the error term. They have 30 clusters (income bins). They use one-sided randomization inference and found that there is no significant effect.

We apply our method and hagemann2024wp. We are interested in testing the two-sided hypothesis as in (ref). Table (ref) summarizes our results. Each column contains the result for a specific focal year and definition of the treatment variable. For instance, column 2 refers to focal year 2020, and the treatment variable is defined to be equal to 1 when the income bin is just above the threshold, i.e., \$100-101.389k. The remaining columns are defined similarly. For each regression specification, we find the largest $\rho$ such that the null is rejected. For entries with “NA” in the table, it refers to the situation where there does not exist a $\rho \geq 0$ such that the null is rejected.

table[table omitted — 1,189 chars of source]

Next, we conduct simultaneous inference for our $t$-test as in the first empirical application. We focus on the two variables for focal year 2021 because we can find $\rho$ such that the null is rejected when $k=1$. In particular, we report the range of $\rho$ for various $k$ as long as there exists $\rho$ such that the null can be rejected. Figure (ref) reports the results for $\alpha \in \{0.01, 0.05, 0.1\}$.

figure[figure omitted — 254 chars of source]

Conclusion

In this paper, we propose a $t$-test to conduct inference with a single treated cluster under a certain relative heterogeneity assumption. The $t$-statistic and the associated critical values are easy to calculate in many empirically relevant applications. We show that our test performs favorably when compared to other methods in the literature. We also show that our test is valid with weaker assumptions.