Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
87,279 characters · 23 sections · 72 citation commands
Cluster-robust inference with a single treated cluster using the t-test
\fontfamily{ppl}
\onehalfspacing
In difference-in-differences designs, it is common for researchers to conduct inference using cluster-robust methods to account for correlation within clusters. However, inference becomes challenging when there is only a single treated cluster. Intuitively, this is because researchers are faced with only one estimate from the treated cluster, making its uncertainty difficult to quantify. Existing methods either impose strong assumptions on having a large number of clusters, require the variances to be homogeneous or estimable, or only work for specific significance levels depending on the number of clusters. In this paper, we develop a $t$-test associated with a suitable critical value to conduct valid inference that relaxes these assumptions.
Having a single treated cluster and a finite number of control clusters is common in empirical work. For instance, this can happen when researchers use data in the United States to examine the impact of a policy that takes place in one particular state but not the other states. The nearby states or the remaining states are commonly used as the control group. In these scenarios, it is reasonable for researchers to believe they are not in the scenario with a “large” number of control clusters. Some recent examples in this context include wangburke2022jfe on the effect of payday loan regulations in Texas, harrislarsen2023jhr on the effect of Hurricane Katrina on student outcomes in New Orleans, dillenderetal2023jpube on the effect of change in health care reimbursement rates in Illinois, alpertetal2024aejep on the impact of Kentucky's prescription drug monitoring programs on opioid prescribing, and kumarliang2024aejep on the labor market effects of constitutional amendments in Texas. hagemann2024wp has documented earlier related examples in the context of a single treated cluster. In the next section of this paper, we demonstrate that our test is also applicable to other empirical designs, in addition to difference-in-differences.
This paper contributes to the literature on cluster-robust inference. We assume that the number of clusters is fixed, as in some related work in this literature such as besteretal2011joe, ibraimovmuller2010jbes, ibraimovmuller2016restat, canayetal2017ecta, hagemann2022wp, and lau2025wp. Although the tests in the aforementioned papers are valid when there is a finite number of clusters, they are not suitable for the problem with a single treated cluster. The first test requires certain homogeneity conditions, and the other tests cannot be applied when there is a single treated cluster. canaysantosshaikh2021restat has showed that Wild cluster bootstrap popularized by cameronetal2008restat can be valid under strong homogeneity conditions when there is a fixed number of clusters. See, for instance, cameronmiller2015jhr, conleyetal2018jar, mackinnonetal2023joe, and alvarezetal2025wp for some surveys on the literature of conducting inference with a fixed number of clusters.
Several tests have been developed to conduct inference when there is a single treated cluster, but with assumptions that can be strong in practice. conleytaber2011restat assume homogeneous errors. fermanpinto2019restat relax the homogeneity assumption and allow for known heteroskedasticity. alvarezferman2023wp relax the homogeneity assumption and allows for spatial correlation. All these tests assume that there is an infinite number of control clusters. Our $t$-test neither assumes the variances are known nor assumes an infinite number of control clusters. Recently, hagemann2024wp develops a novel rearrangement test that relaxes the assumptions mentioned earlier. hagemann2024wp is the most related paper in that his test relies on a relative heterogeneity condition between the treated and control clusters. Both our test and hagemann2024wp work with a fixed number of clusters, allow for arbitrary correlation within clusters, and do not assume that we can estimate or know the cluster variances. He shows his test is valid when the relative heterogeneity condition holds with all or all but one control cluster. However, the validity of his test depends on the number of clusters, the significance level, and the relative heterogeneity condition. For example, when the standard deviation of the treated cluster is bounded by $2$ times all but one of the standard deviations of the control clusters, hagemann2024wp requires at least 14 control clusters to conduct a one-sided test at the 2.5% level using his rearrangement test, and at least 17 control clusters are needed for a 1% level test\footnote{Equivalently, to conduct two-sided tests, there have to be at least 14 control clusters for a 5% level test.}. This potentially limits the applicability of his test. Among the cases where his test is valid, he computes weights for his rearrangement test using numerical optimization that lead to valid tests.
There are several important distinctions between our work and hagemann2024wp. First, we show that our test can be valid for any choice of significance levels and heterogeneity parameters when there are at least two control clusters. For many conventional combinations of significance levels and number of clusters, we derive a closed-form critical value for the $t$-test. No optimization is necessary for these cases, so researchers can apply the test readily. For other cases, we can derive critical values that lead to valid tests through one-dimensional optimizations. Therefore, our test can have broader applicability. Second, we relax the relative heterogeneity condition in hagemann2024wp. Intuitively, the condition restricts the relative variance between the treated and control clusters. hagemann2024wp requires such condition to hold between the treated and all (or all but one) control clusters to show the validity of his test. We weaken this condition by allowing the researcher to impose this restriction between the treated cluster and any number of control clusters. This allows researchers to bound relative heterogeneity on variances depending on their choice on how to restrict the amount of heterogeneity between the treated and control clusters. Third, we allow for simultaneous inference across all the relative heterogeneity assumptions that we impose. Specifically, under the assumption of no treatment effect, we can infer the minimum amount of relative heterogeneity required to explain away the observed association between treatment and outcomes. If the true relative heterogeneity is unlikely to exceed this amount, the treatment is likely to have a significant nonzero effect.
While the $t$-test has been used in Bakirov:2006aa and ibraimovmuller2010jbes, ibraimovmuller2016restat, our proof on the validity of the $t$-test in the current context requires different proof strategies. This is because we only have a single treated cluster and the relative heterogeneity assumption imposes additional structure on the parameter space.
We confirm the performance of our proposed test in simulations. We find that our test controls size under a wide range of significance levels and number of clusters. We also have favorable power performance when compared to other tests in the literature.
The rest of the paper is organized as follows. Section (ref) outlines some popular empirical designs that our test can be applied to. Section (ref) presents our inference procedure. Readers interested in applying the test can refer to Algorithm (ref). Section (ref) presents our main theory. Section (ref) explains how our test can be easily used for simultaneous inference. Section (ref) presents the simulation results. Section (ref) contains two empirical studies. Section (ref) concludes. All proofs can be found in the supplementary material.
We start by describing several empirically relevant designs that are common in applied work and are related to the issue of having a single treated cluster. These examples show that our method applies more broadly apart from standard difference-in-differences designs. The first three examples have also been discussed in hagemann2024wp.
In the following, $i \in \cI \equiv \{1, \ldots, n\}$ indices units, $j \in \cJ \equiv \{1, \ldots, m + 1\}$ indices clusters, and $t \in \cT \equiv \{ 1, \ldots, T\}$ indices time. Only cluster $(m+1)$ is treated, and the remaining clusters are controls. Thus, $\cJ$ can be partitioned as $\cJ_0 \cup \cJ_1$, where $\cJ_0 \equiv \{1, \ldots, m\}$ and $\cJ_1 \equiv \{m + 1\}$ denote the control clusters and treated cluster respectively. Let $D_j \equiv \ind[j \in \cJ_1]$ be a cluster-level treatment indicator across $j \in \cJ$.
As to be described in Section (ref), we require researchers to be able to run regression cluster by cluster, and write the estimator for the target parameter $\Delta$ in terms of the estimator for parameters from cluster-level regressions, denoted by $\{\widehat\theta_j\}^{m+1}_{j=1}$. The following examples explain how to connect $\{\widehat\theta_j\}^{m+1}_{j=1}$ with the estimator for the target parameters in common empirical designs.
For more discussion about recent advances in difference-in-differences and two-way fixed effects, see, for instance, the surveys by dechaisemartindhaultfœuille2023ej, rothetal2023joe, and bakeretal2025wp.
In this section, we present the main assumptions and our algorithm of conducting inference using the $t$-test with a single treated cluster. Readers who are interested in applying our test can directly apply Algorithm (ref). We also present some brief intuition for the validity of our test in Section (ref) to facilitate the theoretical discussion in Section (ref).
We follow the notation used in Section (ref). Let $\{\widehat{\theta}_j\}^{m+1}_{j=1}$ be the cluster-level estimators, $\cJ_0 = \{1, \ldots, m\}$ index the control clusters and $\cJ_1 =\{m + 1\}$ index the treated cluster. For simplicity, we make the dependence of $\{\widehat{\theta}_j\}^{m+1}_{j=1}$ on the cluster size implicit.
Our procedure requires two main assumptions. First, we assume that an appropriate central limit theorem applies to $\{\widehat{\theta}_j\}^{m+1}_{j=1}$ as the sample size within each cluster, denoted by $n$, goes to infinity. A similar assumption is also imposed in related papers on a fixed number of clusters, such as ibraimovmuller2010jbes, ibraimovmuller2016restat, canayetal2017ecta, hagemann2022wp, hagemann2024wp and lau2025wp. Intuitively, this holds when each cluster has a large number of units or consists of a panel with a long time periods. This high-level assumption is formalized as follows.
Below we give a few remarks regarding Assumption (ref). First, it is not necessary for all clusters to have the same size. In cases where clusters have varying sizes, we can define $n$ as the size of the smallest cluster, and our inference essentially requires the sample sizes in all clusters to be large. Second, we assume that the cluster-level estimators are independent, at least asymptotically, so that the asymptotic covariance matrix $\Sigma$ in (ref) is diagonal. This is trivially satisfied if the units are independent across clusters. We emphasize that, while we require independence across clusters, we allow for unknown dependence structure within clusters. Third, we assume that the estimators from the control clusters are consistent for a common parameter $\mu_0$, which can often be justified by model assumptions or study designs as illustrated in Section (ref).
Second, we impose the following relative heterogeneity assumption, which generalizes the relative heterogeneity assumption in hagemann2024wp. Let $\sigma_{(1)} \leq \sigma_{(2)} \leq \cdots \leq \sigma_{(m)}$ be the ordered values of $\{\sigma_j\}^m_{j=1}$ in Assumption (ref).
The above assumption does not require the variances to be known. It only restricts the relative heterogeneity of the standard deviations between the treated and control clusters. In particular, for any $\rho \geq 0$ and $k \in \{1, \ldots, m\}$, it requires the standard deviation of the treated cluster to be less than or equal to $\rho$ times the standard deviations of at least $(m - k + 1)$ control clusters. For example, $k = 1$ requires that the standard deviation of the treated cluster is smaller than or equal to $\rho$ times the standard deviations of all control clusters; when $k=m$, this means that the standard deviation of the treated cluster is smaller than or equal to $\rho$ times the largest standard deviation from the control clusters.
Assumption (ref) is a relaxed version of hagemann2024wp's maximum relative heterogeneity assumption. His assumption is equivalent to Assumption (ref) with $k = 1$ or $k = 2$. Our test allows a general choice of $k \in \{1, \ldots, m\}$. More importantly, as demonstrated in Section (ref), the inference can be simultaneously valid for all $k \in \{1, \ldots, m\}$. In other words, with additional choices of $k$, our test can provide more evidences against the null hypothesis than hagemann2024wp's test, which focuses on $k=1$ or $2$.
Before introducing our algorithm, we end this subsection with some discussion on choosing $(\rho, k)$ in Assumption (ref). If the researcher believes the standard deviation of the treated cluster cannot be larger than the standard deviations of the control clusters, then the researcher can set $k = 1$ and $\rho = 1$.
On the other hand, if the researcher does not want to commit to a particular value of $(\rho, k)$, the researcher can perform simultaneous inference and report the largest value of $\rho$ such that the null can be rejected for each $k$. This approach can also be interpreted as finding the largest $\rho$ such that the conclusion changes, which is related to the idea of finding breakdown points/frontiers in econometrics horowitzmanski1995ecta, klinesantos2013qe, mastenpoirier2020qe. We demonstrate this through two applications in Section (ref).
The goal is to test the following hypothesis:
Our procedure of conducting inference with a single treated cluster is as follows.
To compute the $p$-value, the researcher does not need to run the entire Algorithm (ref). In particular, the researcher only needs to compute the test statistic $\widehat T_m$ in Step 1 of Algorithm (ref) and substitute it into the $p_m$ function in Step 2 of Algorithm (ref) and return the $p$-value as $p_m(\widehat T_m; k, \rho)$.
The researcher can take the critical value $\mathrm{cv}_{m, \alpha, k, \rho}$ from Step 2 of Algorithm (ref) and return $(\widehat\theta_{m+1} - \overline{\widehat{\theta}}_m) \pm \mathrm{cv}_{m, \alpha, k, \rho} \widehat S_m$ as an $(1 - \alpha)$ confidence interval for $\mu_1 - \mu_0$.
Algorithm (ref) is a computationally simple three-step procedure. The first step is to compute the usual $t$-statistic using the cluster-level estimators.
The second step finds the critical value that depends on $(m, \alpha, k, \rho)$. This step aims to find the critical value $\mathrm{cv}_{m, \alpha, k, \rho}$ such that the procedure controls size for any configurations of $\{\sigma_j\}^{m+1}_{j=1}$ that satisfy Assumption (ref). In the first case, a closed-form critical value is available when $\alpha$ is below a certain threshold in Table (ref). If $\alpha$ does not satisfy the condition, we can compute the critical value $\mathrm{cv}_{m, \alpha, k, \rho}$ in the second step through one-dimensional optimization.
The reason for having a closed-form expression for the critical value in case 1 of Step 2 of Algorithm (ref) is as follows. For $k=1$ and any given $(m, \rho)$, for a sufficiently small significance level $\alpha$ (including many conventional choices), we find that the maximum rejection probability over all configurations of $\{\sigma_j\}^{m+1}_{j=1}$ is achieved when $\sigma_{m+1}=\rho \sigma_j$ for all $j = 1, \ldots, m$. The proof is nontrivial and we explain the technical details in Sections (ref). Knowing when the maximum rejection probability is achieved, we are able to derive a closed-form expression for the critical value using properties of the $t$-distribution.
Case 2 of Step 2 is in fact a general version that holds for any $(m, \alpha, k, \rho)$. For a general value of $k \in \{1, \ldots, m\}$, although we cannot find the exact configuration of $\{\sigma_j\}^{m+1}_{j=1}$ that achieves the desired level of rejection probability, we are able to substantially reduce the set of possible values. In particular, there are at most $m^2$ possible cases for the worst-case configuration of $\{\sigma_j\}^{m+1}_{j=1}$, each of which involves at most one unknown parameter. Therefore, we can find the maximum rejection probability through one-dimensional optimization. In addition, we have a good initial guess for the critical value. To facilitate the use of the test in practice, we have tabulated many critical values for researchers' use in Table (ref).
The last step rejects if the absolute value of the test statistic is above the critical value.
In this section, we develop the theory that justifies the validity of Algorithm (ref). In Section (ref), we show that the large-sample behavior of the test can be studied via a fixed number of normal random variables. In Section (ref), we show the general theory on computing the maximum rejection probability of the test statistic. This result is used to search for a critical value such that the test is valid. We show in Section (ref) that we can get a closed-form expression for the critical value as in Step 2 of Algorithm (ref) when $k = 1$ and the significance level is not “too large.” In Section (ref), we discuss the power of our $t$-test.
In this subsection, we show that under Assumption (ref), studying the behavior of the $t$-statistic in equation (ref) of Algorithm (ref) when $n$ is large is equivalent to studying the $t$-statistic of suitably defined $(m+1)$ normal random variables. Formally, consider the following assumption on the normal random variables.
Compared to the notation in Section (ref), $\{\psi_j\}^m_{j=1}$ above correspond to the $m$ control clusters and $\psi_{m+1}$ corresponds to the single treated cluster. As before, the above assumption does not require us to know the variances $\{\sigma_j\}^{m+1}_{j=1}$. The variances can be arbitrarily heterogeneous as long as they satisfy the relative heterogeneity assumption stated in Assumption (ref). Analogous to (ref), we define the following $t$-statistic based on these normal random variables:
where $\overline\psi_m \equiv \frac{1}{m} \sum^m_{j=1} \psi_j$ and $S_m^2 \equiv \frac{1}{m-1} \sum^m_{j=1} (\psi_j - \overline\psi_m)^2$ denote the sample mean and variance for the control clusters, respectively.
The theorem below shows that, to achieve the desired size asymptotically, it suffices to consider a stylized setting with normally distributed random variables. Note that $\widehat{T}_m$ in (ref) depends on the sample size within each cluster $n$ implicitly. In addition, we allow $\mu_1$ and $\mu_0$ to vary with the sample size $n$ as well, which can facilitate power investigation under local alternatives.
Note that in the asymptotics of Theorem (ref), the limit is taken as $n$ goes to infinity, while the number of clusters $(m+1)$ is fixed. In addition, we allow a general $\delta$ for the difference between $\mu_1$ and $\mu_0$ to facilitate the later discussion on the power of the test. To derive valid tests, it suffices to consider the case where $\mu_1 = \mu_0$ and $\delta = 0$. Specifically, to derive large-sample valid $t$-test with a single treated cluster and a finite number of control clusters, it suffices to test the following null hypothesis in the stylized setting with exactly normal observations:
In the remainder of this section, we will study valid $t$-test for (ref) in this stylized setting.
The key to constructing a valid $t$-test is to find an appropriate critical value $\mathrm{cv}_{m, \alpha, k, \rho}$ for Step 2 of Algorithm (ref). This critical value $\mathrm{cv}_{m, \alpha, k, \rho}$ has to be chosen such that for any $\{\sigma_j\}^{m+1}_{j=1}$ that satisfy Assumption (ref), $\bP[|T_m|>\mathrm{cv}_{m, \alpha, k, \rho}]$ under the null hypothesis (ref) is less than or equal to a given significance level $\alpha$. In the following, we will consider the maximum rejection probability at any critical value $c$.
Note that when the treated standard deviation is much larger than the control standard deviations, i.e., $\sigma_{m+1} \gg \max_{1\le j \le m}\sigma_j$, the rejection probability $\bP[|T_m|>c]$ will be close to $1$ for any $c>0$, under which we cannot derive a meaningful $t$-test. Therefore, Assumption (ref) on the relative heterogeneity of $\{\sigma_j\}^{m+1}_{j=1}$ is in some sense necessary.
For descriptive convenience, we introduce $\mathcal{S}_m(k, \rho)$ to denote all possible standard deviations that satisfy Assumption (ref) for a given $m$, $\rho \geq 0$ and $k \in \{1, \ldots, m\}$\footnote{With a slight abuse of notation, we also use $\{\sigma_j\}_{j=1}^{m+1}$ to denote generic values of the standard deviations. Note that $\{\sigma_j\}_{j=1}^{m+1}$ in Assumptions (ref) and (ref) represent the true (asymptotic) standard deviations for all the clusters.}:
Using the above notation, our goal of finding the maximum rejection probability as described above can be formalized as follows
where the $\bP_0$ notation with subscript $0$ indicates that the null hypothesis (ref) holds and $\delta=0$. To conduct a valid test at any given significance level $\alpha\in(0,1)$, it suffices to find the critical value $c$ such that $p_m(c; k, \rho) \le \alpha$ and ideally $p_m(c; k, \rho) = \alpha$. On the other hand, if the goal is to compute a $p$-value, then one can directly compute $p_m(|T_m|; k, \rho)$ without searching for the required critical value as discussed in Section (ref). That is, when we set $c$ to be the observed absolute value of the $t$-statistic, $p_m(c; k, \rho)$ gives a valid $p$-value for testing the null hypothesis in (ref).
In the following, we consider two cases depending on the standard deviation for the treated cluster $\sigma_{m+1}$. Section (ref) considers $\sigma_{m+1} = 0$ and shows that the corresponding maximum rejection probability has a closed-form solution that can be easily computed. Section (ref) considers $\sigma_{m+1} > 0$ and shows that we have an integral representation for the rejection probability. This facilitates its numerical calculation at any given values of $\{\sigma_j\}_{j=1}^{m+1}$. However, directly evaluating the optimization problem in (ref) generally results in solving an $m$-dimensional optimization problem. We discuss how this can be reduced to solving multiple one-dimensional optimization problems in Section (ref).
When $\sigma_{m+1} = 0$, regardless of the choices of the relative heterogeneity parameters $\rho$ and $k$, the control standard deviations $\{\sigma_j\}^m_{j=1}$ can take arbitrary values in $\mathbb{R}^m_{\ge 0}$. Moreover, in this case, except for a multiplicative constant scaling factor of $\sqrt{m}$ in the $t$-statistic, our $t$-statistic essentially reduces to a one-sample $t$-statistic, and our $t$-test is equivalent to testing whether the mean of the control clusters is equal to zero. From Bakirov:2006aa, we know that the maximum rejection probability with a zero $\sigma_{m+1}$ has the following form:
where $R_m(c) \equiv \frac{m^2c^2}{mc^2+m-1}$, and $t_{j-1}$ denotes a random variable following the $t$-distribution with degrees of freedom $(j-1)$. In (ref), the maximum rejection probability is obtained when the control standard deviations $\{\sigma_j\}_{j=1}^m$ are either zero or take some common positive value. It can be efficiently computed by calculating the tail probabilities of at most $(m-1)$ $t$-distributions with various degrees of freedom.
It follows that when $\rho = 0$, we must have $\sigma_{m+1}=0$, and the maximum rejection probability of our $t$-test becomes the same as (ref), i.e., $p_m(c; k, 0) = p_{m,0}(c)$ for any $c>0$ and $k \in \{1, \ldots, m\}$. Thus, in the remainder of this section, we focus on $\rho > 0$.
We now consider the case where $\sigma_{m+1} > 0$. As shown in the following lemma, the rejection probability $\bP_0[|T_m|>c]$ can be written as an integral. This not only facilitates numerical computation, but is also crucial for our later theoretical investigation.
From Lemma (ref), when $\sigma_{m+1}>0$, the rejection probability in (ref) depends only on the ratios between the control standard deviations and the treated standard deviation. The relative heterogeneity assumption in Assumption (ref) equivalently assumes that at least $(m-k+1)$ of $\{\gamma_j\}^m_{j=1}$ is greater than or equal to $\rho^{-1}$. To find the maximum rejection probability, we need to solve an $m$-dimensional optimization over $(\gamma_1, \ldots, \gamma_m) \in [0, \infty)^{k-1} \times [\rho^{-1}, \infty)^{m-k+1}$.\footnote{Note that the rejection probability is invariant to any permutations of the $\{\gamma_j\}^m_{j=1}$. We can, for example, assume that the last $(m-k+1)$ of them is no less than $\rho^{-1}$ without loss of generality.} Optimizing this integral directly can be computationally challenging even for a moderate $m$. As demonstrated in the next subsection, for any $m\ge 2$, we can substantially simplify the $m$-dimensional optimization problem into multiple one-dimensional optimization problems. The reformulated problem can often be efficiently solved numerically.
In this subsection, we discuss how to reduce the optimization for the maximum rejection probability in (ref) into multiple one-dimensional optimization problems. We summarize the key ideas here and defer the technical lemmas to Appendix (ref) of this main text.
First, Lemma (ref) shows that the maximum rejection probability in (ref) must be achieved at some finite values of $(\sigma_1, \ldots, \sigma_m, \sigma_{m+1}) \in \mathcal{S}_m(k, \rho)$. This means we do not need to worry about the boundary case where some of $\{\sigma_j\}_{j=1}^{m+1}$ approach infinity. Next, Lemma (ref) shows the necessary conditions for any $(\sigma_1, \ldots, \sigma_m, \sigma_{m+1}) \in \mathcal{S}_m(k, \rho)$ to be a maximizer. We summarize the implications of Lemma (ref) in the following theorem.
Note that $0$ is always a boundary value of $\sigma_j$ for $1\le j \le m$, and $1$ is also a boundary value for some $\sigma_j$ when $\sigma_{m+1} = \rho$ and Assumption (ref) holds. Therefore, Theorem (ref) essentially indicates that, at the maximizer of the rejection probability in (ref), each $\sigma_j$ must either lie on the boundary or share a common value for all $1 \le j \le m$. Importantly, Theorem (ref) explains why we can simplify the optimization for the rejection probability into multiple one-dimensional optimization problems. This is because for all possible maximizers shown in Theorem (ref), there is only one unknown value, which is the common value of the $\{\sigma_j\}^m_{j=1}$ that are not on the boundary. Note that the maximum rejection probability in case (i) with $\sigma_{m+1} = 0$ can be efficiently computed as shown in (ref). In the following, we will therefore focus on the optimization for the rejection probability in case (ii) with $\sigma_{m+1} > 0$.
Specifically, under Assumption (ref) with any $k \in \{1, \ldots, m\}$ and $\rho > 0$, if $(\sigma_1, \ldots, \sigma_m,$ $\sigma_{m+1}) \in \mathcal{S}_m(k, \rho)$ is the maximizer for the rejection probability and $\sigma_{m+1}>0$, then, for all $1\le j \le m$, $\gamma_j \equiv \frac{\sigma_j}{\sigma_{m+1}}$ is either on the boundary (equals $0$ or $\rho^{-1}$ ), or takes some common value $\gamma \ge 0$. Motivated by this, with a slight abuse of notation, we introduce $\overline{p}_m(c; \rho, \gamma; m_1, m_0)$ to denote the value of $\overline{p}_m(c; \gamma_1, \ldots, \gamma_m)$ in (ref) when $m_1$ of $\{\gamma_j\}_{j=1}^m$ equal $\rho^{-1}$, $m_0$ of them equal zero, and the remaining equal a common value $\gamma$. That is,
where $0 \leq m_0, m_1 \leq m$ and $m_0 + m_1 \leq m$.
Next, define the following as the supremum of $\overline p_m(c; \rho, \gamma; m_1, m_0)$ over $\gamma$, with the support of $\gamma$ depending on $m_1$ and $k$:
where $\underline\rho = 0$ if $m_1 \ge m-k+1$ and $\underline\rho = \rho^{-1}$ if $m_1 < m-k+1$.
With the above notations, we now state our main theorem of how to evaluate the maximum rejection probability for a given critical value $c$, heterogeneity parameters $(k, \rho)$ and the number of clusters $m$.
From Theorem (ref), the key to obtaining the maximum rejection probability is to solve the optimization problem in (ref) for all combinations of $(m_1, m_0)$. Importantly, for each given $(m_1, m_0)$, the optimization in (ref) is an one-dimensional optimization problem, which is computationally much simpler than the original optimization in (ref), and we can solve it using numerical optimization. In particular, for any given $k \in \{1, \ldots, k\}$, we need to solve at most $\frac{1}{2}k(2m+1-k)$ one-dimensional optimization problems\footnote{When $m_1+m_0 = m$, (ref) no longer depends on $\gamma$, and thus no optimization over $\gamma$ is needed.}, which is $m$ when $k=1$, $2m-1$ when $k=2$, \ldots, and $\frac{1}{2}m(m+1)$ when $k=m$. Moreover, at a given $m$, $\rho \geq 0$ and $c > 0$, to find the maximum rejection probability over all $k \in \{1, \ldots, m\}$, we need to solve at most $m(m+1)$ one-dimensional optimization problems of the form (ref), because the optimizations required for different values of $k$ overlap with each other.
We illustrate the above theorem through the following examples. Example (ref) is about $k = 1$ and Example (ref) is about $k = 2$.
Finally, we report the critical values for different numbers of control clusters $m$ and heterogeneity parameters $\rho$ for $\alpha = 0.01$ and $\alpha = 0.05$ in Table (ref) when $k = 1$. We report the critical values for $k = 2$ in the supplementary material.
In this subsection, we consider the case where $k = 1$ in Assumption (ref). This means that $\sigma_{m+1} \leq \rho \sigma_j$ for $j = 1, \ldots, m$, i.e., the treated standard deviation is smaller than or equal to $\rho$ times each of the control standard deviations.
In this case, we can obtain closed-form solutions for the maximum rejection probability when the threshold $c$ is large enough. Equivalently, this corresponds to testing at a significance level that is not “too large.” The theorem is stated as follows:
Theorem (ref) gives a closed-form solution for the maximum rejection probability under the relative heterogeneity assumption with $k=1$ and any given $\rho > 0$. The maximum rejection probability is achieved when $\sigma_1 = \sigma_2 = \ldots = \sigma_m = \rho^{-1}$ and $\sigma_{m+1} = 1$. Importantly, for any given $m$ and $\rho>0$, the cutoff $\underline{c}_{m, \rho}$ can be easily computed numerically, due to the monotonicity of the function $\overline H_m(c, \rho)$ in $c$. Accordingly, Theorem (ref) gives a closed-form critical value of our $t$-test for any significance level less than or equal to
That is, for any significance level $\alpha \in (0, \underline{\alpha}_{m, \rho}]$, a valid critical value for our two-sided $t$-test is $\sqrt{\rho^2 + m^{-1}} \ t_{m-1,1- \frac{\alpha}{2}}$, where $t_{ m-1,1- \frac{\alpha}{2}}$ is the $(1-\frac{\alpha}{2})$ quantile of the $t$-distribution with degree of freedom $m-1$.
Table (ref) reports the largest significance level $\underline{\alpha}_{m, \rho}$ such that our two-sided $t$-test has a simple closed-form expression for the critical value, under Assumption (ref) with $k=1$ and various values of $(m, \rho)$. Note that our $t$-test can handle all values of $m\ge 2$, $k \in \{1, \ldots, m\}$ and $\rho \ge 0$, but it generally involves one-dimensional optimization as described in Section (ref). We can conduct a similar theoretical investigation to simplify the optimization of the rejection probability under Assumption (ref) with $k \geq 2$. However, in this case, the potential maximizers cannot be reduced to a single point, so one-dimensional numerical optimization is still required. We relegate the detailed discussion under Assumption (ref) with $k\geq 2$ to the supplementary material.
We now investigate the power of the proposed $t$-test under the alternative hypothesis $\overline{H}_1$ in (ref) with $\delta \ne 0$. Without loss of generality, we assume that $\delta>0$. The theorem below gives a lower bound for the power of the $t$-test.
The lower bound of power in Theorem (ref) increases with $\delta$, where $\delta$ corresponds to $\sqrt{n}(\mu_1 - \mu_0)$ in our large-sample inference for a finite number of clusters as shown in Theorem (ref). If the gap between the means of the treated and control clusters is bounded away from zero, then, as the sample size $n\longrightarrow \infty$, the power of our $t$-test with any finite critical value will converge to $1$. In addition, Theorem (ref) also provides a rough power estimate when we have some information about the treatment effect size and the variances for the treated and control clusters.
Recall that Assumption (ref) involves two parameters $\rho$ and $k$. They represent the allowable degree of relative heterogeneity between treated and control clusters. Specifically, larger values of $\rho$ and $k$ correspond to a greater degree of allowable heterogeneity. In practice, specifying the values of $(\rho, k)$ might be challenging, and it is often desirable to perform multiple analyses for a wide range of values for $(\rho, k)$. This raises the question of how to interpret the results of the analysis under different values of $(\rho, k)$. In this section, we show that the analyses over all possible values of $(\rho, k)$ can be simultaneously valid, without the need of any adjustment due to multiple analyses. Moreover, the analysis results can be easily visualized and interpreted. Similar simultaneous inference has been used for sensitivity analysis of observational studies cuili2025wp,wuli2025jasa.
We first introduce several notations to denote the true relative heterogeneity of the standard deviations between treated and control clusters. For any $k \in \{1, \ldots, m\}$, define
where $\sigma_{m+1}$ and $\sigma_{(k)}$ in (ref) denote the true standard deviation of the treated cluster and that of the control cluster at rank $k$. Consequently, the values of $\rho_k^\star$ reflect the true relative heterogeneity between the treated and control clusters. In particular, larger values of $\rho_k^\star$ indicate greater heterogeneity between the treated and control clusters.
We now explain the key idea for our simultaneous inference procedure. In Section (ref), we test the null hypothesis in (ref) of no mean difference between the treated and control clusters under Assumption (ref) for some $k \in \{1, \ldots, m\}$ and $\rho \ge 0$, which is equivalent to that $\rho_k^\star \le \rho$ with $\rho_k^\star$ defined as in (ref). Importantly, we can reinterpret the $t$-test in Algorithm (ref) as a valid test for the null hypothesis of $\rho_k^\star \le \rho$ about the true relative heterogeneity in (ref), under the assumption of no mean difference between treated and control clusters (i.e., $\delta = 0$). By standard test inversion, we can construct confidence intervals for $\rho_k^\star$, which must be one-sided confidence intervals with unbounded right endpoints and thus provide essentially lower bounds on the true relative heterogeneity, under the assumption that $\delta=0$. Moreover, these confidence intervals will be simultaneously valid across all $k \in \{1, \ldots, m\}$. We summarize the results in the following theorem, followed by a discussion of its practical implications and interpretation. For any $z\in \mathbb{R}$, we write $(z, \infty] \equiv (z, \infty)\cup \{\infty\}$ and analogously $[z, \infty] \equiv [z, \infty)\cup \{\infty\}$.
The interpretation of Theorem (ref) is as follows. If there is no mean difference between treated and control clusters (i.e., $\delta=0$), then we need to believe that, with $(1-\alpha)$ confidence level, the relative heterogeneity $\rho_k^\star$ in (ref) must be greater than (or equal to) $\hat{\rho}_{m, \alpha, k}$, for all $k \in \{1, \ldots, m\}$. That is, with $(1-\alpha)$ confidence level, the treated standard deviation $\sigma_{m+1}$ must be at least $\hat{\rho}_{m, \alpha, k}$ times the control standard deviation $\sigma_{(k)}$ at rank $k$, for all $k \in \{1, \ldots, m\}$. If we question any of these statements on the true relative heterogeneity, then the assumption of $\delta=0$ is likely to fail, or equivalently there is likely significant mean difference between the treated and control clusters. Obviously, the larger the values of $\{\hat{\rho}_{m, \alpha, k}\}^m_{k=1}$, the stronger the evidence for a nonzero treatment effect. In practice, we can easily visualize these confidence intervals by, say, plotting $k$ against $\hat{\rho}_{m, \alpha, k}$; this is illustrated in the two empirical applications in Section (ref) ahead.
We now discuss the computation for the confidence intervals in Theorem (ref). By the definition of the $p$-value $p_m(c; k, \rho)$ in (ref) and the fact that the set $\mathcal{S}_m(k, \rho)$ in (ref) increases as $k$ or $\rho$ increase, $p_m(c; k, \rho)$ must be nondecreasing in both $k$ and $\rho$. The monotonicity in $\rho$ then explains the equivalent one-sided form of the set in (ref). Thus, we can use the bisection method to find the thresholds $\{\hat{\rho}_{m, \alpha, k}\}^m_{k=1}$. In addition, the monotonicity in $k$ implies the threshold $\hat{\rho}_{m, \alpha, k}$ is nonincreasing with $k$. Hence, we can first find $\hat{c}_{m, \alpha, 1}$, and then sequentially use $\hat{c}_{m, \alpha, k-1}$ as an upper bound for $\hat{\rho}_{m, \alpha, k}$ in the bisection method, for $2\le k \le m$.
In this section, we consider two sets of numerical exercises to compare the performance of the $t$-test against other methods for conducting inference with a single treated cluster.
The first simulation design generates data using normal distributions. The data generating process (DGP) is as follows. We generate normal random variables as in Assumption (ref). In particular, $\{\psi_j\}^m_{j=1}$ are normally distributed with mean 0, $\psi_{m+1}$ has mean $\delta$, $\sigma_{m+1}^2 = \rho^2$, and $\{\sigma_j^2\}^m_{j=1}$ are specified below.
We consider two different DGPs to generate the random variables for the control clusters.
DGP 1 is the homogeneous design in which all control random variables have the same varaiance. We introduce heterogeneity in DGP 2 and allow the variance to vary between 1 and 2. The goal is to test the two-sided hypothesis (ref), under various heterogeneity parameter $\rho$, cluster size $m$, and significance level $\alpha$. We compare the performance of our $t$-test with hagemann2024wp in terms of size and power. This is because both of our tests work with a single treated cluster and a finite number of control clusters, and are valid under a certain relative heterogeneity assumption.
In the simulations, we consider $\rho$ from 0.1 to 2 with step size 0.1, $m \in \{5, 10, 25, 50\}$, and $\alpha \in \{0.01, 0.05, 0.1\}$. We focus on two-sided tests and choose $\rho$ in our $t$-test and hagemann2024wp so that it matches the one used in the DGP, i.e., the relative heterogeneity assumption is always correctly specified. The results are based on 500,000 Monte Carlo replications. Figure (ref) shows the result for $\alpha = 0.05$ under DGP 1 when $k = 1$. The figure plots the rejection rate against various values of $\rho$. Each facet represents a specific combination of $m$ and $\delta$. $\delta = 0$ corresponds to the results under the null. $\delta > 0$ corresponds to the results under the alternative. Note that hagemann2024wp's result does not appear or only partially appear in some of the facets. This is because hagemann2024wp may not necessarily be able to find a weight for his rearrangement test such that his test can be shown to be valid for some combinations of $m$, $\rho$ and $k$.
Figure (ref) shows that our test performs favorably when compared to hagemann2024wp. Both of our tests control size. The $t$-test is more powerful, especially when $\rho$ is small. As $\rho$ increases above 1, the power difference decreases, but we are still more powerful.
Next, Figure (ref) reports the results for DGP 2 that includes more heterogeneity among the control random variables for $\alpha = 0.05$ and $k=1$. As predicted by the theory, our $t$-test becomes more conservative when more heterogeneity are included. Our $t$-test continues to control size and has power against the alternative. Moreover, our test outperforms hagemann2024wp's test in most cases. In the supplementary material, we report the results for $k=1$ and $k=2$ for both DGP at various significance levels.
Next, we conduct a simulation exercise with two-way fixed effects as in Example (ref). We consider a design that is based on the ones used in conleytaber2011restat and hagemann2024wp. As before, let $\cJ_0 \equiv \{1, \ldots, m\}$ be the control clusters and $\cJ_1 \equiv \{m+1\}$ be the treated cluster. Let $\cT \equiv \{1, \ldots, 10\}$ be the total number of time periods. Let $t_0 = 6$ be the intervention period and $D_{jt} \equiv \ind[j \in \cJ_1 \text{ and } t > t_0]$ for $j \in \cJ$ and $t \in \cT$. Each simulated data is generated from the following two-way fixed effects model
where $\alpha_t = 1$ for all $t \in \cT$ and $\gamma_j = 2 \ind[j \leq \frac{m}{2}] - 1$ for all $j \in \cJ$. We test the null of $\theta = 0$ and consider $\theta \in \{0, 1, 2, 3\}$ when generating the data. We consider the following DGPs based on model (ref). The first four DGPs are the same as the ones considered in hagemann2024wp. The remaining two DGPs are based on the error distributions considered in conleytaber2011restat.
For each DGP, we are interested in testing the two-sided hypothesis of \[ H_0: \theta = 0 \quad \text{ vs } \quad H_1 : \theta \neq 0. \] In addition, we consider simulations with $m \in \{5, 10, 25, 50\}$, $ \sigma \in \{0.1, 0.5, 1, 2\}$, and $\alpha = 0.05$. We report the rejection rate curves based on 5,000 Monte Carlo simulations. For each DGP, we consider conducting inference using the following four methods:
Figures (ref) and (ref) show the results for DGPs 1 and 2 under $\alpha = 0.05$ across different numbers of clusters $m$ and true relative heterogeneity $\sigma$. We set the relative heterogeneity parameter $\rho$ for both our and hagemann2024wp's method in all the DGPs to be the corresponding $\sigma$, so that the relative heterogeneity assumption is correctly specified. We also incorporate the correct value of $\sigma$ in applying FP to specify the heterogeneity.
It can be seen that our $t$-test performs well and compares favorably to other methods. In particular, the CT test exhibits inflated Type I error rates when the number of control clusters $m$ is small and $\sigma$ is large. This inflation arises because the validity of CT relies on the assumption of an infinite number of clusters and homogeneous variances between the treated and control clusters. The FP test also shows an inflated Type I error when $m$ is small, with the inflation becoming more pronounced as $\sigma$ decreases. This again reflects its reliance on the asymptotic validity under an infinite number of clusters. Notably, we assume that FP has access to the true heterogeneity between the treated cluster and each control cluster, which is a stronger assumption than those required by our $t$-test or hagemann2024wp's rearrangement test. Both our test and that of hagemann2024wp successfully control Type I error. In contrast, our test is applicable regardless of the number of control clusters and demonstrates higher power across different levels of heterogeneity. The results for DGPs 3 to 5 are contained in the supplementary material.
In this section, we illustrate the $t$-test proposed in this paper with two recent empirical applications. For each of the empirical applications, we report two sets of results. First, we try to find the largest $\rho$ such that the null hypothesis is rejected when $k = 1$ among various regression specifications of the empirical examples. Among those specifications where the null can be rejected at some $\rho$, we conduct simultaneous inference as discussed in Section (ref).
depewswensen2022ej examines the impact of the 1911 New York State Sullivan Act on mortality rate. The Sullivan Act required New York citizens to obtain a permit and license to carry concealable weapon. We consider the following two-way fixed effects model:
where $\text{Treated}_{jt}$ equals 1 if state $j$ is New York and $t$ is the post-treatment period, $\alpha_t$ is the year fixed effects, $\gamma_j$ is the state fixed effects, and $\epsilon_{jt}$ is the residual term. They consider the following four outcome variables: homicide rate, suicide rate, gun suicide rate, and non-gun suicide rate.
They report cluster-robust standard errors and $p$-value from the Wild cluster bootstrap. There is only nine control clusters in this empirical application. Hence, hagemann2024wp's test may not be applicable because there can be no weights such that his rearrangement test is valid for some relative heterogeneity parameters $\rho$.
We are interested in testing
Table (ref) summarizes the estimation and inference results of the baseline model described in (ref) for the four outcomes described above. The point estimates and the Wild cluster bootstrap $p$-values are taken from Table 2 of depewswensen2022ej.\footnote{The point estimates and $p$-values that depewswensen2022ej report are based on weighted least squares.}
The remaining rows of Table (ref) report the largest $\rho$ such that the null can be rejected for significance levels $\alpha = 0.01$, $0.05$, and $0.1$. For entries with “NA” in the table, it refers to the situation where there does not exist a $\rho \geq 0$ such that the null is rejected. If we focus on the gun suicide rate under $\alpha = 0.1$, it means that the null can be rejected when the variance in New York is at most $1.3^2 \approx 1.69$ larger than the smallest variance in the control clusters. The true relative level of heterogeneity such that the null can be rejected may be related to the population size of the various states. In particular, New York state has a larger population than the other control states in 1911 fred2025website: the population of New York state was 9.249 million, and the largest control state was Massachusetts with 3.383 million and the smallest control state was Vermont with 0.358 million.
Next, we conduct simultaneous inference for our $t$-test as in Section (ref). We focus on the outcome “gun suicide rate” as we can find $\rho$ such that the null is rejected when $k=1$. Figure (ref) reports the results for $\alpha \in \{0.01, 0.05, 0.1\}$ on the range of $\rho$ for various $k$ such that the null can be rejected. The shaded area represents the simultaneous confidence region for true relative heterogeneity when there is no treatment effect.
hiraiwaetal2024wp examine whether firms enforce noncompete agreements (NCAs). NCAs are restrictions that prevent workers from joining or starting competing firms. In 2020, Washington state passed a law that prohibits NCAs for workers earning below a certain threshold. The threshold was \$100,000 per year in 2020 and \$101,390 per year in 2021. They study the impact of being above or below the threshold on employment for different earnings bins. They consider the following two-way fixed effects model in their panel B of Table 1: \[ \log \text{Emp}_{b,t} = \beta \text{Treated}_{b,t} + \alpha_t + \gamma_b + \epsilon_{b,t}, \] where $\text{Emp}_{b,t}$ is the employment count of bin $b$ at year $t$, $\text{Treated}_{b,t}$ equals 1 for the focal bin in the focal year, $\alpha_t$ and $\gamma_b$ are the year and bin fixed effects, and $\epsilon_{b,t}$ is the error term. They have 30 clusters (income bins). They use one-sided randomization inference and found that there is no significant effect.
We apply our method and hagemann2024wp. We are interested in testing the two-sided hypothesis as in (ref). Table (ref) summarizes our results. Each column contains the result for a specific focal year and definition of the treatment variable. For instance, column 2 refers to focal year 2020, and the treatment variable is defined to be equal to 1 when the income bin is just above the threshold, i.e., \$100-101.389k. The remaining columns are defined similarly. For each regression specification, we find the largest $\rho$ such that the null is rejected. For entries with “NA” in the table, it refers to the situation where there does not exist a $\rho \geq 0$ such that the null is rejected.
Next, we conduct simultaneous inference for our $t$-test as in the first empirical application. We focus on the two variables for focal year 2021 because we can find $\rho$ such that the null is rejected when $k=1$. In particular, we report the range of $\rho$ for various $k$ as long as there exists $\rho$ such that the null can be rejected. Figure (ref) reports the results for $\alpha \in \{0.01, 0.05, 0.1\}$.
In this paper, we propose a $t$-test to conduct inference with a single treated cluster under a certain relative heterogeneity assumption. The $t$-statistic and the associated critical values are easy to calculate in many empirically relevant applications. We show that our test performs favorably when compared to other methods in the literature. We also show that our test is valid with weaker assumptions.