Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
36,135 characters · 6 sections · 59 citation commands
A Combinatorial Central Limit Theorem for Stratified Randomization
Stratified randomization is frequently utilized to enhance covariate balance in randomized experiments. In this scheme, blocks or strata are initially created based on pre-treatment variables, and randomization inference is subsequently conducted by randomizing units within each stratum Imbens-Rubin(2015). Stratified randomization can be used for obtaining exact tests Imbens-Rosenbaum(2005), Andrews-Marmer(2008), Zhao-Ding2021, DHT2023 and efficient estimators Cytrynbaum2021, Armstrong2022, BLSTM2023, and is a natural inferential choice for models with survey data obtained through stratified sampling.
The key tool used when establishing the asymptotic validity of the permutation tests is the combinatorial central limit theorem (CLT) Hoeffding1951. While the standard combinatorial CLT or finite population CLT can be applied to derive the asymptotic distributions of stratified permutation statistics when there is a finite number of large strata Imbens-Rosenbaum(2005), LiDing2017, its application to general stratified permutations is not straightforward. This complication arises because the number of strata, either large or small, or both, may grow with the sample size (for instance, consider the case of having $n^{1/2}$ many strata, each of size $n^{1/2}$, where $n$ denotes the sample size).
Despite recent contributions by LiuYang2020 and LiuRenYang2022 that propose finite population CLTs tailored for such scenarios, a combinatorial CLT accommodating potentially many strata of unrestricted size remains essential for addressing general stratified permutation/rank statistics. In this context, this paper introduces a general form of the combinatorial CLT for stratified randomization, applicable under a Lindeberg-type condition. The proof is based on Stein's method CGS2011, and in particular, draws from the constructions of Bolthausen1984 and Schneller1988. Leveraging this result, we characterize the asymptotic distribution of a test statistic proposed by Imbens-Rosenbaum(2005) and a variant of a test statistic by Andrews-Marmer(2008) in the instrumental variables (IVs) setting featuring many strata, under weak assumptions.
A substantial body of literature exists on stratified randomized experiments and large sample results for stratified randomization. For empirical practices and an overview, see, for example, BruhnMcKenzie2009, Imbens-Rubin(2015) and Ding2023. Papers that consider a coarse stratification involving a finite number of strata, where asymptotic normality is typically derived via finite population or rank CLTs, include Andrews-Marmer(2008), LiDing2017, Zhao-Ding2021. Papers that consider a finely stratified experiment with many strata, where the asymptotic normality is derived via Lindeberg or Lyapunov CLTs, include Fogarty2018mit, Bai2022 and BaiRomanoShaikh2022.
Furthermore, while LiuYang2020 and LiuRenYang2022 provide finite population CLTs for stratified experiments involving many strata, DHT2023 establish a combinatorial CLT with many strata to examine the properties of a stratified permutation subvector test in linear regressions. The CLT for stratified randomization in this paper is sufficiently general to encompass the aforementioned CLTs as it handles both large and small strata simultaneously and requires a weaker Lindeberg-type condition. In particular, finite population CLTs with many strata follow from the result akin to the derivation of usual finite population CLTs from standard combinatorial or rank CLTs.
Stratified permutation is an instance of permutations with restricted positions Rosenbaum1984,Diaconis2001,LeiBickel2021; thus, the set of such permutations forms a subset of all permutations of observation indices. CGS2011 provide normal approximation results for a class of restricted permutations {\textemdash} in contrast with the set of entire permutations required in Hoeffding1951's CLT {\textemdash} but do not consider stratified permutations. The present paper fills this gap in the literature by developing normal approximation for this important class of permutations.
This paper proceeds as follows. Section (ref) lays out the basic framework and presents the combinatorial CLT for stratified randomization with an arbitrary number of strata. Section (ref) establishes a connection between our result and finite population CLTs. Section (ref) applies the combinatorial CLT to derive the asymptotic distribution of two test statistics in IV settings. The proofs are provided in the appendix.
Let $n$ be a fixed sample size, ${S}$ be an integer-valued random variable (${S}\ge 1$ a.s.), and $s=1,\dots,{S}$ denote strata of sizes $n_s\ge 2$ a.s., with $\sum_{s=1}^{{S}} n_s=n$. Consider a double array of real-valued random variables in stratum $s$: $\{a_{ij}^s\}_{i,j=1}^{n_s}$, and define $\bm{a}_{n}\equiv \{a_{ij}^s\}_{s=1,\dots, {S}; (i,j)\in\{1,\dots, n_s\}^2}$. Here, the distributions of $\{a_{ij}^s\}_{i,j=1}^{n_s}$ and ${S}$ may depend on $n$, and the stratification may be based on random auxiliary covariates. Let $\pi$ be a stratified permutation that permutes the indices within each stratum:
where $\pi_s$ is a permutation of $\{1,\dots, n_s\}$ in stratum $s=1,\dots,{S}$.\footnote{For simplicity, we do not further index the elements of $\{1,\dots, n_s\}$ by $s$ throughout the paper.} We denote by $\mathbb{S}_n$ the set of all such permutations with $|\mathbb{S}_n|=\prod_{s=1}^{{S}}n_s!$. In this paper, we consider permutations uniformly distributed over $\mathbb{S}_n$ for $\bm{a}_{n}$ given: $\pi\sim\mathcal{U}(\mathbb{S}_n)$, where $\mathcal{U}(A)$ denotes the uniform distribution on a finite set $A$. Let $P^\pi$ be the probability measure of $\pi$ conditional on $\bm{a}_{n}$, and $\operatorname{E}_\pi[\cdot]$ and $\operatorname{Var}_\pi[\cdot]$ denote the corresponding expectation and variance operators.
Furthermore, we use the notation $1(\cdot)$ for the indicator function, and $\Phi(\cdot)$ and $\Phi^{-1}(\cdot)$ for the cumulative distribution and quantile function of a standard real normal distribution. $\sum_{i,j=1}^{n_s}$ and $\sum_{i\neq j}$ are the shortcuts for $\sum_{i=1}^{n_s}\sum_{j=1}^{n_s}$ and $\sum_{i,j=1; i\neq j}^{n_s}$, respectively. We abbreviate “left-hand side” and “right-hand side” as LHS and RHS. Let $\Vert\cdot\Vert$ denote the Frobenius norm, $\lambda_{\min}(\cdot)$ denote the minimum eigenvalue of a square matrix, and $\max_{s, i}$ denote for the maximum over $s\in\{1,\dots,{S}\}$ and $i\in\{1,\dots,n_s\}$ for a given $n$.
The main result of this paper is the following combinatorial CLT for stratified randomization which is based on a Lindeberg-type condition.
A leading special case of interest similar to that considered in Remark (ref)(ref) is the asymptotic normality of the sum $n^{-1/2}\sum_{s=1}^{{S}}\sum_{i=1}^{n_s}\tilde{b}_{si}\tilde{c}_{s\pi(i)}$, where $\tilde{b}_{si}\equiv b_{si}-\bar{b}_s$, $\bar{b}_s\equiv n_{s}^{-1}\sum_{i=1}^{n_s}b_{si}$, $b_{si}\in\mathbb{R}^k$ with $k\geq 1$, $\tilde{c}_{si}\equiv c_{si}-\bar{c}_s$, $c_{si}\in\mathbb{R}$ and $\bar{c}_s\equiv n_{s}^{-1}\sum_{i=1}^{n_s}c_{si}$. The following result holds as a corollary to Theorem (ref).
This section explores the connection between the combinatorial CLT of Theorem (ref) and a finite population CLT. Bickel1984 and LiuYang2020 establish finite population CLTs for linear combinations of stratum means under stratified sampling, where the number of strata may diverge. Consider a finite population divided into non-random ${S}$ strata: $(y_{s1},\dots, y_{sn_s}), s=1,\dots, {S}$. The stratum-specific mean and variance are defined respectively as $\bar{y}_s\equiv n_s^{-1}\sum_{i=1}^{n_s}y_{si}$ and $v_{s}^2\equiv \frac{1}{n_s-1}\sum_{i=1}^{n_s}(y_{si}-\bar{y}_s)^2$, while the population mean and weighted variance are given by
where $p_s\equiv n_s/n$ and $n_{1s}, 2\leq n_{1s}\leq n_s-1,$ denotes the number of units sampled without replacement from stratum $s$ in Bickel1984's set-up (the number of treated units in stratum $s$ in LiuYang2020's set-up). The sampling indicators are $(Z_{s1},\dots, Z_{sn_s})\in\{0,1\}^{n_s}, s=1,\dots, S$, where $Z_{si}=1$ if the unit $i$ in stratum $s$ is sampled, and $0$ otherwise. The probability that $\{(Z_{s1},\dots,Z_{sn_s})\}_{s=1}^S$ takes on a value $\{(z_{s1},\dots,z_{sn_s})\}_{s=1}^S$ such that $n_{1s}=\sum_{i=1}^{n_s}z_{si}$ for $s=1,\dots, S,$ is $\prod_{s=1}^Sn_{1s}!(n_s-n_{1s})!/n_s!$. The weighted sample mean and variance are then defined as
where $\hat{y}_s\equiv n_{1s}^{-1}\sum_{i=1}^{n_s}Z_{si}y_{si}$ and $\hat{v}_{s}^2 \equiv \frac{1}{n_s-1}\sum_{i=1: Z_{si}=1}^{n_s}(y_{si}-\hat{y}_s)^2$ are the stratum-specific sample mean and variance. With the notations introduced, the Lindeberg condition employed by LiuYang2020s (Condition A1 and Theorem A1 therein) may be restated as follows: for any $\epsilon>0$
as $n\to\infty$, where
Under the condition in (ref), Bickel1984 show that $(\hat{y}-\bar{y})/{\hat{v}}\displaystyle \stackrel{d}{\longrightarrow} \mathcal{N}\left(0,1\right)$ as $n\to\infty$. To draw a link between the latter result and Theorem (ref), we let $a_{ij}^s=\tilde{b}_{si}\tilde{c}_{sj}$ as in Remark (ref)(ref) and Corollary (ref), where $b_{si}$ and $c_{sj}$ are now defined as
The choice of $b_{si}$ in (ref) appears in Madow1948 and Sen1995, among others. By simple algebra, $\bar{b}_s=n_s^{-1}\sum_{i=1}^{n_s}b_{si}=n_s^{-1}$, $\bar{c}_s= n_s^{-1}\sum_{j=1}^{n_s}c_{sj}=p_s\bar{y}_s$, and
It turns out that the variance $\sigma_n^2$ defined in the condition (ref) of Theorem (ref) and $v^2$ in (ref) are identical:
Then, through a direct calculation given in Section (ref), the Lindeberg condition in Theorem (ref)(ref) follows from the condition in (ref), as summarized in the remark below.
This section presents two applications of the CLT in Theorem (ref). First, we establish the asymptotic distribution of a statistic proposed by Imbens-Rosenbaum(2005). Second, we show that the test of Andrews-Marmer(2008) may be conservative in the presence of small strata. We then consider a simple modification of their statistic that yields a correct asymptotic level and derive its null asymptotic distribution.
Imbens-Rosenbaum(2005) consider an IV setting with non-random $S$ strata $s=1,\dots, S$ each of size $n_s$, given by: $$Y-\beta D=r_C,$$ where $D=[D_1',\dots, D_{S}']\in\mathbb{R}^n$ with $D_s=[D_{s1},\dots, D_{sn_s}]'\in\mathbb{R}^{n_s}$ and $n=\sum_{s=1}^Sn_s$, denotes the vector of doses received by the units, $Y=[Y_1',\dots, Y_S']'\in\mathbb{R}^n$ with $Y_s=[Y_{s1},\dots, Y_{sn_s}]'\in\mathbb{R}^{n_s}$ is the vector of responses exhibited by the units, $r_C=[r_{C1}',\dots, r_{CS}']'\in\mathbb{R}^n$ with $r_{Cs}=[r_{Cs1},\dots, r_{Csn_s}]'\in\mathbb{R}^{n_s}$ is the vector of responses if the units received the control dose, and the scalar parameter $\beta$, with a true value $\beta_0$, measures of the dose-response relationship. Let $h_s=[h_{s1},\dots, h_{sn_s}]'\in\mathbb{R}^{n_s}$ be a sorted and fixed IV vector such that $h_{sj}\leq h_{s(j+1)}$ for each $j=1,\dots, n_s-1$ and $s=1,\dots,S$. For $h\equiv [h_1',\dots, h_S']'\in\mathbb{R}^{n}$, assume that the observed IV $Z$ is randomly assigned i.e. $Z=h_\pi\in\mathbb{R}^n$, where $\pi\sim \mathcal{U}(\mathbb{S}_n)$ and $\mathbb{S}_n$ is the set of all stratified permutations in this setting. Here, $r_C=Y-\beta_0 D$ is assumed to be fixed, hence independent of $Z$. For testing the hypothesis $H_0:\beta=\beta_0$, Imbens-Rosenbaum(2005) propose the following statistic:
where $q(\cdot)\in\mathbb{R}^n$ is some method of scoring responses such as the ranks or the aligned ranks, and $\rho(\cdot)\in\mathbb{R}^n$ is some method of scoring the instruments that satisfies $\rho_\pi(h)=\rho(h_\pi)=\rho(Z)$. Let $\rho(h)=[\rho_{1}',\dots, \rho_S']'$ with $\rho_s=[\rho_{s1},\dots, \rho_{sn_s}]'$, and $q(Y-\beta_0D)=[q_{1}',\dots, q_{S}']'$ with $q_s=[q_{s1},\dots, q_{sn_s}]'$. Before we state the result for the asymptotic distribution of the statistic $T$, we remark that $T=\sum_{s=1}^S\sum_{i=1}^{n_s}q_{si}\rho_{s\pi(i)}$ using the notations above. With a convention that $\operatorname{Var}_\pi[\sum_{i=1}^{n_s}q_{si}\rho_{s\pi(i)}]=0$ for stratum of size $n_s=1$, it follows from Proposition 1 of Imbens-Rosenbaum(2005) that
where $\bar{q}_s\equiv n_s^{-1}\sum_{i=1}^{n_s}q_{si}$, $\bar{\rho}_s\equiv n_s^{-1}\sum_{i=1}^{n_s}\rho_{si}$, and $\tilde{q}_{si}\equiv q_{si}-\bar{q}_s$ and $\tilde{\rho}_{si}\equiv \rho_{si}-\bar{\rho}_s$.
Imbens-Rosenbaum(2005) outline strategies for deriving the asymptotic distribution of the statistic $T$ in the presence of either a few large strata whose sizes tend to infinity, or many small strata, whose sizes remain bounded, and suggest applying the rank-central limit theorem Hajek-Sidak-Sen(1999) to the former and standard CLTs to the latter.
LiDing2017 consider the case where $q(\cdot)$ is an identity function and $\rho(Z)$ is a vector of binary IVs, and derive the null asymptotic distribution of the statistic in (ref). The general CLT provided in Theorem (ref) allows us to derive the asymptotic distribution of the statistic $T$ in (ref) at once under general conditions, allowing for many strata without categorizing them into large and small ones or restricting their sizes. The following proposition gives a formal justification for this argument.
Next, we consider a related rank-based Anderson-Rubin-type statistic Anderson-Rubin(1949) studied in Andrews-Marmer(2008) under a super population framework. The model is
where $\alpha_s$ denotes a stratum-specific scalar parameter, $Y_{si}$ is the response variable, $D_{si}$ is a scalar endogenous regressor, $r_{Csi}$ is the error term which, unlike the previous case, is random, and the number of strata $S$ is not random. In addition, there is a $k$-vector of IVs, $Z_{si}$, with $k\geq 1$, that does not include a constant. The model in (ref) may arise, for example, due to stratification based on empirical support points $\{X_1,\dots, X_S\}$ of auxiliary regressors $X\in\mathbb{R}^p, p\geq 1,$ i.e. $\alpha_s\equiv X_s'\eta_s$, where $X_s\in\mathbb{R}^p$ and a parameter vector $\eta_s\in\mathbb{R}^p$ are constant within stratum $s$, but may vary across strata. The hypothesis of interest is again $H_0:\beta=\beta_0$. Let $R_{si}$ denote the rank of $Y_{si}-\beta_0 D_{si}$ among $Y_{s1}-\beta_0 D_{s1},\dots, Y_{sn_s}-\beta_0 D_{sn_s}$, and $\varphi:(0,1)\mapsto \mathbb{R}$ be a score function. The Andrews-Marmer(2008) statistic (Equation (3.3) therein) may be defined as
where $A_{n}=\sum_{s=1}^SA_{ns}$, $A_{ns}\equiv n^{-1}\sum_{i=1}^{n_s}(Z_{si}-\bar{Z}_s)\varphi\left(\frac{R_{si}}{n_s+1}\right)$, $\bar{Z}_s\equiv n_s^{-1}\sum_{i=1}^{n_s}Z_{si}$, and
with $\bar{\varphi}\equiv\int_{0}^{1}\varphi(x)dx$. Remark here that if $\{r_{Csi}\}_{i=1}^{n_s}$ are i.i.d. with a continuous distribution, $\{R_{si}\}_{i=1}^{n_s}$ are uniformly distributed over $\{1,2,\dots, n_s\}$ vanderVaart(1998), so we can reformulate $A_n$ as a stratified permutation statistic. We deviate from Andrews-Marmer(2008) in a minor way by considering the statistic
with a covariance matrix estimator $\Omega_n^{*}$ defined as
where ${\varphi}_{si}\equiv \varphi\left(\frac{i}{n_s+1}\right)$ and $\bar{\varphi}_s\equiv n_{s}^{-1}\sum_{i=1}^{n_s}\varphi_{si}$. The estimator in (ref) is preferred to that in (ref) because the former is the variance of $n^{1/2}A_{n}$ (with respect to the distribution of the ranks), whereas the covariance matrix in (ref) is an approximation valid only when the strata are large, a scenario considered by Andrews-Marmer(2008). In fact, when there are many small strata, the term $\int_{0}^{1}(\varphi(x)-\bar{\varphi})^2dx$ will involve a nonnegligible error as an approximation of $(n_s-1)^{-1}\sum_{i=1}^{n_s}\left(\varphi_{si}-\bar{\varphi}_s\right)^2$. By way of example, let $n_s=2$ for all $s=1,\dots, S$ and $\varphi(x)=\Phi^{-1}(x)$. Simple calculations yield $\bar{\varphi}_s=(\Phi^{-1}(1/3)+\Phi^{-1}(2/3))/2=0$ and
This is far below the value $\int_{0}^{1}(\varphi(x)-\bar{\varphi})^2dx=1$, so the covariance matrix estimator in (ref) is biased upward as an estimator of the asymptotic variance of $n^{1/2}A_{n}$ and the level-$\alpha$ test that rejects when $B_n$ exceeds the $1-\alpha$ quantile of $\chi^2_k$ distribution will be conservative. The LHS of (ref) (after rounding) is 0.56, 0.69, 0.89, 0.96, 0.98 for $n_s=5, 10, 50, 200, 500$, respectively, indicating a reduction in the approximation error with increasing strata sizes. A similar result holds for the Wilcoxon score as well.
The null asymptotic distribution of the statistic $B_{n}^{*}$ is provided in the proposition below.
As argued above, the test based on the statistic in (ref), which employs a $\chi^2_k$ critical value, may be conservative in the presence of many small strata, whereas the test based on the modified statistic in (ref) maintains an asymptotically correct level. This discrepancy is illustrated in a simple numerical experiment based on the following design:
where we first make $n$ independent draws of $Z_i\sim t_5(0, I_k)$ with $k=1,3$ corresponding to just and over-identified cases.\footnote{Here, $t_5(0, I_k)$ stands for the $k$-variate $t$-distribution with degrees of freedom $5$, zero mean and covariance matrix $I_k$.} For each $k$, we draw the stratification variable $X_i\sim \mathcal{U}(\{1,\dots, \frac{n}{r}\})/( \frac{n}{r})$, where $r$ varies over $\{2,5,10,25,80\}$. A smaller value of $r$ leads to a greater number of strata, $S$, equal to the empirical support size of $X_i$. We keep $\{(Z_i',X_i)'\}_{i=1}^n$ thus generated fixed over the simulation replications. For each replication, we let $u_i\sim t_1$, $\epsilon_i\sim t_1$, and $v_i=\rho u_i+\sqrt{1-\rho^2}\epsilon_i,$ where $\rho=0.5$ and $u_i$ and $\epsilon_i$ are mutually independent and i.i.d. across $i=1,\dots, n$. The values $\beta,\gamma_1,\gamma_2, \psi_1$ and $\psi_2$ are set to $0$, and $\pi=(1,\dots, 1)'\sqrt{\lambda/(nk)}\in\mathbb{R}^k$ with $\lambda=9$, indicating moderate identification strength of the IVs. The sample size is set relatively large at $n=400$ to elucidate the effect of the strata sizes embodied in $r$. After stratifying the synthetic data according to the realized values of $X_i$, we compute the following four statistics to test the null hypothesis $H_0:\beta=0$: the statistic in (ref) using both the normal and Wilcoxon score functions, denoted as $B_n^N$ and $B_n^W$, respectively, and their modified counterparts $B_n^{N*}$ and $B_n^{W*}$ based on (ref). Table (ref) displays the empirical size of the tests based on 5000 replications. The maximum stratum size $\max n_s$ and the number of strata $S$ are proportional and inversely proportional to $r$ respectively. The tests $B_n^N$ and $B_n^W$ under-reject when the strata sizes are small, although the latter shows less size distortion. Not surprisingly, the rejection rates of the tests improve as the strata size increases. In contrast, the modified versions $B_n^{N*}$ and $B_n^{W*}$ exhibit reasonably accurate rejection rates in all cases considered, consistent with the theoretical predictions.\footnote{The result remains qualitatively similar when $Z_i$ and $X_i$ are randomly drawn in each replication, and is available upon request.}