The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
36,135 characters
A Combinatorial Central Limit Theorem for Stratified Randomization
{\title{{A Combinatorial Central Limit Theorem for Stratified Randomization}\thanks{The author thanks Xavier D'Haultf\oe{}uille for helpful discussions that led to this work.}}
\date{}
\author{
Purevdorj Tuvaandorj\thanks{York University, {[email removed].}}
}
}
\maketitle
\begin{abstract}
This paper establishes a combinatorial central limit theorem for stratified randomization, which holds under a Lindeberg-type condition. The theorem allows for an arbitrary number or sizes of strata, with the sole requirement being that each stratum contains at least two units. This flexibility accommodates both a growing number of large and small strata simultaneously, while imposing minimal conditions. We then apply this result to derive the asymptotic distributions of two test statistics proposed for instrumental variables settings in the presence of potentially many strata of unrestricted sizes.
\begin{description}
\item[Keywords:] Combinatorial central limit theorem, permutation test, randomization inference, stratification.
\end{description}
\end{abstract}
\section{Introduction}
Stratified randomization is frequently utilized to enhance covariate balance in randomized experiments.
In this scheme, blocks or strata are initially created based on pre-treatment variables, and randomization inference is subsequently conducted by randomizing units within each stratum \citep[see e.g.,][Chapter~9]{Imbens-Rubin(2015)}. Stratified randomization can be used for obtaining exact tests \citep{Imbens-Rosenbaum(2005), Andrews-Marmer(2008), Zhao-Ding2021, DHT2023} and efficient estimators \citep{Cytrynbaum2021, Armstrong2022, BLSTM2023}, and is a natural inferential choice for models with survey data obtained through stratified sampling.\par
The key tool used when establishing the asymptotic validity of the permutation tests is
the combinatorial central limit theorem (CLT) \citep{Hoeffding1951}.
While the standard combinatorial CLT or finite population CLT can be applied to derive the asymptotic distributions of stratified permutation statistics when there is a finite number of large strata
\citep{Imbens-Rosenbaum(2005), LiDing2017}, its application to general stratified permutations is not straightforward. This complication arises because the number of strata, either large or small, or both, may \emph{grow} with the sample size (for instance, consider the case of having $n^{1/2}$ many strata, each of size $n^{1/2}$, where $n$ denotes the sample size).\par
Despite recent contributions by \cite{LiuYang2020} and \cite{LiuRenYang2022} that propose finite population CLTs tailored for such scenarios, a combinatorial CLT accommodating potentially many strata of unrestricted size remains essential for addressing general stratified permutation/rank statistics. In this context, this paper introduces a general form of the combinatorial CLT for stratified randomization, applicable under a \emph{Lindeberg-type} condition. The proof is based on Stein's method \citep[see, e.g.,][for a detailed account]{CGS2011}, and in particular, draws from the constructions of \cite{Bolthausen1984} and \cite{Schneller1988}. Leveraging this result, we characterize the asymptotic distribution of a test statistic proposed by \cite{Imbens-Rosenbaum(2005)} and a variant of a test statistic by \cite{Andrews-Marmer(2008)} in the instrumental variables (IVs) setting featuring many strata, under weak assumptions.\par
A substantial body of literature exists on stratified randomized experiments and large sample results for stratified randomization. For empirical practices and an overview, see, for example, \cite{BruhnMcKenzie2009}, \cite{Imbens-Rubin(2015)} and \cite{Ding2023}. Papers that consider a coarse stratification involving a finite number of strata, where asymptotic normality is typically derived via finite population or rank CLTs, include \cite{Andrews-Marmer(2008)}, \cite{LiDing2017}, \cite{Zhao-Ding2021}. Papers that consider a finely stratified experiment with many strata, where the asymptotic normality is derived via Lindeberg or Lyapunov CLTs, include \cite{Fogarty2018mit}, \cite{Bai2022} and \cite{BaiRomanoShaikh2022}.\par
Furthermore, while \cite{LiuYang2020} and \cite{LiuRenYang2022} provide finite population CLTs for stratified experiments involving many strata, \cite{DHT2023} establish a combinatorial CLT with many strata to examine the properties of a stratified permutation subvector test in linear regressions. The CLT for stratified randomization in this paper is sufficiently general to encompass the aforementioned CLTs as it handles both large and small strata simultaneously and requires a weaker Lindeberg-type condition. In particular, finite population CLTs with many strata follow from the result akin to the derivation of usual finite population CLTs from standard combinatorial or rank CLTs.\par
Stratified permutation is an instance of permutations with \emph{restricted positions} \citep[see, e.g.,][]{Rosenbaum1984,Diaconis2001,LeiBickel2021}; thus, the set of such permutations forms a subset of all permutations of observation indices. \citet[][Chapter 6]{CGS2011} provide normal approximation results for a class of restricted permutations {\textemdash} in contrast with the set of entire permutations required in \citet{Hoeffding1951}'s CLT {\textemdash} but do not consider stratified permutations. The present paper fills this gap in the literature by developing normal approximation for this important class of permutations.\par
This paper proceeds as follows. Section \ref{sec: Main} lays out the basic framework and presents the combinatorial CLT for stratified randomization with an arbitrary number of strata. Section \ref{sec: link FPCLT} establishes a connection between our result and finite population CLTs. Section \ref{sec: IV} applies the combinatorial CLT to derive the asymptotic distribution of two test statistics in IV settings. The proofs are provided in the appendix.
\section{Main result}\label{sec: Main}
Let $n$ be a fixed sample size, ${S}$ be an integer-valued random variable (${S}\ge 1$ a.s.), and $s=1,\dots,{S}$ denote strata of sizes $n_s\ge 2$ a.s., with $\sum_{s=1}^{{S}} n_s=n$. Consider a double array of real-valued random variables in stratum $s$: $\{a_{ij}^s\}_{i,j=1}^{n_s}$, and define $\bm{a}_{n}\equiv \{a_{ij}^s\}_{s=1,\dots, {S}; (i,j)\in\{1,\dots, n_s\}^2}$. Here, the distributions of $\{a_{ij}^s\}_{i,j=1}^{n_s}$ and ${S}$ may depend on $n$, and the stratification may be based on random auxiliary covariates. \par
Let $\pi$ be a \emph{stratified} permutation that permutes the indices within each stratum:
\begin{equation*}
\pi=
\underbrace{\begin{pmatrix}
1&\dots &n_1\\
\pi_1(1)&\dots &\pi_1(n_1)
\end{pmatrix}}_{\text{Stratum $1$}}
\dots
\underbrace{
\begin{pmatrix}
1&\dots &n_{{S}}\\
\pi_S(1)&\dots &\pi_{{S}}(n_{{S}})
\end{pmatrix}}_{\text{Stratum ${S}$}},
\end{equation*}
where $\pi_s$ is a permutation of $\{1,\dots, n_s\}$ in stratum $s=1,\dots,{S}$.\footnote{For simplicity, we do not further index the elements of $\{1,\dots, n_s\}$ by $s$ throughout the paper.}
We denote by $\mathbb{S}_n$ the set of all such permutations with $|\mathbb{S}_n|=\prod_{s=1}^{{S}}n_s!$. In this paper, we consider permutations uniformly distributed over $\mathbb{S}_n$ for $\bm{a}_{n}$ given: $\pi\sim\mathcal{U}(\mathbb{S}_n)$, where $\mathcal{U}(A)$ denotes the uniform distribution on a finite set $A$. Let $P^\pi$ be the probability measure of $\pi$ conditional on $\bm{a}_{n}$, and $\operatorname{E}_\pi[\cdot]$ and $\operatorname{Var}_\pi[\cdot]$ denote the corresponding expectation and variance operators.
Furthermore, we use the notation $1(\cdot)$ for the indicator function, and
$\Phi(\cdot)$ and $\Phi^{-1}(\cdot)$ for the cumulative distribution and quantile function of a standard real normal distribution. $\sum_{i,j=1}^{n_s}$ and $\sum_{i\neq j}$ are the shortcuts for $\sum_{i=1}^{n_s}\sum_{j=1}^{n_s}$ and $\sum_{i,j=1; i\neq j}^{n_s}$, respectively. We abbreviate ``left-hand side'' and ``right-hand side'' as LHS and RHS. Let $\Vert\cdot\Vert$ denote the Frobenius norm, $\lambda_{\min}(\cdot)$ denote the minimum eigenvalue of a square matrix, and $\max_{s, i}$ denote for the maximum over $s\in\{1,\dots,{S}\}$ and $i\in\{1,\dots,n_s\}$ for a given $n$. \par
\par
The main result of this paper is the following combinatorial CLT for stratified randomization which is based on a Lindeberg-type condition.
\begin{theorem}[Combinatorial CLT for stratified randomization]\label{HoeffdingCLT}
Let ${S}$ be an integer-valued random variable such that ${S}\ge 1$ a.s. and $s=1,\dots,{S}$ denote strata of sizes $n_s\ge 2$ a.s., with $\sum_{s=1}^{{S}} n_s=n$. Let $\pi\sim\mathcal{U}(\mathbb{S}_n)$ and for each $s$, $\{a_{ij}^s\}_{i,j=1}^{n_s}$ be $n_s^2$ scalar random variables satisfying:
\begin{enumerate}[label=(\alph*)]
\item\label{cond: centered}$\sum_{i=1}^{n_s}a_{ij}^{s}=\sum_{j=1}^{n_s}a_{ij}^s=0$ for $1\leq i,j\leq n_s$ and $s=1,\dots, {S}$ a.s.;
\item\label{cond: Lindeberg} For any $\epsilon>0$, $\sigma_n^{-2}\sum_{s=1}^{{S}} \frac{1}{n_s}\left(\sum_{i,j=1}^{n_s}a_{ij}^{s2}1(\sigma_n^{-1}|a_{ij}^s|>\epsilon)\right)\stackrel{a.s.}{\longrightarrow} 0$ as $n\to\infty$, where
$\sigma_n^2\equiv \sum_{s=1}^{{S}} \frac{1}{n_s-1}\left(\sum_{i,j=1}^{n_s}a_{ij}^{s2}\right)$ and $a_{ij}^s\neq 0$ for some $(i,j)$ and $s$.
\end{enumerate}
Let $T^\pi\equiv \sum_{s=1}^{{S}} \sum_{i=1}^{n_s} a^{s}_{i\pi(i)}/\sigma_n$. Then, as $n\to\infty$
\begin{equation}\label{eq: CCLT}
P^\pi(T^\pi\leq t)\stackrel{a.s.}{\longrightarrow} \Phi(t)\ \ \text{for any}\ \ t\in\mathbb{R}.
\end{equation}
\end{theorem}
\begin{remark}\label{remark: R}~
\normalfont
\begin{enumerate}[label={(\arabic{enumi})},leftmargin=*]
\item\label{R1}When ${S}=1$ with probability 1, $n_{{S}}=n$ and the result specializes to Theorem 1 of \cite{Motoo1956}. The latter provides a version of \cite{Hoeffding1951}'s CLT based on a Lindeberg-type condition that corresponds to the condition in \ref{cond: Lindeberg}.
\item\label{R2} The key ingredient in the proof of Theorem \ref{HoeffdingCLT} is the random selection of a stratum $s$ with probability $p_s\equiv n_s/n$. We then bound $|\operatorname{E}_\pi[T^\pi f(T^\pi)]-\operatorname{E}_\pi[f'(T^\pi)]|$, where $f(\cdot)$ is defined in \eqref{def: f} below. This is achieved by perturbing the sum $\sum_{i=1}^{n_s} a^{s}_{i\pi(i)}$ within the stratum $s$, following the approach in \cite{Bolthausen1984} and \cite{Schneller1988}, and then ``averaging'' over the strata $s=1,\dots, S$.
\item\label{R3} Since $n_s\geq 2$ a.s., ${S}\leq n/2$ and $n-{S}\geq n/2$ a.s. On noting that $n-{S}=\sum_{s=1}^{S}(n_s-1)\leq \prod_{s=1}^{{S}}n_s\leq \vert\mathbb{S}_n\vert$ a.s. and $n/2\to \infty$, we have $|\mathbb{S}_n|\stackrel{a.s.}{\longrightarrow} \infty$.
\item\label{R4} A Lyapunov-type sufficient condition for the Lindeberg condition in \ref{cond: Lindeberg} is
\begin{equation}\label{cond: Lyap}
\frac{1}{\sigma_n^{2+\delta}}\sum_{s=1}^{{S}} \frac{1}{n_s}\sum_{i,j=1}^{n_s}|a_{ij}^{s}|^{2+\delta}\stackrel{a.s.}{\longrightarrow} 0
\end{equation}
as $n\to\infty$, for some $\delta>0$.
\item\label{R5} The randomness of $T^\pi$ stems from $\bm{a}_{n}$ and $\pi$. If $\bm{a}_{n}$ are random variables such that the condition \ref{cond: centered} holds a.s. and the convergence in \ref{cond: Lindeberg} holds in probability, then using a subsequencing argument, the convergence in \eqref{eq: CCLT} can be restated as
\begin{equation}\label{eq: CCLT2}
P^\pi(T^\pi\leq t)\stackrel{p}{\longrightarrow} \Phi(t).
\end{equation}
\item\label{R6}
Consider the case $a_{ij}^s=\tilde{b}_{si}\tilde{c}_{sj}$, where $\tilde{b}_{si}\equiv b_{si}-\bar{b}_s$, $\bar{b}_s\equiv n_{s}^{-1}\sum_{i=1}^{n_s}b_{si}$, $\tilde{c}_{sj}\equiv c_{sj}-\bar{c}_s$, $\bar{c}_s\equiv n_{s}^{-1}\sum_{j=1}^{n_s}c_{sj}$ and $b_{si}$ and $c_{sj}$ are scalars.
\cite{DHT2023} (see Lemma 4 therein) establish a combinatorial CLT for stratified randomization based on the following assumptions: as $n\to\infty$
\begin{enumerate}[label=(\alph*)]
\item\label{cond: sigcon} For $\sigma_n^2\equiv\sum_{s=1}^{{S}} \frac{1}{n_s-1}\left(\sum_{i=1}^{n_s}\tilde{b}_{si}^{2}\right)\left(\sum_{i=1}^{n_s}\tilde{c}_{si}^{2}\right)$,
$\sigma_n^2\stackrel{p}{\longrightarrow} \sigma^2>0$;
\item \label{cond: 3mom} $\sum_{s=1}^{{S}}n_s^{-1} \left(\sum_{i=1}^{n_s} |\tilde{b}_{si}|^3\right)
\left(\sum_{i=1}^{n_s} |\tilde{c}_{si}|^3\right)\stackrel{p}{\longrightarrow} 0$ ;
\item\label{cond: Uvarcon}
$\sum_{s=1}^{{S}}(n_s-1)^{-1}\left(\sum_{i=1}^{n_s}\tilde{b}_{si}^{4}\right)\left(\sum_{i=1}^{n_s}\tilde{c}_{si}^{4}\right)\stackrel{p}{\longrightarrow} 0$.
\end{enumerate}
From the conditions \ref{cond: sigcon} and \ref{cond: 3mom}, we have
\begin{equation}\label{cond: Lyap2}
\frac{1}{\sigma_n^{3}}\sum_{s=1}^{{S}} \frac{1}{n_s}\sum_{i,j=1}^{n_s}|a_{ij}^{s}|^{3}
=\frac{1}{\sigma_n^{3}}\sum_{s=1}^{{S}} \frac{1}{n_s}\left(\sum_{i=1}^{n_s}|\tilde{b}_{si}|^{3}\right)\left(\sum_{i=1}^{n_s}|\tilde{c}_{si}|^{3}\right)\stackrel{p}{\longrightarrow} 0.
\end{equation}
Thus, the condition in \eqref{cond: Lyap} holds in probability with $\delta=1$, and \eqref{eq: CCLT2} follows from a subsequencing argument.
\end{enumerate}
\end{remark}
\noindent A leading special case of interest similar to that considered in Remark \ref{remark: R}\ref{R6} is the asymptotic normality of the sum $n^{-1/2}\sum_{s=1}^{{S}}\sum_{i=1}^{n_s}\tilde{b}_{si}\tilde{c}_{s\pi(i)}$, where $\tilde{b}_{si}\equiv b_{si}-\bar{b}_s$, $\bar{b}_s\equiv n_{s}^{-1}\sum_{i=1}^{n_s}b_{si}$, $b_{si}\in\mathbb{R}^k$ with $k\geq 1$, $\tilde{c}_{si}\equiv c_{si}-\bar{c}_s$, $c_{si}\in\mathbb{R}$ and $\bar{c}_s\equiv n_{s}^{-1}\sum_{i=1}^{n_s}c_{si}$. The following result holds as a corollary to Theorem \ref{HoeffdingCLT}.
\begin{corollary}\label{cor: vector bc}
Suppose that $\lambda_{\min}(\Sigma_n)>\lambda>0$ for $\Sigma_n\equiv n^{-1}\sum_{s=1}^{{S}} \frac{1}{n_s-1}\left(\sum_{i=1}^{n_s}\tilde{b}_{si}\tilde{b}_{si}'\right)\left(\sum_{j=1}^{n_s}\tilde{c}_{sj}^{2}\right)$, and for some positive constants $\delta$ and $M_0$, either
\begin{enumerate}[label=(\alph*)]
\item\label{cond: Vector bc1}
$n^{-1}\sum_{s=1}^{{S}}\sum_{i=1}^{n_s}\Vert {b}_{si}\Vert^{4+\delta}<M_0$ a.s. and
$n^{-1}\sum_{s=1}^{{S}}\sum_{i=1}^{n_s}|{c}_{si}|^{4+\delta}<\infty$ a.s. for all $n$; or
\item\label{cond: Vector bc2} $n^{-1/2}\max_{s,i}\Vert b_{si}\Vert\stackrel{a.s.}{\longrightarrow} 0$,
$n^{-1}\sum_{s=1}^{{S}}\sum_{i=1}^{n_s}\Vert {b}_{si}\Vert^{2}<M_0$ a.s.
and $\max_{1\leq s\leq S}n_s^{-1}\sum_{i=1}^{n_s}\vert {c}_{si}\vert^{2+\delta}<M_0$ a.s. for all $n$.
\end{enumerate}
Then, as $n\to\infty$
\begin{equation}\label{eq: Vector AN}
\Sigma_n^{-1/2}n^{-1/2}\sum_{s=1}^{{S}}\sum_{i=1}^{n_s}\tilde{b}_{si}\tilde{c}_{s\pi(i)}\displaystyle \stackrel{d}{\longrightarrow} \mathcal{N}\left(0, I_k\right)\quad\text{a.s}.
\end{equation}
\end{corollary}
\section{Relation to finite population CLT for stratified randomization}\label{sec: link FPCLT}
This section explores the connection between the combinatorial CLT of Theorem \ref{HoeffdingCLT} and a finite population CLT.
\cite{Bickel1984} and \cite{LiuYang2020} establish finite population CLTs for linear combinations of stratum means under stratified sampling, where the number of strata may diverge.\par
Consider a finite population divided into non-random ${S}$ strata: $(y_{s1},\dots, y_{sn_s}), s=1,\dots, {S}$. The stratum-specific mean and variance are defined respectively as $\bar{y}_s\equiv n_s^{-1}\sum_{i=1}^{n_s}y_{si}$ and $v_{s}^2\equiv \frac{1}{n_s-1}\sum_{i=1}^{n_s}(y_{si}-\bar{y}_s)^2$, while the population mean and weighted variance are given by
\begin{equation}\label{eq: popmeanvar}
\bar{y}\equiv n^{-1}\sum_{s=1}^S\sum_{i=1}^{n_s}y_{si}=\sum_{s=1}^Sp_s\bar{y}_s\ \ \text{and}\ \
v^2\equiv \sum_{s=1}^{S}p_s^2v_s^2\frac{n_s-n_{1s}}{n_{1s}n_s},
\end{equation}
where $p_s\equiv n_s/n$ and $n_{1s}, 2\leq n_{1s}\leq n_s-1,$ denotes the number of units sampled without replacement from stratum $s$ in \cite{Bickel1984}'s set-up (the
number of treated units in stratum $s$ in \cite{LiuYang2020}'s set-up). The sampling indicators are $(Z_{s1},\dots, Z_{sn_s})\in\{0,1\}^{n_s}, s=1,\dots, S$, where $Z_{si}=1$ if the unit $i$ in stratum $s$ is sampled, and $0$ otherwise. The probability that $\{(Z_{s1},\dots,Z_{sn_s})\}_{s=1}^S$ takes on a value
$\{(z_{s1},\dots,z_{sn_s})\}_{s=1}^S$ such that $n_{1s}=\sum_{i=1}^{n_s}z_{si}$ for $s=1,\dots, S,$ is $\prod_{s=1}^Sn_{1s}!(n_s-n_{1s})!/n_s!$. The weighted sample mean and variance are then defined as
\begin{equation*}
\hat{y}\equiv \sum_{s=1}^Sp_s\hat{y}_s\ \ \text{and}\ \ \hat{v}^2\equiv \sum_{s=1}^Sp_s^2\hat{v}_{s}^2\frac{n_s-n_{1s}}{n_{1s}n_s},
\end{equation*}
where $\hat{y}_s\equiv n_{1s}^{-1}\sum_{i=1}^{n_s}Z_{si}y_{si}$ and $\hat{v}_{s}^2
\equiv \frac{1}{n_s-1}\sum_{i=1: Z_{si}=1}^{n_s}(y_{si}-\hat{y}_s)^2$ are the stratum-specific sample mean and variance.
With the notations introduced, the Lindeberg condition employed by \cite{LiuYang2020s} (Condition A1 and Theorem A1 therein) may be restated as follows: for any $\epsilon>0$
\begin{equation}\label{eq: LYLind}
\sum_{s=1}^{{S}} \frac{1}{n_s-1}\left(\sum_{i=1}^{n_s}w_s^2\frac{(y_{si}-\bar{y}_s)^2}{v_s^2}1\left(w_s\frac{|y_{si}-\bar{y}_s|}{v_s}>\epsilon \sqrt{\frac{n_{1s}(n_s-n_{1s})}{n_s}}\right)\right)\to 0
\end{equation}
as $n\to\infty$, where
\begin{equation}\label{eq: ws Vs}
w_s^2\equiv \frac{p_s^2v_s^2(n_s-n_{1s})}{v^2n_sn_{1s}}.
\end{equation}
Under the condition in \eqref{eq: LYLind}, \cite{Bickel1984} show that $(\hat{y}-\bar{y})/{\hat{v}}\displaystyle \stackrel{d}{\longrightarrow} \mathcal{N}\left(0,1\right)$ as $n\to\infty$. To draw a link between the latter result and Theorem \ref{HoeffdingCLT}, we let $a_{ij}^s=\tilde{b}_{si}\tilde{c}_{sj}$ as in
Remark \ref{remark: R}\ref{R6} and Corollary \ref{cor: vector bc}, where $b_{si}$ and $c_{sj}$ are now defined as
\begin{equation}\label{eq: LYb}
b_{si}\equiv
\begin{cases}
\frac{1}{n_{1s}}&\text{if}\quad i=1,\dots, n_{1s},\\
0&\text{if}\quad i=n_{1s}+1,\dots, n_s
\end{cases},
\quad c_{sj}\equiv p_sy_{sj}.
\end{equation}
The choice of $b_{si}$ in \eqref{eq: LYb} appears in \cite{Madow1948} and \cite{Sen1995}, among others.
By simple algebra, $\bar{b}_s=n_s^{-1}\sum_{i=1}^{n_s}b_{si}=n_s^{-1}$, $\bar{c}_s= n_s^{-1}\sum_{j=1}^{n_s}c_{sj}=p_s\bar{y}_s$, and
\begin{equation}\label{eq: LYb2}
\tilde{b}_{si}= b_{si}-\bar{b}_s=
\begin{cases}
\frac{1}{n_{1s}}-\frac{1}{n_s}&\text{if}\quad i=1,\dots, n_{1s},\\
-\frac{1}{n_s}&\text{if}\quad i=n_{1s}+1,\dots, n_s
\end{cases},\quad
\tilde{c}_{sj}= c_{sj}-\bar{c}_s= p_s({y}_{sj}-\bar{y}_s).
\end{equation}
It turns out that the variance
$\sigma_n^2$ defined in the condition \ref{cond: Lindeberg} of Theorem \ref{HoeffdingCLT} and $v^2$ in \eqref{eq: popmeanvar} are identical:
\begin{equation}\label{eq: vareq}
\sigma_n^2
=\sum_{s=1}^{{S}} \frac{1}{n_s-1}\left(\sum_{i=1}^{n_s}\tilde{b}_{si}^{2}\right)\left(\sum_{i=1}^{n_s}\tilde{c}_{si}^{2}\right)=\sum_{s=1}^{{S}} \frac{1}{n_s-1}\frac{(n_s-n_{1s})p_s^2}{n_{1s}n_s}\left(\sum_{i=1}^{n_s}({y}_{si}-\bar{y}_s)^{2}\right)=v^2.
\end{equation}
Then, through a direct calculation given in Section \ref{sec:R2}, the Lindeberg condition in Theorem \ref{HoeffdingCLT}\ref{cond: Lindeberg} follows from the condition in \eqref{eq: LYLind}, as summarized in the remark below.
\begin{remark}\label{remark: R2}
\normalfont
The condition in \eqref{eq: LYLind} implies the Lindeberg condition in Theorem \ref{HoeffdingCLT}\ref{cond: Lindeberg}.
\end{remark}
\section{Stratified randomization inference with IVs}\label{sec: IV}
This section presents two applications of the CLT in Theorem \ref{HoeffdingCLT}. First, we establish the asymptotic distribution of a statistic proposed by
\cite{Imbens-Rosenbaum(2005)}. Second, we show that the test of \cite{Andrews-Marmer(2008)} may be conservative in the presence of small strata. We then consider a simple modification of their statistic that yields a correct asymptotic level and derive its null asymptotic distribution.\par
\subsection{\cite{Imbens-Rosenbaum(2005)} statistic}
\cite{Imbens-Rosenbaum(2005)} consider an IV setting with non-random $S$ strata $s=1,\dots, S$ each of size $n_s$, given by: $$Y-\beta D=r_C,$$ where
$D=[D_1',\dots, D_{S}']\in\mathbb{R}^n$ with $D_s=[D_{s1},\dots, D_{sn_s}]'\in\mathbb{R}^{n_s}$ and $n=\sum_{s=1}^Sn_s$, denotes the vector of doses received by the units,
$Y=[Y_1',\dots, Y_S']'\in\mathbb{R}^n$ with $Y_s=[Y_{s1},\dots, Y_{sn_s}]'\in\mathbb{R}^{n_s}$ is the vector of responses exhibited by the units,
$r_C=[r_{C1}',\dots, r_{CS}']'\in\mathbb{R}^n$ with $r_{Cs}=[r_{Cs1},\dots, r_{Csn_s}]'\in\mathbb{R}^{n_s}$ is the vector of responses if the units received the control dose, and the scalar parameter $\beta$, with a true value $\beta_0$, measures of the dose-response relationship.
\par
Let $h_s=[h_{s1},\dots, h_{sn_s}]'\in\mathbb{R}^{n_s}$ be a sorted and fixed IV vector such that
$h_{sj}\leq h_{s(j+1)}$ for each $j=1,\dots, n_s-1$ and $s=1,\dots,S$. For $h\equiv [h_1',\dots, h_S']'\in\mathbb{R}^{n}$, assume that the observed IV $Z$ is randomly assigned i.e.
$Z=h_\pi\in\mathbb{R}^n$, where $\pi\sim \mathcal{U}(\mathbb{S}_n)$ and $\mathbb{S}_n$ is the set of all stratified permutations in this setting.\par
Here, $r_C=Y-\beta_0 D$ is assumed to be fixed, hence independent of $Z$. For testing the hypothesis $H_0:\beta=\beta_0$,
\cite{Imbens-Rosenbaum(2005)} propose the following statistic:
\begin{equation}\label{def: T}
T=q(Y-\beta_0D)'\rho(Z),
\end{equation}
where $q(\cdot)\in\mathbb{R}^n$ is some method of scoring responses such as the ranks or the aligned ranks, and
$\rho(\cdot)\in\mathbb{R}^n$ is some method of scoring the instruments that satisfies $\rho_\pi(h)=\rho(h_\pi)=\rho(Z)$. Let $\rho(h)=[\rho_{1}',\dots, \rho_S']'$ with $\rho_s=[\rho_{s1},\dots, \rho_{sn_s}]'$, and $q(Y-\beta_0D)=[q_{1}',\dots, q_{S}']'$ with $q_s=[q_{s1},\dots, q_{sn_s}]'$. \par
Before we state the result for the asymptotic distribution of the statistic $T$, we remark that
$T=\sum_{s=1}^S\sum_{i=1}^{n_s}q_{si}\rho_{s\pi(i)}$ using the notations above. With a convention that $\operatorname{Var}_\pi[\sum_{i=1}^{n_s}q_{si}\rho_{s\pi(i)}]=0$ for stratum of size $n_s=1$, it follows from Proposition 1 of \cite{Imbens-Rosenbaum(2005)} that
\begin{align}
\mu
&\equiv \operatorname{E}_\pi[T]=\sum_{s=1}^S{n_s}\bar{q}_{s}\bar{\rho}_{s},\label{eq: MeanT}\\
\sigma^2\equiv \operatorname{Var}_\pi[T]
&=\sum_{s=1, n_s\geq 2}^S\frac{1}{n_s-1}\left(\sum_{i=1}^{n_s}\tilde{q}_{si}^2\right)\left(\sum_{i=1}^{n_s}\tilde{\rho}_{si}^2\right),\label{eq: varT}
\end{align}
where $\bar{q}_s\equiv n_s^{-1}\sum_{i=1}^{n_s}q_{si}$, $\bar{\rho}_s\equiv n_s^{-1}\sum_{i=1}^{n_s}\rho_{si}$,
and $\tilde{q}_{si}\equiv q_{si}-\bar{q}_s$ and $\tilde{\rho}_{si}\equiv \rho_{si}-\bar{\rho}_s$. \par
\cite{Imbens-Rosenbaum(2005)} outline strategies for deriving the asymptotic distribution of the statistic $T$ in the presence of either a few large strata whose sizes tend to infinity, or many small strata, whose sizes remain bounded, and suggest applying the rank-central limit theorem \citep[see e.g.,][Chapter 6]{Hajek-Sidak-Sen(1999)} to the former and standard CLTs to the latter.
\cite{LiDing2017} consider the case where $q(\cdot)$ is an identity function and $\rho(Z)$ is a vector of binary IVs, and derive the null asymptotic distribution of the statistic in \eqref{def: T}. The general CLT provided in Theorem \ref{HoeffdingCLT} allows us to derive the asymptotic distribution of the statistic $T$ in \eqref{def: T} at once under general conditions, allowing for many strata without categorizing them into large and small ones or restricting their sizes. The following proposition gives a formal justification for this argument.
\begin{proposition}\label{prop: SRIVa}
Suppose that $\displaystyle \operatornamewithlimits{\lim\inf \ }_{n\to\infty}\sigma^2>\lambda>0$, and for some positive constants $M_0$ and $\delta$, either
\begin{enumerate}[label=(\alph*)]
\item\label{Asbound2} $n^{-1}\sum_{s=1}^S\sum_{i=1}^{n_s}|\rho_{si}|^{4+\delta}<M_0$ and $n^{-1}\sum_{s=1}^S\sum_{i=1}^{n_s}|q_{si}|^{4+\delta}<M_0$ for all $n$; or
\item\label{Asbound1} $n^{-1/2}\max_{s,i}|\rho_{si}|\to 0$, $n^{-1}\sum_{s=1}^S\sum_{i=1}^{n_s}|\rho_{si}|^{2}<M_0$ and $\max_{1\leq s\leq S}n_s^{-1}\sum_{i=1}^{n_s}|{q}_{si}|^{2+\delta}<M_0$ for all $n$.
\end{enumerate}
Then, under $H_0:\beta=\beta_0$, ${\sigma}^{-1}(T-\mu)\displaystyle \stackrel{d}{\longrightarrow} \mathcal{N}\left(0,1\right)$ as $(n-S)\to \infty$.
\end{proposition}
\begin{remark}\label{rem: ImbensRosenbaum}~
\normalfont
\begin{enumerate}[label={(\arabic{enumi})},leftmargin=*]
\item The strata of size $n_s=1$ are discarded in the statistic ${\sigma}^{-1}(T-\mu)$. Since $n-S$, which corresponds to the effective sample size, tends to infinity, so does the number of observations stratified into strata of sizes $2$ or greater.
Consequently, the condition $n\to\infty$ in Theorem \ref{HoeffdingCLT} can be replaced by $(n-S)\to \infty$.
For $S$ random, a sufficient condition for $(n-S)\stackrel{a.s.}{\longrightarrow} \infty$ is provided in Lemma 2 of \cite{DHT2023}.
\item Assumption \ref{Asbound1} on $q_{si}$ holds, for example, if $q_{si}=\varphi(i/(n_s+1))$, where $\varphi:(0,1)\mapsto \mathbb{R}$ is a score function
satisfying $\int_{0}^{1}\vert {\varphi}(x)\vert^{2+\delta}dx<M_0/2$. This is because $(n_s+1)/n_s\leq 2$ and
\begin{equation}\label{eq: int ineq}
(n_s+1)^{-1}\sum_{i=1}^{n_s}|\varphi(i/(n_s+1))|^{2+\delta}\leq \int_{0}^{1}|\varphi(x)|^{2+\delta}dx,
\end{equation}
as the LHS of the inequality above is the sum of the areas of $n_s$
rectangles with sides $|\varphi(i/(n_s+1))|^{2+\delta}$ and $[i/(n_s+1), (i+1)/(n_s+1)]$, $i=1,\dots, n_s,$ which lie between the function $|\varphi(x)|^{2+\delta}$ and $[0,1]$.\par
For the Wilcoxon score $\varphi(x)=x$, letting $\delta=1$, the RHS of \eqref{eq: int ineq} is $\int_{0}^{1}x^3dx=1/4$. For the normal (or van der Waerden) score $\varphi(x)=\Phi^{-1}(x)$, letting $\delta=2$, we have, through a change of variables,
$\int_{0}^{1}(\Phi^{-1}(x))^4dx=3$. So the condition on $q_{si}$ in Assumption \ref{Asbound1} is satisfied for these score functions.
\end{enumerate}
\end{remark}
\subsection{\cite{Andrews-Marmer(2008)} statistic}
Next, we consider a related rank-based Anderson-Rubin-type statistic \citep{Anderson-Rubin(1949)} studied in \citet[][Section~3]{Andrews-Marmer(2008)} under a super population framework.
The model is
\begin{equation}\label{eq: AMmodel}
Y_{si}=\alpha_s+D_{si}\beta+r_{Csi},\quad i=1,\dots, n_s;\ s=1,\dots, S,
\end{equation}
where $\alpha_s$ denotes a stratum-specific scalar parameter, $Y_{si}$ is the response variable, $D_{si}$ is a scalar endogenous regressor, $r_{Csi}$ is the error term which, unlike the previous case, is random, and the number of strata $S$ is not random. In addition, there is a $k$-vector of IVs, $Z_{si}$, with $k\geq 1$, that does not include
a constant. The model in \eqref{eq: AMmodel} may arise, for example, due to stratification based on empirical support points
$\{X_1,\dots, X_S\}$ of auxiliary regressors $X\in\mathbb{R}^p, p\geq 1,$
i.e.
$\alpha_s\equiv X_s'\eta_s$, where $X_s\in\mathbb{R}^p$ and a parameter vector $\eta_s\in\mathbb{R}^p$ are constant within stratum $s$, but may vary across strata. The hypothesis of interest is again $H_0:\beta=\beta_0$.\par
Let $R_{si}$ denote the rank of $Y_{si}-\beta_0 D_{si}$ among $Y_{s1}-\beta_0 D_{s1},\dots, Y_{sn_s}-\beta_0 D_{sn_s}$, and
$\varphi:(0,1)\mapsto \mathbb{R}$ be a score function. The
\cite{Andrews-Marmer(2008)} statistic (Equation (3.3) therein) may be defined as
\begin{equation}\label{def: AMRAR}
B_{n}
\equiv n\,A_{n}'\Omega_{n}^{-1}A_{n},
\end{equation}
where $A_{n}=\sum_{s=1}^SA_{ns}$, $A_{ns}\equiv n^{-1}\sum_{i=1}^{n_s}(Z_{si}-\bar{Z}_s)\varphi\left(\frac{R_{si}}{n_s+1}\right)$,
$\bar{Z}_s\equiv n_s^{-1}\sum_{i=1}^{n_s}Z_{si}$, and
\begin{equation}\label{eq: AMcovm}
\Omega_{n}
\equiv n^{-1}\sum_{s=1}^{S}\sum_{i=1}^{n_s}(Z_{si}-\bar{Z}_s)(Z_{si}-\bar{Z}_s)'\int_{0}^{1}(\varphi(x)-\bar{\varphi})^2dx,
\end{equation}
with $\bar{\varphi}\equiv\int_{0}^{1}\varphi(x)dx$. Remark here that if $\{r_{Csi}\}_{i=1}^{n_s}$ are i.i.d. with a continuous distribution, $\{R_{si}\}_{i=1}^{n_s}$ are uniformly distributed over $\{1,2,\dots, n_s\}$ \citep[see e.g.,][Lemma 13.1]{vanderVaart(1998)}, so we can reformulate $A_n$ as a stratified permutation statistic. We deviate from \cite{Andrews-Marmer(2008)} in a minor way by considering the statistic
\begin{equation}\label{def: Bnstar}
B_{n}^{*}\equiv n\,A_n'\Omega_n^{*-1}A_n,
\end{equation}
with a covariance matrix estimator $\Omega_n^{*}$ defined as
\begin{equation}\label{eq: AMmcovm}
\Omega_{n}^{*}
\equiv n^{-1}\sum_{s=1, n_s\geq 2}^{S}\left\{\sum_{i=1}^{n_s}(Z_{si}-\bar{Z}_s)(Z_{si}-\bar{Z}_s)'\right\}
\left\{\frac{1}{n_s-1}\sum_{i=1}^{n_s}\left(\varphi_{si}-\bar{\varphi}_s\right)^2\right\},
\end{equation}
where ${\varphi}_{si}\equiv \varphi\left(\frac{i}{n_s+1}\right)$ and $\bar{\varphi}_s\equiv n_{s}^{-1}\sum_{i=1}^{n_s}\varphi_{si}$.
The estimator in \eqref{eq: AMmcovm} is preferred to that in \eqref{eq: AMcovm} because
the former is the variance of $n^{1/2}A_{n}$ (with respect to the distribution of the ranks), whereas the covariance matrix in \eqref{eq: AMcovm} is an approximation valid only when the strata are large, a scenario considered by \cite{Andrews-Marmer(2008)}. In fact, when there are many small strata, the term $\int_{0}^{1}(\varphi(x)-\bar{\varphi})^2dx$
will involve a nonnegligible error as an approximation of $(n_s-1)^{-1}\sum_{i=1}^{n_s}\left(\varphi_{si}-\bar{\varphi}_s\right)^2$. By way of example, let $n_s=2$ for all $s=1,\dots, S$ and $\varphi(x)=\Phi^{-1}(x)$. Simple calculations yield $\bar{\varphi}_s=(\Phi^{-1}(1/3)+\Phi^{-1}(2/3))/2=0$ and
\begin{equation}\label{eq: step}
\frac{1}{n_s-1}\sum_{i=1}^{n_s}\left(\varphi_{si}-\bar{\varphi}_s\right)^2
=(\Phi^{-1}(1/3))^2+(\Phi^{-1}(2/3))^2\approx 0.37.
\end{equation}
This is far below the value $\int_{0}^{1}(\varphi(x)-\bar{\varphi})^2dx=1$, so the covariance matrix estimator in \eqref{eq: AMcovm} is biased upward as an estimator of the asymptotic variance of $n^{1/2}A_{n}$ and the level-$\alpha$ test that rejects when $B_n$ exceeds the $1-\alpha$ quantile of $\chi^2_k$ distribution will be conservative. The LHS of \eqref{eq: step} (after rounding) is 0.56, 0.69, 0.89, 0.96, 0.98 for $n_s=5, 10, 50, 200, 500$, respectively, indicating a reduction in the approximation error with increasing strata sizes. A similar result holds for the Wilcoxon score as well.
The null asymptotic distribution of the statistic $B_{n}^{*}$ is provided in the proposition below.
\begin{proposition}\label{prop: AMstat}
Assume that $\displaystyle \operatornamewithlimits{\lim\inf \ }_{n\to\infty}\lambda_{\min}(\Omega_{n}^{*})>\lambda>0$, and
\begin{enumerate}[label=(\alph*)]
\item\label{AM1} $n^{-1/2}\max_{s,i}\Vert Z_{si}\Vert\to 0$, $n^{-1}\sum_{s=1}^S\sum_{i=1}^{n_s}\Vert Z_{si}\Vert^{2}<C_0$ for all $n$, and
$\int_{0}^{1}\vert {\varphi}(x)\vert^{2+\delta}dx<C_0$
for some positive constants $\delta$ and $C_0$;
\item\label{AM2} $\{r_{Csi}: 1\leq i\leq n_s, 1\leq s\leq S\}$ are independent random variables and for each $s$, $\{r_{Csi}\}_{i=1}^{n_s}$ have an identical continuous distribution.
\end{enumerate}
Then, under $H_0:\beta=\beta_0$, $B_{n}^{*}\displaystyle \stackrel{d}{\longrightarrow} \chi^2_k$ as $(n-S)\to \infty$.
\end{proposition}
\begin{remark}\label{rem: AndrewMarmer}~
\normalfont
The assumptions on the IVs and error terms in \ref{AM1} and \ref{AM2} are comparable to Assumptions C1 and 4 of \cite{Andrews-Marmer(2008)}, and the integrability assumption on the score function
in \ref{AM1} is slightly stronger than that in Assumption 3 of \cite{Andrews-Marmer(2008)} but is not overly restrictive.
\end{remark}
As argued above, the test based on the statistic in \eqref{def: AMRAR}, which employs a $\chi^2_k$ critical value, may be conservative in the presence of many small strata, whereas the test based on the modified statistic in \eqref{def: Bnstar} maintains an asymptotically correct level. This discrepancy is illustrated in a simple numerical experiment
based on the following design:
\begin{align*}
Y_i&=D_i\beta+\gamma_1+X_i\gamma_2+u_i,\\
D_i&=Z_i'\pi+\psi_1+X_i\psi_2+v_i,\quad i=1,\dots, n,
\end{align*}
where we first make $n$ independent draws of $Z_i\sim t_5(0, I_k)$ with $k=1,3$ corresponding to just and over-identified cases.\footnote{Here, $t_5(0, I_k)$ stands for the $k$-variate $t$-distribution with degrees of freedom $5$, zero mean and covariance matrix $I_k$.} For each $k$, we draw the stratification variable $X_i\sim \mathcal{U}(\{1,\dots, \frac{n}{r}\})/( \frac{n}{r})$,
where $r$ varies over $\{2,5,10,25,80\}$. A smaller value of $r$ leads to a greater number of strata, $S$, equal to the empirical support size of $X_i$. We keep $\{(Z_i',X_i)'\}_{i=1}^n$ thus generated fixed over the simulation replications. For each replication, we let $u_i\sim t_1$, $\epsilon_i\sim t_1$, and $v_i=\rho u_i+\sqrt{1-\rho^2}\epsilon_i,$ where $\rho=0.5$ and $u_i$ and $\epsilon_i$ are mutually independent and i.i.d. across $i=1,\dots, n$. The values $\beta,\gamma_1,\gamma_2, \psi_1$ and $\psi_2$ are set to $0$, and $\pi=(1,\dots, 1)'\sqrt{\lambda/(nk)}\in\mathbb{R}^k$ with $\lambda=9$, indicating moderate identification strength of the IVs. The sample size is set relatively large at $n=400$ to elucidate the effect of the strata sizes embodied in $r$.\par
After stratifying the synthetic data according to the realized values of $X_i$, we compute the following four statistics to test the null hypothesis $H_0:\beta=0$: the statistic in \eqref{def: AMRAR} using both the normal and Wilcoxon score functions, denoted as $B_n^N$ and $B_n^W$, respectively, and their modified counterparts $B_n^{N*}$ and $B_n^{W*}$ based on \eqref{def: Bnstar}.\par
Table \ref{tab: size} displays the empirical size of the tests based on 5000 replications. The maximum stratum size $\max n_s$ and the number of strata $S$ are proportional and inversely proportional to $r$ respectively.
The tests $B_n^N$ and $B_n^W$ under-reject when the strata sizes are small, although the latter shows less size distortion. Not surprisingly, the rejection rates of the tests improve as the strata size increases.
In contrast, the modified versions $B_n^{N*}$ and $B_n^{W*}$ exhibit reasonably accurate rejection rates in all cases considered, consistent with the theoretical predictions.\footnote{The result remains qualitatively similar when $Z_i$ and $X_i$ are randomly drawn in each replication, and is available upon request.}
\begin{table}[htbp!]
\begin{center}
\begin{threeparttable}
\caption{Null rejection rates for $H_0:\beta=0$ at $5\%$ level} \label{tab: size}
\begin{tabular}{crrrrrrrrrr}
\toprule
&\multicolumn{5}{c}{$k=1$}&\multicolumn{5}{c}{$k=3$}\\
\cmidrule(lr){2-6} \cmidrule(lr){7-11}
$r$ & 2 & 5 & 10 & 25 & 80 & 2 & 5 & 10 & 25 & 80 \\
\midrule
$B_n^{N*}$ & 4.74 & 4.72 & 5.30 & 4.80 & 4.64 & 4.96 & 4.76 & 4.68 & 5.02 & 4.62 \\
$B_n^{W*}$ & 4.76 & 4.66 & 5.44 & 4.92 & 4.84 & 4.78 & 4.98 & 4.46 & 4.90 & 4.74 \\
$B_n^{N}$& 0.44 & 0.92 & 2.06 & 3.08 & 4.10 & 0.02 & 0.32 & 0.94 & 2.52 & 3.64 \\
$B_n^{W}$ & 2.12 & 2.96 & 4.20 & 4.60 & 4.76 & 1.54 & 2.32 & 3.02 & 4.16 & 4.60 \\ \midrule
$\max n_s$ & 7.00 & 10.00 & 15.00 & 31.00 & 83.00 & 7.00 & 10.00 & 16.00 & 31.00 & 86.00 \\
$S$ & 162.00 & 79.00 & 40.00 & 16.00 & 5.00 & 174.00 & 80.00 & 40.00 & 16.00 & 5.00 \\
\bottomrule
\end{tabular}
\begin{tablenotes}
\footnotesize{\item Notes: $B^N_n$ and $B_n^W$ denote the normal and Wilcoxon score rank statistics, respectively, and $B^{N*}_n$ and $B^{W*}_n$ denote their modified versions. 5000 replications.}
\end{tablenotes}
\end{threeparttable}
\end{center}
\end{table}
\newpage