EconBase
← Back to paper

Partial Identification under Stratified Randomization

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

71,104 characters · 15 sections · 44 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Partial Identification under Stratified Randomization

bibunit[chicago] \setstretch{1} \newsavebox{\tablebox} \newlength{\tableboxwidth} \begin{center} First Draft: \monthyeardate\today; This Draft: \monthyeardate\today \ {Abstract} \end{center} This paper develops a unified framework for partial identification and inference in stratified experiments with attrition, accommodating both equal and heterogeneous treatment shares across strata. For equal-share designs, we apply recent theory for finely stratified experiments to Lee bounds, yielding closed-form, design-consistent variance estimators and properly sized confidence intervals. Simulations show that the conventional formula can overstate uncertainty, while our approach delivers tighter intervals. When treatment shares differ across strata, we propose a new strategy, which combines inverse probability weighting and global trimming to construct valid bounds even when strata are small or unbalanced. We establish identification, introduce a moment estimator, and extend existing inference results to stratified designs with heterogeneous shares, covering a broad class of moment-based estimators which includes the one we formulate. We also generalize our results to designs in which strata are defined solely by observed labels. \ Keywords: Partial Identification, Lee Bounds, Attrition, Stratification, Design-Consistent Inference, Inverse Probability Weighting. \ JEL Codes: C12, C14, C24, C93. \doublespacing

Introduction

Randomized experiments are central to causal inference, but their practical implementation often faces two complications: differential outcome attrition and stratified (blocked) randomization. Attrition can undermine the identification of average treatment effects (ATEs), particularly when missingness is related to treatment status. At the same time, various experiments now employ stratified or blocked assignment, randomizing units within strata to ensure balance and improve efficiency. The combination of these features creates unresolved challenges for valid inference, especially when treatment shares vary across blocks.\footnote{Examples of randomized experiments facing both stratified assignment and differential outcome attrition (or selection into observed outcomes) include the National Job Corps evaluation, where randomization is stratified by center with known, potentially varying assignment probabilities, and wages are observed only for the employed lee2009training; the Colombian job-training program Jóvenes en Acción, where randomization is stratified by training site-by-class and follow-up differs by treatment status for men attanasio2011subsidizing; the Colombian school voucher program (PACES), where lottery assignment occurs within city and cohort-specific lotteries, and ICFES exam-taking differs by voucher status angrist2006long; and Project STAR, where randomization is stratified by school, and ACT/SAT outcomes are observed only for test takers, with class size affecting test taking krueger2001effect. For further discussions of stratified randomization and its prevalence in field experiments, see duflo2007using and bruhn2009pursuit.} In many of these settings, researchers turn to Lee monotone-selection bounds lee2009training to address differential attrition. However, the original Lee bounds framework was developed for settings with independent treatment assignment and does not account for stratified randomization, so additional theory is needed to adapt it to such designs.

To address these issues, we develop a unified framework for partial identification and inference in stratified experiments with outcome attrition, accommodating both equal and heterogeneous treatment shares across blocks. In the equal-share case, we show that Lee’s procedure continues to identify the bounds, but that the conventional standard errors are conservative under this design. To conduct valid inference, we rely on the finely stratified experiment theory of bai2023efficiency and apply it to the context of Lee bounds. This provides a closed-form, design-consistent variance estimator that accurately reflects the negative correlation among assignments induced by equal-share randomization, yielding tighter confidence intervals for Lee bounds. We also extend this approach to encompass designs where blocks are formed either on observed covariates or purely on discrete labels such as schools, clinics, or randomization waves, even when no additional covariate balancing occurs within blocks.

In the heterogeneous-shares case, the standard unconditional Lee bounds are invalid, and the conditional version, based on trimming within each stratum, is often unstable, biased, or undefined in finely stratified experiments with small or sparsely treated blocks. To address these issues, we introduce a novel bounding strategy that uses inverse probability weighting and a single global trimming share to construct bounds that remain valid and robust even with small or unbalanced blocks. We establish identification, propose a moment estimator, and develop an inference framework that generalizes to arbitrary, stratum-specific treatment shares.

While our equal-share results rely on a direct application of the framework in bai2023efficiency, the heterogeneous-shares case requires a separate analysis. Thus, we build on their results to extend inference theory to stratified designs with heterogeneous treatment shares, encompassing a broad class of moment-based estimators that includes the one we introduce. Stratified assignment with varying treatment shares requires additional arguments and leads to a modified variance decomposition. The resulting inference procedure remains tractable even in complex experiments with many small or unbalanced blocks.

This study contributes to several strands of the economics literature, including partial identification with Lee bounds in the presence of attrition and design-consistent inference under stratified randomization. On the partial identification side, early work on nonparametric bounds for randomized experiments with missing data established core tools for analyzing treatment effects under attrition horowitz2000nonparametric,manski2003partial. Recent surveys provide broad overviews of this literature and discuss its implications for empirical work tamer2010partial,kline2023recent. Our contribution within this strand is to focus on a simple and widely used monotone-selection procedure and to show how it can be implemented in stratified randomized experiments with attrition in a way that respects the randomization and delivers design-consistent variance estimation.

Within this toolkit, lee2009training remains especially influential, offering sharp bounds for the average treatment effect among always-observed units under random assignment and monotonicity. Subsequent work extends this approach to settings where selection patterns vary with covariates. semenova2025generalized generalizes Lee’s monotonicity to a covariate-indexed version, and related work develops debiased machine-learning estimators for set- and intersection-bound problems semenova2023debiased,semenova2025aggint. Despite these developments, Lee bounds remain widespread in applied work due to their transparency and practicality, which motivates our study of their implementation in stratified experiments.

In parallel, stratified randomization has become standard in field experiments, improving efficiency with fixed treatment shares within covariate-homogeneous or common-shock blocks bruhn2009pursuit, bai2023efficiency. Nevertheless, most empirical applications with Lee bounds continue to estimate their standard errors as if treatment were assigned completely at random, ignoring the dependence introduced by stratification. As we show through theoretical analysis and simulation evidence, this practice can substantially overstate uncertainty in finely stratified experiments, leading to confidence intervals that are markedly more conservative than intended and highlighting the need for variance estimators that are consistent with the experimental design.

At the same time, the literature on inference for stratified experiments has primarily focused on point-identified average treatment effects bai2022inference, although recent advances, such as the GMM-based framework of bai2023efficiency, provide a sufficiently general theoretical foundation for working with the estimators considered in this paper. Existing approaches have yet to combine monotone-selection bounds with the dependence structure induced by stratification, or to explicitly address partial identification when treatment shares vary across blocks, a case in which standard Lee bounds are no longer valid and Lee bounds that include block fixed effects as covariates may be infeasible in finely stratified designs with small or sparsely treated blocks. These gaps motivate the contributions for inference under stratified assignment and attrition introduced in this paper.

The remainder of the paper is organized as follows. Section (ref) lays out our framework for stratified experimental designs. Section (ref) presents results for designs with homogeneous treatment shares across strata. Section (ref) develops our approach for designs that allow heterogeneous treatment shares. Section (ref) derives inference for these general stratified designs. Section (ref) concludes.

General Setting under Heterogeneous-Shares Stratification

In this section, we consider a stratified design in which treatment shares may vary across blocks, or heterogeneous-shares stratification. The study population consists of $n$ units, indexed by $i = 1, \dots, n$, each associated with observed baseline covariates $X_i \in \mathbb{R}^{d_x}$. Each unit is assigned a binary treatment $D_i \in \{0,1\}$, and we define potential outcomes $(Y_i^*(1), Y_i^*(0))$ and potential selection indicators $(S_i(1), S_i(0)) \in \{0,1\}^2$. The observed outcome and selection indicator are then $S_i = D_i S_i(1) + (1-D_i) S_i(0)$ and $Y_i = S_i\left[D_i Y_i^*(1) + (1-D_i) Y_i^*(0)\right]$, so $Y_i$ is only observed if $S_i = 1$. Let $Q_n$ denote the joint distribution of $\big(Y^{*(n)}(1), Y^{*(n)}(0), S^{(n)}(1),\\ S^{(n)}(0), X^{(n)}\big)$, where the superscript $(n)$ indicates stacking over units $i = 1,\dots,n$. Throughout, we assume i.i.d.\ sampling from a superpopulation, that is, $Q_n = Q^n$, where $Q$ is the marginal distribution of $\big(Y_i^*(1), Y_i^*(0), S_i(1), S_i(0), X_i\big)$.

Heterogeneous-shares stratification refers to partitioning the sample into $G$ blocks, $\Lambda(X^{(n)}) \\= \{\lambda_g : g = 1, \dots, G\}$, with (possibly unequal) sizes $|\lambda_g| = N_g$ so that $\sum_{g=1}^G N_g = n$. Let $b_i \in \{1, \ldots, G\}$ denote the stratum membership of unit $i$, so that $i \in \lambda_{b_i}$. Blocks are constructed so that units within each block are similar in observed covariates, a design known as covariate stratification. Within each block $g$, treatment is assigned so that exactly $T_g \in \{1,\dots, N_g-1\}$ units receive treatment, with block-specific treatment share $\eta_g := T_g / N_g \in (0,1)$. Importantly, we allow $\{\eta_g\}_{g=1}^G$ (equivalently, $\eta_i=\Pr(D_i=1\mid X^{(n)})=\eta_{b_i}$) to depend on the full realized covariate matrix $X^{(n)}$ through the stratification rule $\Lambda(X^{(n)})$, and not only on $X_i$. Formally, \[ (D_i : i \in \lambda_g) \,\Big|\, X^{(n)} \sim \operatorname{Unif} \left\{ \mathbf{d} \in \{0,1\}^{N_g} : \sum_{i \in \lambda_g} d_i = T_g \right\}, \qquad g = 1, \dots, G, \] where assignments are independent across blocks and each block treats exactly a share $\eta_g$ of its units (with $\eta_g$ allowed to vary across $g$). This flexible design allows both block sizes and assignment rates to reflect the logistical, ethical, or other substantive features of real-world experiments. Strata might correspond to geographic regions, clinics, schools, or other covariate-based groups, with treatment shares $\eta_g$ chosen to meet power or policy targets.

This blockwise randomization ensures that, conditional on the full set of observed covariates, the collection of potential outcomes and selection indicators $\{Y_i^*(1), Y_i^*(0), S_i(1), S_i(0)\}_{i=1}^n$ is jointly independent of the treatment vector $(D_1, \dots, D_n)$. Each unit in a given block $g$ has assignment probability $\eta_g$. Given this assignment mechanism and the conditions described above, the marginal distribution of the observed data $(X_i, D_i, Y_i, S_i)$ is identical across $i = 1, \dots, n$, and does not depend on $n$. However, due to the fixed numbers of treated units within blocks, the observed data are not generally independent across units. As a result, the collection $\{(X_i, D_i, Y_i, S_i)\}_{i=1}^n$ forms an identically distributed, but dependent, sample from a common distribution. When the $\eta_g$ differ across blocks, unconditional independence between $D_i$ and potential outcomes need not hold at the sample level because treated and control groups weight blocks differently.

Importantly, fixing $T_g$ within each block introduces negative correlation among assignments, which has substantial consequences for inference. Accounting for the exact assignment mechanism is therefore critical for valid inference, as it ensures that variance estimators properly reflect this dependence structure and fully realize the efficiency gains available under design-consistent variance estimation.

To establish the large-sample results used later for inference with heterogeneous treatment shares, we state a set of design and regularity conditions under which our Lee-type estimators satisfy a central limit theorem we will propose later. The conditions are analogous to those in bai2023efficiency but extended to allow block-specific treatment shares, specifically with a change in Assumption (ref). They pertain only to the design and regularity required for inference, and identification assumptions are presented in Section (ref).

assumption[i.i.d.\ sampling] Let $Q$ be the marginal distribution of $(Y_i^*(1),Y_i^*(0),S_i(1),\\S_i(0),X_i)$ and let $Q_n$ be the joint distribution of $\big(Y^{*(n)}(1),Y^{*(n)}(0),S^{(n)}(1),S^{(n)}(0),X^{(n)}\big)$. We assume i.i.d.\ sampling from a superpopulation, i.e.\ $Q_n = Q^n$.
assumption[Covariate balancing] As $n\to\infty$ with $G\to\infty$, block sizes are uniformly bounded, i.e.\ there exists a constant $\bar N<\infty$ such that \[ \Pr\!\left(\max_{1\le g\le G} N_g \le \bar N\right)\ \longrightarrow\ 1. \] Moreover, \[ \frac{1}{n}\sum_{g=1}^{G}\ \max_{i,j\in\lambda_g}\|X_i-X_j\|^2 \xrightarrow{p} 0, \] so that blocks become asymptotically homogeneous in covariates.
assumption[Stratification] Given the full covariate matrix $X^{(n)}=(X_1,\dots,X_n)$ (equivalently, conditioning on stratum membership $b_i$ and covariates $X_i$), the collection $\{Y_i^*(1),Y_i^*(0),\\S_i(1),S_i(0)\}_{i=1}^n$ is jointly independent of the treatment vector $(D_1,\dots,D_n)$. Furthermore, conditional on $X^{(n)}$, the blockwise assignment vectors $\{(D_i: i\in\lambda_g)\}_{g=1}^G$ are independent across blocks, and within each block $\lambda_g$ treatment is assigned uniformly over all allocations with exactly $T_g$ treated units: \[ (D_i : i \in \lambda_g)\,\big|\,X^{(n)} \sim \mathrm{Unif}\!\left\{\mathbf d\in\{0,1\}^{N_g} : \sum_{i\in\lambda_g} d_i=T_g\right\},\qquad g=1,\dots,G, \] with possibly heterogeneous shares $\eta_g:=T_g/N_g\in(0,1)$. Here $\eta_g$ (equivalently, $\eta_i=\Pr(D_i=1\mid X^{(n)})=\eta_{b_i}$) is allowed to depend on the full covariate matrix $X^{(n)}$ through the design map $X^{(n)}\mapsto \big(\Lambda(X^{(n)}),\{T_g\}_{g=1}^G\big)$, rather than only on $X_i$.
assumption[Moment regularity] Let $m(\cdot)=(m_s(\cdot):1\le s\le d_\theta)'$. The moment functions are such that \begin{enumerate}[label=(\alph*)] • For every $\epsilon>0$, \[ \inf_{\theta\in\Theta:\,\|\theta-\theta_0\|>\epsilon}\ \big\|\operatorname{\mathbb{E}}\big[m(X_i,D_i,Y_i;\theta)\big]\big\|\;>\;0. \]$\operatorname{\mathbb{E}}[m(X_i,D_i,Y_i;\theta)]$ is differentiable at $\theta_0$ with a nonsingular derivative \[ M \;=\; \left.\frac{\partial}{\partial\theta'}\,\operatorname{\mathbb{E}}\!\left[m(X,D,Y;\theta)\right]\right|_{\theta=\theta_0}\, . \] • For $1\le s\le d_\theta$ and $d\in\{0,1\}$, \[ \operatorname{\mathbb{E}}\!\left[\big(m_s(X,d,Y(d),\theta)-m_s(X,d,Y(d),\theta_0)\big)^2\right]\ \to\ 0 \quad\text{as }\theta\to\theta_0. \] • For $1\le s\le d_\theta$, $\{m_s(x,d,y,\theta):\theta\in\Theta\}$ is pointwise measurable in the sense that there exists a countable set $\Theta^\ast$ such that for each $\theta\in\Theta$, there exists a sequence $\{\theta_m\}\subset\Theta^\ast$ such that $m_s(x,d,y,\theta_m)\to m_s(x,d,y,\theta)$ as $m\to\infty$ for all $x,d,y$. • (i) $\displaystyle \sup_{\theta\in\Theta}\ \operatorname{\mathbb{E}}\!\left[\|m(X,d,Y(d),\theta)\|\right]<\infty$ for $d\in\{0,1\}$. (ii) $\{m(x,1,y,\theta):\theta\in\Theta^\ast\}$ and $\{m(x,0,y,\theta):\theta\in\Theta^\ast\}$ are $Q$-Donsker. • For $d\in\{0,1\}$, let $\mu_d(x,\theta):=\operatorname{\mathbb{E}}[m(X,d,Y(d),\theta)\mid X=x]$. There exists $L_d<\infty$ such that for all $x,x'\in\mathbb R^{d_x}$, \[ \sup_{\theta\in\Theta^\ast}\,|\mu_d(x,\theta)-\mu_d(x',\theta)| \ \le\ L_d\|x-x'\|. \] \end{enumerate}
remark[Label-based stratification] The setting in this section does not directly extend to designs with many small blocks defined purely by discrete labels, such as schools, clinics, or randomization waves. Appendix (ref) treats this case separately, using a label-based triangular-array framework appropriate for experiments with many small strata. The results are somewhat similar, despite different proofs.

Inference under Equal-Share Stratification

Setup: Equal-Share Stratification

This section specializes the design to a common treatment share across strata. Let $\Lambda(X^{(n)})=\{\lambda_g: g=1,\dots,G\}$ denote the observable blocks with equal sizes $N_g=k$ and labels $b_i\in\{1,\dots,G\}$. Within block $g$, exactly $\ell$ units are assigned to treatment without replacement, with a common share $\eta:=\ell/k\in(0,1)$ for all $g$. To formally characterize the assignment mechanism in this equal-share stratification setting, we state the following assumption: \newcounter{saveassumpES} \setcounter{saveassumpES}{\value{assumption}}

assumption[Equal-share stratification] Given the full covariate matrix $X^{(n)}=(X_1,\\\dots,X_n)$ (equivalently, conditioning on stratum membership $b_i$ and covariates $X_i$), the collection $\{Y_i^*(1),Y_i^*(0),S_i(1),S_i(0)\}_{i=1}^n$ is jointly independent of the treatment vector $(D_1,\dots,D_n)$. Furthermore, conditional on $X^{(n)}$, the blockwise assignment vectors $\{(D_i: i\in\lambda_g)\}_{g=1}^G$ are independent across blocks, and within each block $\lambda_g$ treatment is assigned uniformly over all allocations with exactly $\ell$ treated units: \[ (D_i: i\in\lambda_g)\,\big|\,X^{(n)} \sim \operatorname{Unif}\!\left\{\mathbf d\in\{0,1\}^{k}:\ \sum_{i\in\lambda_g} d_i=\ell\right\},\qquad g=1,\dots,G, \] with $\eta:=\ell/k\in(0,1)$ common to all $g$.

\setcounter{assumption}{\value{saveassumpES}}

Partial Identification and Consistency of Lee Bounds under Equal-Share Stratification

We now discuss the conditions used for the partial identification of the always-observed effect in stratified designs. In lee2009training, the unconditional Lee bounds are derived under an independence condition in which treatment is unconditionally independent of the vector of potential outcomes and selection indicators. In our setting, Lemma (ref) shows that this unconditional independence is implied by the equal-share stratified randomization design. In addition, we impose the monotonicity in selection condition from the original Lee framework, which does not follow from the experimental design and must be justified on empirical grounds.

assumption[Monotone selection] \[ \Pr\!\big[S_i(1)\ge S_i(0)\big]=1. \] Equivalently, one may assume the reverse monotonicity $\Pr[S_i(1) \leq S_i(0)] = 1$; all results then follow by swapping treated and control in the trimming step. We maintain $\Pr[S_i(1) \geq S_i(0)] = 1$ for exposition.

Under equal-share stratified randomization, treated and control units draw the same mixture over strata, so treatment is unconditionally independent of the vector of potential outcomes and selection indicators. The next result formalizes this implication. At the same time, fixing $T_g$ within each block induces negative within-block correlation in assignments. Thus, observations are identically distributed but dependent, and inference must account for this design-induced dependence.

lemma[Equal-share stratification implies unconditional independence] Suppose Assumption (ref) holds. Then \[ (Y_i^*(1),Y_i^*(0),S_i(1),S_i(0))\ \perp\!\!\!\perp\ D_i. \]
proofBlock randomization implies stratum-level independence, \( (Y_i^*(1),Y_i^*(0),S_i(1),S_i(0))\ \perp\!\!\!\perp\ D_i \mid b_i, \) where $b_i$ denotes stratum membership. Equal shares give $\Pr(D_i=1\mid b_i=g)=\eta$ for all $g$, hence $D_i\perp b_i$. Let $W_i:=(Y_i^*(1),Y_i^*(0),S_i(1),S_i(0))$. Then \[ \Pr(D_i=1\mid W_i) = \sum_{g=1}^G \Pr(D_i=1\mid b_i=g)\,\Pr(b_i=g\mid W_i) = \eta \sum_{g=1}^G \Pr(b_i=g\mid W_i) = \eta, \] which is constant in $W_i$. Therefore $(Y_i^*(1),Y_i^*(0),S_i(1),S_i(0))\ \perp\!\!\!\perp\ D_i$.

In the blocked design studied here, each stratum uses a common treatment share. By Lemma (ref), block randomization delivers unconditional independence at the sample level, while selection monotonicity (Assumption (ref)) remains a substantive identifying restriction. Under these conditions, the unconditional Lee bounds interval continues to characterize the identified set for the ATE among always-observed units, and the sharpness argument is unchanged. The next result states partial identification and sharpness in this setting. We then turn to consistency.

proposition[Partial ID and sharpness under equal-share stratification] Suppose Assumption (ref) holds, and Assumptions (ref), (ref), (ref), and (ref) also hold, which specify an equal-share stratification design. Let \[ q \;=\; 1 - \frac{\Pr(S=1\mid D=0)}{\Pr(S=1\mid D=1)} \in [0,1], \] let $y_\alpha$ denote the $\alpha$-quantile of $Y$ among $\{D=1,S=1\}$, and define \begin{align*} \mu_0 &= \operatorname{\mathbb{E}}\big[\,Y \mid D=0, S=1\,\big],\\ \mu_1^{LB} &= \operatorname{\mathbb{E}}\big[\,Y \mid D=1, S=1,\, Y \le y_{1-q}\,\big],\\ \mu_1^{UB} &= \operatorname{\mathbb{E}}\big[\,Y \mid D=1, S=1,\, Y \ge y_{q}\,\big]. \end{align*} Then the average treatment effect among always-observed units, \[ \Delta_{AO} \;=\; \operatorname{\mathbb{E}}\!\left[\,Y^*(1)-Y^*(0)\mid S(1)=S(0)=1\,\right], \] is partially identified by the Lee bounds interval \[ \Delta_{AO} \in \big[\,\mu_1^{LB}-\mu_0,\; \mu_1^{UB}-\mu_0\,\big], \] and these bounds are sharp. Hence, with equal-share stratification, the identified set coincides with the i.i.d.\ case, and Lee bounds remain sharp.
proofThe argument follows lee2009training, which relies on unconditional independence and monotonicity in selection. Under equal-share stratification, Lemma (ref) shows that Assumptions (ref), (ref), (ref), and (ref) imply unconditional independence. Together with Assumption (ref), this verifies all conditions used in lee2009training, so his proof applies verbatim, including the least-favorable construction that yields sharpness.

With identification in place, the only departure from the i.i.d.\ case concerns the uniform convergence step used in the consistency proof. The i.i.d.\ ULLN invoked by newey1994large, cited by lee2009training, is replaced by a version valid for block randomization with a common share, built from finite-population limit theory and uniform convergence results. The proposition below records the resulting consistency claim, and the proof supplies only this substitution.

proposition[Consistency under equal-share stratification] Suppose Assumption (ref) holds, and Assumptions (ref), (ref), (ref) and (ref) also hold, which specify an equal-share stratification design with a common treatment share and independent blocks. Let $\hat q = 1 - \widehat{\Pr}(S=1\mid D=0)/\widehat{\Pr}(S=1\mid D=1)$, let $\hat y_\alpha$ denote the empirical $\alpha$-quantile of $Y$ among $\{D=1,S=1\}$, and define the sample analogs \begin{align*} \hat\mu_0 &= \operatorname{\mathbb{E}}_n[\,Y \mid D=0,S=1\,],\\ \hat\mu_1^{LB} &= \operatorname{\mathbb{E}}_n[\,Y \mid D=1,S=1,\, Y \le \hat y_{1-\hat q}\,],\\ \hat\mu_1^{UB} &= \operatorname{\mathbb{E}}_n[\,Y \mid D=1,S=1,\, Y \ge \hat y_{\hat q}\,]. \end{align*} Then \[ \big(\hat\mu_1^{LB} - \hat\mu_0,\; \hat\mu_1^{UB} - \hat\mu_0\big) \;\xrightarrow{p}\; \big(\mu_1^{LB} - \mu_0,\; \mu_1^{UB} - \mu_0\big). \]
proofThe argument adapts lee2009training to equal-share stratification by replacing the i.i.d.\ ULLN with a version valid for block randomization. Under Assumptions (ref), (ref), (ref) and (ref), Lemma (ref) shows that unconditional independence holds, so the conditions used in lee2009training are satisfied under Assumption (ref). Blockwise LLNs follow from hajek1960limiting, and uniform convergence results (including triangular arrays) follow from andrews1992generic. With this substitution, newey1994large yields consistency exactly as in lee2009training. See Appendix (ref) for a detailed proof.

Hence, partial identification and sharpness carry over to the equal-share stratified design, and consistency follows once the uniform-convergence step is adapted to blocked assignment. What changes is sampling behavior, because assignment without replacement creates within-block dependence, so inference must use variance formulas that respect this feature.

Asymptotic Distribution and Variance for Lee Bounds under Equal-Share Stratification

Understanding the asymptotic behavior of Lee bounds under equal-share stratification is crucial for valid and efficient inference in finely stratified experimental designs. This setting is characterized by partitioning the population into blocks with a common treatment share, inducing within-block dependence that alters the variance relative to i.i.d.\ sampling. To rigorously justify inference, we work under Assumption (ref) together with the design and regularity conditions in Assumptions (ref), (ref) and (ref). These are the same as in bai2023efficiency, so we first represent the Lee lower bound as a just-identified GMM estimator and then apply their central limit theorem and variance results to this moment representation.

Let $Z_i = (Y_i,S_i,D_i,X_i)$ and define $L_i = \mathbf{1}\{D_i=1,\,S_i=1,\,Y_i \le y_{1-p}\}$ and $U_i = \mathbf{1}\{D_i=1,\,S_i=1,\,Y_i > y_{1-p}\}$, where $y_{1-p}$ is the $(1-p)$ quantile of $Y_i$ among $\{D_i=1,S_i=1\}$. Parameterize $\theta = (\Delta_{AO}^{LB},\mu_0,p,\alpha)^\prime$, where $\Delta_{AO}^{LB}$ is the Lee lower bound, $\mu_0 = \operatorname{\mathbb{E}}[Y_i \mid D_i=0,S_i=1]$ is the control mean, $p$ is the trimming share, and $\alpha = \Pr(S_i=1\mid D_i=0)$ is the control-group selection rate, so that the trimmed treated mean is $\mu_1^{LB} = \mu_0 + \Delta_{AO}^{LB}$. The corresponding moment vector is \[ m(Z_i,\theta) =

pmatrix[pmatrix omitted — 248 chars of source]

, \qquad \operatorname{\mathbb{E}}[m(Z_i,\theta_0)] = 0. \] In this parametrization, the Lee lower bound is the first component of $\theta$, and an analogous moment system characterizes the upper bound by using indicators that trim the lower tail of the treated outcome distribution instead of the upper tail.

proposition[Asymptotics of Lee bounds under equal-share stratification] Sup\-pose Assumptions (ref), (ref), (ref) and (ref) hold. Then, the Lee bounds estimator obeys the following large-sample distribution: \[ \sqrt{n}\,(\hat\theta_n - \theta_0) \;\xrightarrow{d}\; \mathcal{N}(0, V_\eta), \] where $V_\eta = M^{-1}\Omega_\eta M^{-T}$, with $M = \partial_\theta \operatorname{\mathbb{E}}[m(Z_i, \theta)]\big|_{\theta = \theta_0}$, and the variance $\Omega_\eta$ decomposes as \[ \Omega_\eta = \operatorname{\mathbb{E}}\left[ \eta\,\operatorname{Var}(m_1 \mid X) + (1-\eta)\,\operatorname{Var}(m_0 \mid X) \right] + \operatorname{Var}\left[ \eta\,\mu_1(X) + (1-\eta)\,\mu_0(X) \right], \] where $m_d = m(Z_i, \theta_0) \mathbf{1}\{D_i = d\}$ and $\mu_d(X) = \operatorname{\mathbb{E}}[m_d \mid X]$ for $d \in \{0,1\}$. In particular, the asymptotic variance of the Lee lower bound corresponds to the $(1,1)$ element of $V_\eta$.
proofThe proof is immediate from Theorem 3.1 of bai2023efficiency, since the assumptions made are equivalent to Assumptions 3.1-3.3 (plus i.i.d.\ sampling) in that paper. Covariate balance ensures that block-level imbalances vanish at the $\sqrt{n}$ rate, allowing the asymptotic variance to be accurately captured by $V_\eta$. Furthermore, the Lee bounds estimator meets the required moment regularity conditions because its moments are defined only by linear and indicator functions.

An implementation caveat for variance estimation concerns the moment vector. In lee2009training, under i.i.d.\ assignment, the moment vector used in the sandwich variance targets only the trimmed treated mean, and the control mean is incorporated only when assembling the final variance of the estimator. Under equal-share stratification, the trimmed treated mean and the control mean are correlated by within-block assignment. Omitting the control mean from the moment vector ignores this covariance and can bias the estimated standard errors. Accordingly, we specify the first moment condition as the difference between the trimmed treated mean and the control mean and then estimate the variance with the design-consistent formula of bai2023efficiency.

Given this moment specification, the variance decomposition presented in bai2023efficiency applies directly to the Lee bounds estimator and enables consistent estimation even with small block sizes. Notably, omitting the block correction term, as in conventional i.i.d.\ variance estimation, leads to conservative standard errors. These observations reinforce an important theoretical result: the variance under equal-share stratification is always less than or equal to that under i.i.d.\ assignment bai2023efficiency. Incorporating the blocked assignment structure into variance estimation is thus essential, not only to avoid overstating uncertainty but also to fully realize the efficiency gains in empirical applications.

remark[Equal shares and label-based designs] With equal shares, unconditional independence holds by design and standard Lee bounds apply in both covariate-based and label-based stratified designs. The main difference lies in how precision can be improved. When strata are formed by grouping units with similar covariates and within-stratum diameters shrink as $n$ grows, it is possible to borrow information from nearby strata. In label-based designs with many small blocks, there are no meaningful neighbors to use for pooling. Appendix (ref) considers this setting and requires at least two treated and two control observations in each stratum.

Simulation Results

We now present simulation evidence with two purposes. First, we assess the performance of the design-consistent variance estimator for Lee bounds under equal-share stratification. Second, we study how conditional Lee bounds, obtained by trimming within each stratum and then aggregating the resulting intervals, perform in stratified designs with small strata. The estimation problems that arise in this setting motivate a pooled alternative that draws on information across strata.

In an equal-share matched-pair design, the design-consistent variance estimator based on bai2023efficiency delivers standard errors that closely match the Monte Carlo standard deviation of the Lee bounds estimator. By contrast, the conventional i.i.d.\ variance estimator substantially overstates uncertainty. Its average standard error is $0.0569$ against an empirical standard deviation of $0.0397$, inflating coverage to about $99.5\%$ instead of the nominal $95\%$. This confirms that incorporating the blocked assignment structure is essential for accurate inference under equal-share stratification.

A second simulation considers stratified designs with many small strata. A natural alternative in this setting is to apply Lee bounds conditional on strata and then aggregate the resulting intervals. In our designs, the difficulty lies in estimation rather than identification. With only a few treated units and controls in each stratum, the cell-level trimming share and the associated trimmed means are based on very little information and can be sensitive to single observations, including in sparse or nearly empty cells. This produces conditional Lee bounds that are noisy and sometimes poorly centered, even when the assumptions behind Lee bounds hold. Appendix (ref) reports the simulation designs, numerical results and a broader discussion of additional limitations of conditional Lee bounds in sparse stratified designs.

Lee–IPW Bounds under Heterogeneous-Shares Stratification

Partial Identification Strategy with Lee-IPW Bounds

Many field trials and randomized experiments assign different fractions of units to treatment across strata. Examples include oversampling in small sites to ensure power, demographic quotas that create unequal assignment rates, and logistical constraints that yield classes, villages, or clinics of different sizes and treatment proportions. Under stratified designs with heterogeneous treatment shares, the design yields only stratum-level (conditional) independence, so unconditional independence between treatment and potential outcomes generally fails. Consequently, conventional unconditional Lee bounds are invalid. The following lemma formalizes this point.

lemma[Heterogeneous-shares stratification implies stratum-level independence] Suppose Assumption (ref) holds. Then \[ (Y_i^*(1),Y_i^*(0),S_i(1),S_i(0))\ \perp\!\!\!\perp\ D_i \ \big|\ b_i . \] Moreover, writing $W_i:=(Y_i^*(1),Y_i^*(0),S_i(1),S_i(0))$, \[ \Pr(D_i=1\mid W_i)=\sum_{g=1}^G \eta_g\,\Pr(b_i=g\mid W_i), \] so unconditional independence $(W_i\perp\!\!\!\perp D_i)$ generally fails unless the right-hand side is constant in $W_i$ (e.g., equal shares or stratum composition invariant in $W_i$). In heterogeneous-shares settings, this stratum-level independence is therefore the appropriate independence condition.
proofUniform assignment within each $\lambda_g$ implies \( (Y_i^*(1),Y_i^*(0),S_i(1),S_i(0))\ \perp\!\!\!\perp\ D_i \mid b_i, \) establishing stratum-level independence. For the marginal statement, \[ \Pr(D_i=1\mid W_i) =\sum_{g=1}^G \Pr(D_i=1\mid b_i=g)\,\Pr(b_i=g\mid W_i) =\sum_{g=1}^G \eta_g\,\Pr(b_i=g\mid W_i), \] which is constant in $W_i$ only in the noted special cases. Otherwise, unconditional independence fails.

Building on Lemma (ref), the partial identification of treatment effects with heterogeneous treatment shares must work with this stratum-level condition that reflects the block randomization design. The Lee-IPW estimator adopts the stratum-level independence implied by Assumption (ref), together with Assumption (ref), allowing for flexible assignment probabilities while restoring valid partial identification.

Our target estimand is the average treatment effect among always-observed units, $ \Delta_{AO} = \mathbb{E}\left[ Y^*(1) - Y^*(0) \,\big|\, S(1) = S(0) = 1 \right]. $ The Lee–IPW method provides a principled strategy for handling stratified experiments with heterogeneous treatment shares. By employing inverse probability weighting (IPW), it re-balances the data to account for varying assignment proportions across strata and then applies a single trimming share to the weighted sample. This procedure yields more stable and reliable bounds, particularly in studies with many small or sparsely treated strata, where traditional approaches can be biased or infeasible.

In this framework, let $p = \Pr(D = 1)$ denote the overall treated share, and let $w_{c,g} = (1-p)/(1-\eta_g)$ denote the weight applied to control units in stratum $g$. After weighting by $w_{c,g}$, the distribution of control observations matches the distribution of controls in the overall sample. This adjustment corrects for the over- or under-representation of controls from strata with heterogeneous treatment shares. The always-observed control mean is then identified as \[ \mu_{0}=E[Y^{*}(0)\mid AO \text{ (Always-observed)}]=\frac{E[(1-D)S\,w_{c,g}Y]}{E[(1-D)S\,w_{c,g}]}. \]

Similarly, for the treated–observed units, we define $\delta = \Pr(D_{g,i} = 1 \mid AO)$ as the probability that a unit is treated, conditional on being always-observed, and reweight each treated–observed outcome by $h_{g,i} = \delta / \eta_g$, so that $\widetilde{Y}_{g,i} = h_{g,i} Y_{g,i}$. This ensures that the mean of the reweighted treated outcomes corresponds to the mean for always-observed treated potential outcomes: $\mathbb{E}[\widetilde{Y} \mid D=1, AO] = \mathbb{E}[Y^*(1) \mid AO]$. It is the same principle that motivated $w_{c,g}$, now applied to the treated arm with a focus on always-observed units. The weight can be written as a functional of the data and design (see Appendix (ref) for details), thus it is identified.

To identify the appropriate trimming share, monotonicity implies that the excess response rate in the treated group, after control reweighting to mirror the treated units distribution across strata, corresponds to the proportion $q$ of “induced” outcomes to be trimmed. With $w_{q,g} = \frac{\eta_g (1-p)}{(1-\eta_g)p}$ denoting the weight applied to control units in stratum $g$ when constructing $q$, the trimming share is given by \[ q = \frac{\Pr[S=1 \mid D=1] - \mathbb{E}[S\,w_{q,g} \mid D=0]}{\Pr[S=1 \mid D=1]}. \]

Finally, the bounds for the always-observed average treatment effect are constructed by trimming the lower (or upper) $q$ fraction from the distribution of reweighted treated–observed outcomes. Let $\tilde{y}_\alpha$ be the $\alpha$-quantile of $\widetilde{Y}$. The bounds are given by \[ \mu_1^{LB} = \mathbb{E}[\widetilde{Y} \mid D=1, S=1, \widetilde{Y} \le \tilde{y}_{1-q}], \quad \mu_1^{UB} = \mathbb{E}[\widetilde{Y} \mid D=1, S=1, \widetilde{Y} \ge \tilde{y}_{q}], \] and thus, \[ \Delta_{AO} \in [\mu_1^{LB} - \mu_0,\, \mu_1^{UB} - \mu_0]. \]

Considering the preceding definitions and construction, the identified set for the always-observed treatment effect is characterized by the following result.

proposition[Partial Identification with Lee-IPW Bounds] Under the stratum-level independence condition implied by the stratified assignment (Lemma (ref)) and Assumption (ref) (monotonicity), the average treatment effect among always-observed units, \[ \Delta_{AO} = \mathbb{E}\left[ Y^*(1) - Y^*(0) \mid S(1) = S(0) = 1 \right], \] is partially identified by the interval \[ [\mu_1^{LB} - \mu_0,\ \mu_1^{UB} - \mu_0], \] where \begin{align*} \mu_0 &= \frac{\mathbb{E}[(1-D)S w_{c,g} Y]}{\mathbb{E}[(1-D)S w_{c,g}]}, \\ \mu_1^{LB} &= \mathbb{E}[\widetilde{Y} \mid D=1, S=1,\, \widetilde{Y} \leq \tilde{y}_{1-q}], \\ \mu_1^{UB} &= \mathbb{E}[\widetilde{Y} \mid D=1, S=1,\, \widetilde{Y} \geq \tilde{y}_q], \\ q &= \frac{\Pr[S=1 \mid D=1] - \mathbb{E}[S w_{q,g} \mid D=0]}{\Pr[S=1 \mid D=1]}. \end{align*} All weights and quantiles are as defined in the prior discussion.
proofSee Appendix (ref) for a detailed proof and derivation.

Conditional Lee bounds are sharp under stratum-level randomization and monotonicity, as they fully exploit differential selection within each block, but their reliability breaks down when strata are small, sparsely treated, or contain outliers, often resulting in bias, high variance, or undefined bounds. The Lee–IPW procedure, while not sharp because it does not use all within-block selection information and may yield somewhat wider intervals, addresses these limitations by pooling information and applying a unified trimming rule. This approach provides stable and feasible inference without ad hoc adjustments for ill-defined trimming shares, offering robust partial identification even in finely stratified or sparsely populated designs. In practice, Lee–IPW intervals are especially valuable in applied settings where strata are small or imbalanced, and offer a reliable tool for empirical analysis in challenging experimental designs.

Lee-IPW Estimator and Moment Conditions

The Lee–IPW estimator translates the identification strategy from the previous section into a practical procedure for bounding the average treatment effect among always-observed units. All calculations rely solely on observed data and the known randomization shares for each stratum, and the estimator is consistent even with small or highly unbalanced blocks.

The estimation proceeds by first computing the realized treatment rate $\hat{p} = n^{-1} \sum_{i=1}^n D_i$ and the design share $\eta_g = T_g / N_g$ within each stratum. For each unit, block weights are assigned to controls and for trimming purposes as $\hat{w}_{c,g} = (1-\hat{p})/(1-\eta_g)$ and $\hat{w}_{q,g} = \eta_g (1-\hat{p}) / [(1-\eta_g)\hat{p}]$. In the following, all summations are over the entire sample.

To estimate the trimming share, which is fundamental for partial identification, the procedure calculates \[ \hat{q} = 1 - \frac{\hat{p} \sum_{i} (1-D_i) S_i \hat{w}_{q,b_i}}{(1-\hat{p}) \sum_{i} D_i S_i}, \qquad \hat{q} \in [0, 1]. \] This value represents the proportion of treated–observed outcomes to trim, corresponding to the excess response rate attributable to treatment, after adjusting for heterogeneous assignment.

Next, the always-observed treatment probability is estimated by first computing, within each stratum, the proportion of controls observed: $\hat{m}_g = \sum_{i\in\lambda_g} (1-D_i) S_i / \sum_{i\in\lambda_g} (1-D_i)$. The overall probability is then \[ \hat{\delta} = \frac{\sum_i D_i \hat{m}_{b_i}}{\sum_i \hat{m}_{b_i}}. \]

Each treated, observed outcome is then reweighted as $\widetilde{Y}_i = (\hat{\delta} / \eta_{b_i}) Y_i$, aligning the mean of the reweighted treated sample with that for always-observed units. The empirical quantiles $\hat{\tilde{y}}_{1-\hat{q}}$ and $\hat{\tilde{y}}_{\hat{q}}$ among these values set the cutoffs for trimming. Then, the lower and upper bounds for the treated mean are \[ \hat{\mu}_1^{LB} = \frac{\sum_{i} D_i S_i\, \mathbbm{1}\{\widetilde{Y}_i \leq \hat{\tilde{y}}_{1-\hat{q}}\} \, \widetilde{Y}_i}{\sum_{i} D_i S_i\, \mathbbm{1}\{\widetilde{Y}_i \leq \hat{\tilde{y}}_{1-\hat{q}}\}}, \qquad \hat{\mu}_1^{UB} = \frac{\sum_{i} D_i S_i\, \mathbbm{1}\{\widetilde{Y}_i \geq \hat{\tilde{y}}_{\hat{q}}\} \, \widetilde{Y}_i}{\sum_{i} D_i S_i\, \mathbbm{1}\{\widetilde{Y}_i \geq \hat{\tilde{y}}_{\hat{q}}\}}, \] while the control mean is estimated by \[ \hat{\mu}_0 = \frac{\sum_{i} (1-D_i) S_i\, \hat{w}_{c,b_i}\, Y_i}{\sum_{i} (1-D_i) S_i\, \hat{w}_{c,b_i}}. \] The Lee–IPW bounds for the always-observed average treatment effect are hence given by \[ \hat{\Delta}^{LB} = \hat{\mu}_1^{LB} - \hat{\mu}_0, \qquad \hat{\Delta}^{UB} = \hat{\mu}_1^{UB} - \hat{\mu}_0. \]

To place the estimator within a more formal framework and to connect it with the inference procedures discussed in the next section, it is necessary to express the Lee–IPW bounds estimator as the solution to a system of moment equations. The characterization is summarized below.

proposition[Moment Characterization of the Lee-IPW Lower Bound] Let $\theta = (\mu_{1}^{LB}, \mu_{0}, \tilde{y}_{1-q}, \delta, q)^{\top} \in \Theta \subset \mathbb{R}^5$ collect the parameters for the lower Lee-IPW bound, and let $Z_i = (Y_i, S_i, D_i, b_i)$ collect the observed outcome, selection indicator, treatment assignment, and stratum membership for unit $i$. Define the moment vector $m(Z_i, \theta) = (m_{1i}(\theta), m_{2i}(\theta), m_{3i}(\theta),\\ m_{4i}(\theta), m_{5i}(\theta))^\top$ by \begin{align*} m_{1i}(\theta) &= (\widetilde{Y}_i - \mu_{1}^{LB})\, D_i S_i\, \mathbbm{1}\{\widetilde{Y}_i \le \tilde{y}_{1-q}\}, \\ m_{2i}(\theta) &= (Y_i - \mu_0)\, (1-D_i) S_i\, w_{c,b_i}, \\ m_{3i}(\theta) &= [\mathbbm{1}\{\widetilde{Y}_i > \tilde{y}_{1-q}\} - q]\, D_i S_i, \\ m_{4i}(\theta) &= r_{b_i} (D_i - \delta), \\ m_{5i}(\theta) &= \frac{1-q}{p} D_i S_i - \frac{1}{1-p} (1-D_i) S_i\, w_{q,b_i}, \end{align*} where $p = \Pr(D = 1)$ and \[ r_{b_i} := \frac{\sum_{j \in \lambda_{b_i}} (1-D_j) S_j}{\sum_{j \in \lambda_{b_i}} (1-D_j)} \] denotes the observed selection rate among controls in the stratum containing unit $i$. Let $\Delta^{LB}$ denote the lower Lee-IPW bound constructed in Section (ref). Under stratified randomization in Assumption (ref), so that Lemma (ref) holds, and monotone selection in Assumption (ref), there exists $\theta_0 = (\mu_{1}^{LB}, \mu_{0}, \tilde{y}_{1-q}, \delta, q)^{\top} \in \Theta$ such that \[ \mathbb{E}[m(Z_i, \theta_0)] = 0 \qquad\text{and}\qquad \Delta^{LB} = \mu_{1}^{LB} - \mu_{0}. \] Moreover, any $\theta \in \Theta$ satisfying $\mathbb{E}[m(Z_i, \theta)] = 0$ has its first two components equal to the trimmed treated mean and control mean that define $\Delta^{LB}$ and its fifth component equal to the trimming share $q$. Hence the system $\mathbb{E}[m(Z_i, \theta)] = 0$ is equivalent to the Lee-IPW lower-bound construction.

The lower Lee–IPW bound is constructed using the system above; the upper bound replaces the indicator $\mathbbm{1}\{\widetilde{Y}_i \le y_{1-q}\}$ with $\mathbbm{1}\{\widetilde{Y}_i \ge y_{q}\}$ in the relevant moments. The just-identified GMM estimator $\hat{\theta}_n$ solves $\frac{1}{n} \sum_{i=1}^n m(Z_i, \hat{\theta}_n) = 0$, with Jacobian matrix $\hat{M}_n = \frac{1}{n} \sum_{i=1}^n \partial_{\theta} m(Z_i, \hat{\theta}_n)$. The estimator for the lower bound is then given by $\hat{\Delta}^{LB} = e_1^\top \hat{\theta}_n - e_2^\top \hat{\theta}_n$, where $e_1$ and $e_2$ are standard basis vectors selecting the appropriate components.

This unified approach allows the Lee–IPW estimator to be implemented efficiently in both large and small samples, providing robust inference for the always-observed average treatment effect even when standard approaches fail.

remark[Heterogeneous shares and label-based designs] Under heterogeneous treatment shares, randomization delivers independence within strata but not unconditional independence, so unconditional Lee bounds are not valid. The Lee-IPW bounds developed in this section restore validity under the covariate-based stratification framework used here. When strata are defined by many discrete labels, we consider the same Lee-IPW bounds, and Appendix (ref) establishes the corresponding partial identification and inference results in a label-based triangular-array setting.

Inference for Moment-Based Estimators under Heterogeneous-Shares Stratification

Asymptotic Distribution under Heterogeneous-Shares Stratification

This section develops inference procedures for the Lee-IPW bounds introduced in Section (ref), focusing on the statistical challenges that arise when treatment shares differ across strata. Such heterogeneity is common in empirical studies, but falls outside the equal-share framework discussed in Section (ref) and in bai2023efficiency. Consequently, standard variance formulas and asymptotic results do not generally apply, making it necessary to extend existing theory for valid inference. The main contribution here, then, is to generalize the central limit theorem of bai2023efficiency to accommodate stratum-specific treatment shares, covering any moment estimator that satisfies the required regularity conditions and yielding valid, design-consistent inference in complex settings. The Lee-IPW bounds estimator, introduced in this article, demonstrates the usefulness of this broader framework.

Now, we formally analyze asymptotic inference for a broad class of generalized method of moments (GMM) estimators under stratified experimental designs with possibly heterogeneous treatment shares, $\eta_g$, across strata. The approach applies to any estimator whose moment vector satisfies the regularity conditions in Assumption (ref), which ensures the validity of the central limit theorem in this setting. The Lee-IPW bounds introduced in Section (ref) represent an important instance of this general framework, but the results here accommodate any such moment-based estimator.

Let $m(Z_i,\theta)$ denote a vector of moment functions that may depend on observed data and known design shares, and let \[ M := \left.\frac{\partial}{\partial\theta'}\,\operatorname{\mathbb{E}}\!\left[m(Z_i,\theta)\right]\right|_{\theta=\theta_0} \] be the corresponding population Jacobian. Allowing for heterogeneous treatment shares leads to an asymptotic “meat” matrix of the form

align*[align* omitted — 355 chars of source]

where $\mu_d(X) := \operatorname{\mathbb{E}}[m(X,d,Y(d);\theta_0)\mid X]$ and $\eta(X):=\Pr(D=1\mid X)$. In this notation, $X$ denotes the observed covariates used by the design, which may include stratum indicators.

theorem[Asymptotics under heterogeneous-shares stratification] Let $\hat{\theta}_n$ be the GMM estimator defined by the sample analog of $m(Z_i,\theta)$. Suppose Assumptions (ref), (ref), (ref), (ref), (ref), and (ref) hold. Then, \[ \sqrt{n}\,(\hat\theta_n-\theta_0)\ \xrightarrow{d}\ \mathcal N(0,V_\ast), \qquad V_\ast = M^{-1}\Omega_{\eta(\cdot)} M^{-1\prime}, \] where $M$ is nonsingular.
proofSee Appendix (ref) for a detailed proof.

This result establishes asymptotic normality and provides a variance formula for any moment estimator that satisfies the required regularity conditions. In particular, when applied to the Lee-IPW bounds estimator, it yields the asymptotic distribution and variance expressions for the lower and upper bounds.

Variance Decomposition and Estimation under Heterogeneous-Shares Stratification

The asymptotic covariance matrix $V_{\eta(\cdot)}$ from Theorem (ref) extends the equal-share variance structure to handle arbitrary, stratum-specific treatment shares. This formulation is valid for any estimator whose moment vector satisfies the regularity conditions outlined earlier, including the Lee-IPW bounds as our example.

To establish consistency of the variance estimator under heterogeneous shares, we impose two additional regularity conditions that are directly analogous to Assumptions 3.4--3.5 in bai2023efficiency. The first controls the construction of the involution $\pi(\cdot)$ used to handle singleton within-arm cross-products by requiring paired blocks to be asymptotically close in covariates. The second provides local regularity (integrability, empirical process, and smoothness) conditions ensuring the stability of the variance components that involve conditional means and second moments.

assumption[Paired-block covariate closeness] Let $\Lambda(X^{(n)})=\{\lambda_g:g=1,\dots,G\}$ be the covariate-based blocks from Section (ref), and let $\pi:\{1,\dots,G\}\to\{1,\dots,G\}$ be a fixed involution (a pairing map) such that $\pi(\pi(g))=g$ and $\pi(g)\neq g$ for all $g$. Assume that the paired blocks are asymptotically close in baseline covariates: \[ \frac{1}{G}\sum_{g=1}^{G}\ \max_{i\in\lambda_g,\ i'\in\lambda_{\pi(g)}} \|X_i-X_{i'}\|^2 \ \xrightarrow{p}\ 0. \]

\noindentRemark. Under Assumption (ref) (covariate balancing), it is typically possible to re-order the blocks and then construct $\pi(\cdot)$ so that $(g,\pi(g))$ pairs blocks with similar covariates. For example, one may apply a nonbipartite matching procedure to the block means $\{\bar X_g\}_{g=1}^G$, where $\bar X_g := N_g^{-1}\sum_{i\in\lambda_g} X_i$, to obtain a pairing $\pi(\cdot)$ such that Assumption (ref) holds.

assumption[Local regularity for variance components] There exists $\delta>0$ such that, for each $d\in\{0,1\}$: \begin{enumerate}[label=(\alph*)] • (Local uniform integrability). \begin{align*} \lim_{\lambda\to\infty}\ \operatorname{\mathbb{E}}\Bigg[ &\sup_{\theta\in\Theta:\ \|\theta-\theta_0\|<\delta} \big\|m\!\left(X_i,d,Y_i(d);\theta\right)\big\|^2 \\ &\qquad\times \mathbf 1\Bigg\{ \sup_{\theta\in\Theta:\ \|\theta-\theta_0\|<\delta} \big\|m\!\left(X_i,d,Y_i(d);\theta\right)\big\| >\lambda \Bigg\} \Bigg] =0. \end{align*} • (Local Glivenko--Cantelli for conditional means and second moments). For each $1\le s\le d_\theta$, the classes \begin{align*} \Big\{ &\operatorname{\mathbb{E}}\!\big[m_s\!\left(X_i,d,Y_i(d);\theta\right)\mid X_i=x\big] :\ \|\theta-\theta_0\|<\delta \Big\} \end{align*} and \begin{align*} \Big\{ &\operatorname{\mathbb{E}}\!\big[ m_s\!\left(X_i,d,Y_i(d);\theta\right)\, m\!\left(X_i,d,Y_i(d);\theta\right)^\top \mid X_i=x\big] :\ \|\theta-\theta_0\|<\delta \Big\} \end{align*} are $Q$-Glivenko--Cantelli. • (Local uniform Lipschitzness in covariates). Each component of $\operatorname{\mathbb{E}}[m(X,d,Y(d);\theta)\mid X=x]$ and $\operatorname{\mathbb{E}}[m(X,d,Y(d);\theta)m(X,d,Y(d);\theta)^\top\mid X=x]$ is Lipschitz in $x$ with a constant that is uniform over $\{\theta\in\Theta:\ \|\theta-\theta_0\|<\delta\}$; that is, there exists $L_d<\infty$ such that for all $x,x'\in\mathbb R^{d_x}$, \begin{align*} \sup_{\theta\in\Theta:\ \|\theta-\theta_0\|<\delta} \Big| &\operatorname{\mathbb{E}}\!\big[m_s(X,d,Y(d);\theta)\mid X=x\big] - \operatorname{\mathbb{E}}\!\big[m_s(X,d,Y(d);\theta)\mid X=x'\big] \Big| \\ &\le L_d\|x-x'\|, \end{align*} and similarly for each component of $\operatorname{\mathbb{E}}[m(X,d,Y(d);\theta)m(X,d,Y(d);\theta)^\top\mid X=x]$. \end{enumerate}

An important contribution of this section is a decomposition of the “meat” of the sandwich variance in terms of unconditional moments, extending beyond the conditional variance formulation\footnote{We propose a novel decomposition for the heterogeneous-shares variance by rewriting $\Omega_{\eta(\cdot)}$ as in (ref). This differs from the variance decomposition for the equal-share case in bai2023efficiency and yields a slightly different plug-in implementation. bai2023efficiency discusses heterogeneous assignment probabilities and states the asymptotic variance in its original form with conditional variances, whereas our decomposition removes these conditional terms, allowing for tractable and consistent plug-in variance estimators.}. This expression is derived by applying the definition of conditional variance and properties of expectations to the original asymptotic variance formula:

align[align omitted — 808 chars of source]

This unconditional decomposition is particularly important in practice, since directly estimating conditional variances within small strata can be highly variable or even infeasible when some strata contain very few treated or control units. By expressing the variance entirely in terms of unconditional means and cross-products, this approach enables reliable and consistent variance estimation even in finely stratified designs with heterogeneous treatment shares.

To facilitate practical estimation, the variance formula can be expressed as a collection of sample averages and cross-products that are straightforward to compute. Specifically, let $\hat m_i = m(Z_i, \hat\theta_n)$ denote the estimated moment for each unit. For each stratum $g$, let $N_{d,g}$ represent the number of units assigned to arm $d$, $w_g = N_g/n$ the weight of stratum $g$ in the sample, and $\eta_g = T_g/N_g$ the observed treatment share. For empirical implementation, the variance components can be written as functions of sample averages and cross-products computed from the observed data. The primary unit-level quantities are

align*[align* omitted — 311 chars of source]

In addition to these unit-level terms, cross-products within each block are incorporated to reduce noise and stabilize variance estimation by pooling information across pairs, which is especially important in finely stratified designs with small or uneven strata. These are defined as

align*[align* omitted — 719 chars of source]

Here, for $d\in\{0,1\}$, the within-arm cross-product component $\widehat\varsigma_{g,n}(d,d)$ is defined as \[ \widehat\varsigma_{g,n}(d,d) :=

cases\frac{2}{N_{d,g}(N_{d,g}-1)} \sum_{\substack{i<i'\\ i,i'\in\lambda_g:\,D_i=D_{i'}=d}} \hat m_i\,\hat m_{i'}^{\top} , & if N_{d,g}\ge 2,\\[1.1em] \hat m_{i_g(d)}\,\hat m_{i_{\pi(g)}(d)}^{\top}, & if N_{d,g}=1,

\] where $i_g(d)$ denotes the (unique) index in block $g$ such that $D_{i_g(d)}=d$ when $N_{d,g}=1$. In the case $N_{d,g}=1$, the between-block product uses a paired block $\pi(g)$, where $\pi:\{1,\dots,G\}\to\{1,\dots,G\}$ is a fixed involution with $\pi(\pi(g))=g$ and $\pi(g)\neq g$ for all $g$.\footnote{After re-ordering blocks, one may construct $\pi(\cdot)$ so that paired blocks $(g,\pi(g))$ are asymptotically close in covariates (cf.\ Assumption 3.4 of bai2024inference).}

With these components, the estimated variance “meat” is assembled as \[ \widehat\Omega_n = \hat{A}_{1,n} + \hat{A}_{0,n} + \hat{B}_n - \hat{A}_{3,n}, \] and the full estimated covariance matrix is \[ \widehat V_{\eta(\cdot)} = \hat M_n^{-1} \widehat\Omega_n\, \hat M_n^{-T}, \qquad \hat M_n = \frac{1}{n} \sum_{i=1}^n \partial_\theta m(Z_i, \hat\theta_n). \] For the Lee-IPW lower and upper bounds, the standard error is given by the variance of a difference, that is, \[ \hat\sigma_{\bullet}^2 = \widehat V_{\eta(\cdot),11}^{\,\bullet} + \widehat V_{\eta(\cdot),22}^{\,\bullet} - 2\,\widehat V_{\eta(\cdot),12}^{\,\bullet}, \] where $\widehat V_{\eta(\cdot),ij}^{\,\bullet}$ denotes the $(i,j)$ entry of the estimated covariance matrix, constructed using the appropriate moment conditions for the lower or upper bound.

theorem[Consistency of $\widehat V_{\eta(\cdot)}$ under heterogeneous-shares stratification] Suppose Assumptions (ref), (ref), (ref), (ref), (ref), and (ref) hold, and $\hat M_n \xrightarrow{p} M$. Let $\widehat\Omega_n$ and $\widehat V_{\eta(\cdot)}$ be defined as in this section. Then, \[ \widehat V_{\eta(\cdot)} \xrightarrow{p} V_{\eta(\cdot)}:=M^{-1}\Omega_{\eta(\cdot)}M^{-T}. \] Consequently, $n^{-1} \hat\sigma_{\bullet}^2$ delivers consistent variances for $\hat\Delta^{LB}$ and $\hat\Delta^{UB}$, using the appropriate moment vectors.
proofSee Appendix (ref) for a detailed proof.

The estimator relies on sample averages and block-level cross-products, making it broadly applicable and efficient in computation. The correction term $\hat{B}_n$ is particularly important when treatment shares are highly variable or block sizes are small, as it adjusts for the negative covariance introduced by fixing treated counts within strata. Neglecting this feature can lead to overconservative inference. The variance decomposition and estimation methods presented here apply not only to Lee-IPW bounds, but more generally to any estimator meeting the required regularity conditions in stratified experimental designs with heterogeneous shares.

Empirical Implementation of Standard Errors for Lee-IPW Bounds

The analytic sandwich variance estimator developed in Section (ref) is recommended as the primary method for inference, as it is consistent under a wide range of stratification and blocking schemes, and directly incorporates the negative covariance that arises when the number of treated units per stratum is fixed by design. This feature distinguishes it from classical i.i.d.\ and equal-share variance estimators, which do not fully capture the dependence structure present in complex experimental designs with heterogeneous treatment shares. In empirical applications, the sandwich estimator is generally preferable for constructing confidence intervals and conducting hypothesis tests, especially when treatment shares or block sizes are unbalanced.

In all cases, Lee-IPW bounds should be accompanied by confidence intervals constructed using the recommended procedures lee2009training and the standard error estimator presented here. When the confidence limits agree in sign, conclusions regarding the sign of the treatment effect are robust to monotone selection. For designs with highly variable treatment shares or small blocks, special attention should be paid to the correction term in the variance formula, as neglecting this adjustment may result in overly conservative inference.

Overall, the methods presented in this section extend an existing inference framework to accommodate arbitrary heterogeneity in treatment assignment and block structure. They provide a coherent and flexible approach for valid uncertainty quantification, applicable not only to Lee-IPW bounds but to a broad class of moment estimators in stratified experiments.

Conclusion

This paper develops a unified, design-consistent approach for valid inference on Lee bounds in randomized experiments subject to differential outcome attrition and stratified randomization. We establish that, under equal-share stratification, the variance reduction induced by fixing the treatment allocation within blocks can be fully captured using recent theory for finely stratified experiments. Our results clarify that conventional variance estimators based on i.i.d.\ assumptions systematically overstate uncertainty when block-level assignment is present, leading to unnecessarily conservative inference. By instead using design-consistent variance estimators that incorporate the block structure and fixed treatment shares, practitioners can achieve tighter, more informative confidence intervals for Lee bounds.

Among its contributions, our work introduces a formal innovation by extending the scope of valid inference beyond covariate-based stratification to include designs where stratification is determined solely by discrete block labels, such as schools or survey waves, a setting not previously explicitly accommodated in the literature bai2023efficiency. We show that the assumptions required for sharp identification and consistency of Lee bounds are preserved under equal-share stratification, and that the efficiency gains from block randomization are readily realized under both covariate-based and label-based stratification settings.

When treatment shares vary across blocks, conventional Lee bounds based on unconditional trimming are no longer valid, and conditional bounds constructed within small or sparsely treated strata can be highly variable, biased, or undefined. To address these challenges, we introduce the Lee–IPW bounds, which employ inverse probability weighting and a global trimming rule to restore validity and robustness in heterogeneous-shares stratified designs. For these bounds, we provide a closed-form, design-consistent variance estimator that ensures reliable inference even with small or unbalanced blocks.

Beyond this specific partial identification setting, we extend existing inference results to accommodate a broad class of moment-based estimators under heterogeneous-shares stratified designs. This extension builds on the literature on finely stratified experiments, enabling valid and efficient inference in a wider range of complex experimental settings.

Practically, our recommendations are as follows. Researchers should always retain and account for the full block structure of their experimental design, record actual treatment shares within each stratum, and use the appropriate design-consistent variance estimator. Specifically, $\widehat V_{\eta}$ should be used for equal shares, and $\widehat V_{\eta(\cdot)}$ for heterogeneous shares. When strata are small or sparsely treated, the Lee–IPW bounds provide a robust and well-defined alternative to conditional Lee bounds, as they avoid the instability and potential bias that can arise from such procedures in these settings. In all cases, relying on standard i.i.d.\ variance formulas ignores negative assignment dependence and typically yields overly conservative confidence intervals.

In summary, this work provides a general, practical, and efficient set of tools for valid inference on both conventional Lee bounds and the newly introduced Lee–IPW bounds in the presence of outcome attrition under a wide range of stratified experimental designs. Beyond these partial identification settings, our results further extend to inference for a broad class of moment-based estimators, enabling design-consistent variance estimation even in experiments with complex or heterogeneous-shares stratification. By connecting partial identification with more contemporaneous inference theory and robust variance estimation, these advances help ensure that treatment effect analyses remain both credible and informative across diverse empirical research contexts.