The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
95,882 characters
Bootstrap Inference with a Randomly Assigned Regressor: Covariance Filtering and the Limits of Marginal Resampling
\maketitle
\begin{abstract}
Random assignment can make ordinary least squares (OLS) inference insensitive to outcome dependence, yet iid resampling can still fail because assignment and resampling need not remove the same covariance terms. With binary treatment, the iid variance target differs from the sampling variance by exactly minus aggregate cross-unit covariance of treatment effects. Under a Gaussian first-order limit, positive covariance leads to over-rejection and negative covariance to under-rejection. Two designs can generate the same distribution for each observation but require variance corrections of opposite signs, so no marginal-only variance correction is first-order exact for both. For first-order Gaussian inference, only one covariance component---the covariance carried by the randomized-regressor score---must be recovered. Standard dependence estimators on that score restore validity under suitable ordered or grouped dependence conditions. In simulations with ordered data, fixed-bandwidth calibration reduces several large-bandwidth size distortions.
\end{abstract}
\medskip\noindent\textbf{JEL codes:} C12, C13, C15, C21.\qquad
\textbf{Keywords:} bootstrap; random assignment; heterogeneous treatment effects; covariance filtering; long-run variance; cross-sectional dependence.
\section{Introduction}
Random assignment can make inference much simpler than the dependence structure of the observed data suggests. \citet{CHLS2026} (hereafter CHLS) show that, for the ordinary least squares (OLS) coefficient on a randomly assigned regressor, residualizing the regressor against fixed-dimensional controls is asymptotically absorbed by the assignment mechanism: under their conditions, the leading score is the centered randomized regressor multiplied by the regression error. Broad dependence in outcomes, controls, and regression errors can therefore disappear from the first-order variance or survive only through a restricted covariance component. That simplification might suggest that iid resampling is automatically valid. Our question is sharper: \emph{does independent resampling reproduce exactly the covariance cancellations induced by random assignment in the sampling law?}
The answer is no in general. An iid bootstrap deletes cross-observation covariance from the resampled score by construction. Random assignment also filters covariance, but it need not delete the same terms. Random assignment may therefore simplify the sampling variance even though a seemingly natural iid bootstrap still targets the wrong first-order variance. This problem is distinct from the bootstrap-moment pathology of \citet{HahnLiao2021}: the variance target can already be wrong before conditional second moments become an issue.
\paragraph{Scope of the assignment design.}
Throughout the formal results, assignment is iid at the unit level. Fixed-count complete randomization, stratified or covariate-adaptive assignment, matched designs, and cluster-level assignment alter the assignment covariance structure and can change the relevant filter at first order; Remark~\ref{rem:designboundary} gives an explicit fixed-count counterexample. The formulas below do not carry over automatically to those designs. For such experiments, one must first derive the design-specific score filter and then choose an appropriate resampling or variance correction; see, for example, \citet{BugniCanayShaikh2018} for covariate-adaptive randomization.
The mechanism is most transparent with binary treatment. Suppose $D_i$ is independently randomized and
\[
Y_i=Y_i(0)+D_i\tau_i,
\qquad
\tau_i=Y_i(1)-Y_i(0).
\]
Let $V_n$ denote the first-order sampling variance of the leading OLS statistic and $V_n^\#$ the diagonal variance targeted by iid score resampling. Define
\[
\Delta_{\tau,n}=\frac1n\sum_{i\ne j}\operatorname{Cov}(\tau_i,\tau_j).
\]
We prove the exact leading-score identity
\[
\boxed{V_n^\#-V_n=-\Delta_{\tau,n}.}
\]
Under the standardized Gaussian first-order limit used for inference below, positive aggregate covariance makes iid inference anti-conservative; negative covariance makes it conservative; covariance cancellation can leave it first-order correct despite dependence. The variance discrepancy itself does not depend on the Bernoulli assignment probability $p$, even though both $V_n$ and $V_n^\#$ do.
Correlated treatment effects under individual assignment are economically natural when units share \emph{effect modifiers}, even without interference. Returns to an individually randomized training program can co-move within local labor markets because local demand conditions affect the payoff to training. Returns to a household subsidy can co-move across nearby cohorts because macroeconomic conditions affect take-up returns. Treatment effects can also co-move within schools because students face common teachers, curricula, or institutional environments. In each example assignment can remain iid at the unit level while the response to treatment is correlated across units.
The general covariance filter nests the binary identity. When the regression error takes the assignment-dependent form
\[
\varepsilon_i=\lambda(D_i)'e_i,
\]
with a mean-zero latent vector $e_i$, let
\[
\sigma_D^2=\operatorname{Var}(D_i),
\qquad
\mu_A=E[(D_i-E D_i)\lambda(D_i)].
\]
Random assignment eliminates every off-diagonal covariance direction except the scalar projection $\mu_A'e_i$. If
\[
\Delta_{A,n}=\frac1n\sum_{i\ne j}\operatorname{Cov}(\mu_A'e_i,\mu_A'e_j),
\]
then
\[
\boxed{V_n^\#-V_n=-\sigma_D^{-4}\Delta_{A,n}.}
\]
The bootstrap problem is therefore lower-dimensional than the observed-data dependence problem: only the covariance surviving the assignment-induced projection matters for the coefficient of interest.
The information-limit result makes this restriction explicit. We construct two Gaussian moving-average (MA(1)) designs satisfying the maintained assumptions with exactly the same one-observation law of $(D_i,Y_i)$ and the same diagonal bootstrap target, but opposite treatment-effect covariance and different first-order sampling variances. Any variance rule asymptotically determined only by the one-observation marginal law must fail for at least one design. For centered-Gaussian bootstraps this rules out universal first-order validity over the pair; under the conditional moment condition of Theorem~\ref{thm:ui}, the same information limit applies to bootstrap second moments more generally. Ordered or grouped observations can distinguish the designs. The missing information is therefore genuinely joint rather than marginal.
A complementary minimality result shows that full dependence estimation is unnecessary. Conditional on consistently estimating the observation-by-observation variance component, a centered-Gaussian procedure is first-order valid if and only if it recovers the single surviving covariance contribution at the first-order variance scale. The associated covariance quotient is one-dimensional whenever $\mu_A\ne0$. The information lower bound and the dimensionality result therefore line up: marginals are insufficient, but one projected covariance contribution is enough.
\paragraph{Relation to design-based clustering and causal bootstrap.}
\citet{AbadieEtAl2023} show in a design-based framework that clustering can matter when assignment is clustered or heterogeneous treatment effects co-move within clusters. Our binary identity is related to that heterogeneity channel; our question is different: whether independent resampling removes covariance that random assignment leaves in the sampling law. Our comparison leads to the general covariance-filter identity, the marginal-information limit, and the one-dimensional minimality result. \citet{ZhangEtAl2025} provide a sampling-side counterpart by showing that least-squares inference with random regressors can be robust to correlated errors. Our strong-exogeneity benchmark shares that intuition; our contribution begins when the regression error responds to assignment.
\citet{AbadieAtheyImbensWooldridge2020} distinguish sampling- and design-based uncertainty, while \citet{ImbensMenzel2021} develop a finite-population causal bootstrap. Bootstrap inference under matched-pair or covariate-adaptive assignment is studied by \citet{JiangEtAl2024,ZhangZheng2020}, with general covariate-adaptive inference in \citet{BugniCanayShaikh2018}. Those designs make assignment itself dependent; here assignment is iid, so surviving covariance comes from potential outcomes or errors. Our constructive step then uses standard heteroskedasticity-and-autocorrelation-consistent (HAC), block, cluster, or dependent-bootstrap methods \citep{Kunsch1989,PolitisRomano1994,Shao2010,Hounyo2023,CameronEtAl2008}; related cluster guidance includes \citet{DjogbenouEtAl2019,CanaySantosShaikh2021,MacKinnonNielsenWebb2023}. Group-based $t$ and network-HAC alternatives include \citet{IbragimovMuller2010,KojevnikovEtAl2021}.
Ordinary pairs resampling does not evade the problem. Its linearized target is the usual heteroskedasticity-robust HC0 variance for the residualized randomized-regressor score and therefore inherits the same diagonal discrepancy. Under their respective consistency conditions, HAC, cluster, or spatial variance estimators applied to that score can recover $V_n$. The problem is dependence-ignorant diagonal inference, whether analytic or based on iid row resampling.
The simulations illustrate these implications. Time-series designs cover positive, negative, and cancellation cases. In a cluster design with iid individual assignment and correlated treatment effects, iid score inference rejects $18.5\%$ at a nominal $5\%$ level when cluster size is eight, while the correctly grouped projected correction rejects about $5.1\%$. Deliberate grouping misspecification shows that the correction is only as good as the dependence structure supplied to it. For ordered data we document bandwidth sensitivity and compare small-bandwidth normal with fixed-bandwidth calibration. Because the relevance ratio is noisier, we treat it only as a descriptive summary rather than a pretest.
The remainder of the paper first defines the sampling and bootstrap targets, establishes the strong-exogeneity benchmark, and derives the covariance filter. It then proves the marginal information limit and one-dimensional minimality result, specializes to binary treatment and ordinary pairs resampling, develops the projected-score correction and diagnostic, and reports Monte Carlo evidence. We also show that once the iid and projected tests are calibrated to the same first-order size, their Gaussian local-power curves coincide. Proofs, primitive sufficient conditions, and additional simulations are collected in the numbered Online Appendix.
\section{Setup and inferential target}\label{sec:setup}
We use a triangular array where needed and suppress $n$ subscripts when no ambiguity arises. Within each row, the assignment variables $\{D_{i,n}\}_{i=1}^n$ are iid with a common law and the full assignment vector is independent of the latent/control array $\{(e_{i,n}',W_{i,n}')'\}_{i=1}^n$. Dependence across observations in the latent variables and controls is otherwise allowed subject to the conditions of each result. The dimensions of $W_{i,n}$ and $e_{i,n}\in\mathbb R^L$ are fixed as $n$ grows. The first component of $W_i$ is a nonzero constant. We assume $\sigma_{D,n}^2>0$ for all sufficiently large $n$; stronger lower bounds are imposed only where needed. Thus $\mu_D$ and $\sigma_D^2$ below denote the row-$n$ quantities $\mu_{D,n}$ and $\sigma_{D,n}^2$ whenever a triangular-array statement is in force. Sample-matrix inverses are understood on their nonsingularity events. On exceptional samples the associated statistic may be assigned an arbitrary fixed value; the relevant theorems ensure that such events have probability tending to zero. Define
\[
\mu_D=\mathbb E[D_i],\qquad \sigma_D^2=\operatorname{Var}(D_i),\qquad D_i^c=D_i-\mu_D.
\]
The regression model is
\[
Y_i=D_i\beta+W_i'\gamma+\varepsilon_i.
\]
Stack the observations as $Y=(Y_1,\ldots,Y_n)'$, $D=(D_1,\ldots,D_n)'$, and $\varepsilon=(\varepsilon_1,\ldots,\varepsilon_n)'$, and let $W$ be the $n\times K$ matrix with $i$th row $W_i'$, where the fixed control dimension is $K$. Let $I_n$ denote the $n\times n$ identity matrix and define the residual-maker
\[
M_W=I_n-W(W'W)^{-1}W'.
\]
The scalar $\widehat\beta$ denotes the OLS coefficient on $D$. By the Frisch--Waugh--Lovell theorem,
\[
\widehat\beta-\beta=(D'M_WD)^{-1}D'M_W\varepsilon.
\]
The Frisch--Waugh--Lovell formula above is an exact finite-sample identity for OLS. The next step is asymptotic. We take from CHLS the representation
\begin{equation}
\sqrt n(\widehat\beta-\beta)=\sigma_{D,n}^{-2}\frac1{\sqrt n}\sum_{i=1}^n D_{i,n}^c\varepsilon_{i,n}+o_p(1).
\label{eq:chlsrep}
\end{equation}
Equation~\eqref{eq:chlsrep} requires one clarification about the controls. In the exact Frisch--Waugh--Lovell formula, $D$ is residualized against $W$. Under the CHLS conditions, iid assignment makes the normalized sample projection of the centered assignment draw on the fixed-dimensional controls asymptotically negligible, while $n^{-1}D'M_WD$ converges to the assignment variance. Consequently, the control-projection terms contribute only $o_p(1)$ to the normalized numerator and the first-order score reduces to $(D_{i,n}-\mu_{D,n})\varepsilon_{i,n}$. This is a first-order statement, not a claim that the controls are absent from the finite-sample OLS formula or from the joint distribution of the data.
Our target is the unconditional CHLS superpopulation sampling law, not the finite-population randomization distribution; this distinction follows the sampling-versus-design perspective emphasized by \citet{AbadieAtheyImbensWooldridge2020}. Accordingly, the variance objects below are deterministic variance sequences for the \emph{leading score} in \eqref{eq:chlsrep}; they are not asserted to equal the exact finite-sample variance of the nonlinear OLS estimator.
Define the first-order score
\[
\psi_i=D_i^c\varepsilon_i,
\]
its sampling quadratic scale
\[
\sigma_{D\varepsilon,n}^2
=
\operatorname{Var}\!\left(
n^{-1/2}\sum_{i=1}^n\psi_i
\right),
\]
the diagonal score scale
\[
s_{\#,n}^2
=
\frac1n\sum_{i=1}^nE[\psi_i^2],
\]
the corresponding diagonal bootstrap target
\[
V_n^\#
=
\sigma_{D,n}^{-4}s_{\#,n}^2,
\]
and the first-order sampling-variance sequence
\[
V_n=\sigma_{D,n}^{-4}\sigma_{D\varepsilon,n}^2.
\]
Thus $V_n$ is exactly the variance of the leading linearized statistic
\[
\sigma_{D,n}^{-2}n^{-1/2}\sum_{i=1}^n\psi_{i,n},
\]
while its interpretation as the first-order variance of $\sqrt n(\widehat\beta-\beta)$ uses \eqref{eq:chlsrep}. If $V_n$ can shrink, the unscaled remainder $o_p(1)$ in \eqref{eq:chlsrep} is not by itself enough to obtain a standardized limit: one needs the stronger remainder $o_p(V_n^{1/2})$ or, equivalently for our purposes, a standardized CHLS conclusion stated directly. Every theorem below that permits a shrinking variance scale therefore assumes the standardized sampling law explicitly. These objects are defined before specializing the error structure, so the strong-exogeneity and assignment-dependent-heteroskedasticity branches share the same normalization. Throughout, $P^*$, $E^*$, and $\operatorname{Var}^*$ denote probability, expectation, and variance conditional on the observed sample. We write $\Phi$ for the standard normal distribution function and $z_q=\Phi^{-1}(q)$ for its $q$th quantile. To avoid collisions, $L$ is reserved for the latent-vector dimension, $m_n$ for a theoretical kernel denominator, $\ell$ for the largest included positive lag reported in the Bartlett simulations, $\varrho_{\psi,n}(h)$ and $\varrho_{\tau,n}(h)$ for score and treatment-effect autocovariances, and $h_{\mathrm{loc}}$ for a local-alternative shift.
\begin{assumption}[Stable scale]\label{ass:stable}
There exist $0<c<C<\infty$ such that for all sufficiently large $n$,
\[
c\le \sigma_{D,n}^2\le C,
\qquad
c\le \sigma_{D\varepsilon,n}^2\le C,
\qquad
c\le s_{\#,n}^2\le C.
\]
\end{assumption}
\begin{lemma}[Randomized-design LLN]\label{lem:design}
Let $\{D_{i,n}\}_{i\le n}$ be row-wise iid with mean $\mu_{D,n}$ and variance $\sigma_{D,n}^2$, and define
\[
\bar D_n=n^{-1}\sum_{i=1}^nD_{i,n},
\qquad
\widehat Q_n=n^{-1}\sum_{i=1}^n(D_{i,n}-\bar D_n)^2.
\]
Suppose for some $\delta\in(0,2]$,
\[
\sup_n E|D_{1,n}-\mu_{D,n}|^{2+\delta}<\infty.
\]
Then
\[
\bar D_n-\mu_{D,n}=O_p(n^{-1/2}),
\qquad
\widehat Q_n-\sigma_{D,n}^2=o_p(1).
\]
If, in addition, $\inf_n\sigma_{D,n}^2>0$, then $\widehat Q_n/\sigma_{D,n}^2\to_p1$.
\end{lemma}
\subsection{Assignment-dependent heteroskedasticity and heterogeneous effects}
Let $\lambda(d)\in\mathbb R^L$ denote the assignment-dependent loading vector and write
\[
\varepsilon_i=\lambda(D_i)'e_i,
\qquad A_i=(D_i-\mu_D)\lambda(D_i),
\qquad \mu_A=\mathbb E[A_i],
\]
with a mean-zero latent vector $e_i\in\mathbb R^L$, as in CHLS, so that $\mathbb E e_i=0$. Put $\widetilde A_i=A_i-\mu_A$. Then
\[
\sigma_{D\varepsilon,n}^2
=
\sigma_{e,1,n}^2+\sigma_{e,2,n}^2,
\]
where
\[
\sigma_{e,1,n}^2=\frac1n\sum_{i=1}^n\mathbb E[(\widetilde A_i'e_i)^2],
\qquad
\sigma_{e,2,n}^2=\operatorname{Var}\left(n^{-1/2}\sum_{i=1}^n\mu_A'e_i\right).
\]
Thus the CHLS conditional-heteroskedastic representation decomposes the generic score variance rather than defining a separate normalization.
\subsection{Gaussian multiplier bootstrap target and null-imposed score}
Using the score $\psi_i$ defined above, let $\xi_i\sim N(0,1)$ be iid and independent of the data. Define
\[
T_{n,o}^\#=\sigma_{D,n}^{-2}n^{-1/2}\sum_{i=1}^n\xi_i\psi_{i,n},
\qquad
\widehat V_{n,o}^\#=\sigma_{D,n}^{-4}n^{-1}\sum_{i=1}^n\psi_{i,n}^2.
\]
This is an infeasible device that isolates the covariance effect.
For testing the null hypothesis $H_0:\beta=\beta_0$, define the sample assignment mean $\bar D=n^{-1}\sum_{i=1}^nD_i$ and the null-imposed nuisance estimator
\[
\widetilde\gamma(\beta_0)=(W'W)^{-1}W'(Y-D\beta_0),
\]
the restricted residual
\[
\widetilde\varepsilon_i(\beta_0)=Y_i-D_i\beta_0-W_i'\widetilde\gamma(\beta_0),
\]
and the restricted score
\[
\widetilde\psi_i(\beta_0)=(D_i-\bar D)\widetilde\varepsilon_i(\beta_0).
\]
Let
\[
\widehat Q_n=n^{-1}\sum_{i=1}^n(D_i-\bar D)^2,
\]
and
\[
T_n^\#(\beta_0)=\widehat Q_n^{-1}n^{-1/2}\sum_{i=1}^n\xi_i\widetilde\psi_i(\beta_0).
\]
Conditional on the sample, this statistic is exactly Gaussian with variance
\[
\widetilde V_n^\#(\beta_0)=\widehat Q_n^{-2}n^{-1}\sum_{i=1}^n\widetilde\psi_i(\beta_0)^2.
\]
\paragraph{Test inversion and confidence sets.}
The feasible score and its variance are null imposed: both are recomputed at the hypothesized value $\beta_0$. Accordingly, the formal results below justify tests of $H_0:\beta=\beta_0$. A $(1-\alpha)$ confidence set is obtained by inversion,
\[
\mathcal C_{1-\alpha}
=
\{\beta_0:\text{the level-}\alpha\text{ null-imposed test does not reject}\}.
\]
In general this set need not equal a Wald interval of the form $\widehat\beta\pm z_{1-\alpha/2}\widehat{\mathrm{se}}$ because the variance estimate may vary with $\beta_0$. A conventional null-independent Wald interval would require a separate consistency argument for an unrestricted-score variance estimator; we do not infer that result from the null-imposed theory.
The following branch-neutral lemma records the only nuisance-replacement result needed by both the strong-exogeneity benchmark and the feasible conditional-heteroskedastic analysis.
\begin{lemma}[Restricted-score replacement]\label{lem:restricted}
Under $H_0:\beta=\beta_0$, suppose
\[
\|(n^{-1}W'W)^{-1}\|=O_p(1),
\qquad
n^{-1}W'\varepsilon=o_p(1),
\]
\[
\frac1n\sum_i\varepsilon_i^2=O_p(1),
\qquad
\frac1n\sum_i\|W_i\|^2=O_p(1),
\qquad
\sup_n\operatorname{Var}(D_{1,n})<\infty.
\]
Under the assignment structure in Section~\ref{sec:setup},
\[
\frac1n\sum_{i=1}^n
\{\widetilde\psi_i(\beta_0)-\psi_i\}^2=o_p(1).
\]
Under $H_0$,
\[
\widetilde\gamma(\beta_0)-\gamma
=(n^{-1}W'W)^{-1}(n^{-1}W'\varepsilon)=o_p(1)
\]
under the stated conditions. More precisely,
\[
\frac1n\sum_i(\widetilde\psi_i-\psi_i)^2
\le
2(\bar D-\mu_D)^2\frac1n\sum_i\varepsilon_i^2
+
2\|\widetilde\gamma-\gamma\|^2 B_n,
\]
where
\[
B_n=\frac1n\sum_i(D_i-\bar D)^2\|W_i\|^2=O_p(1).
\]
\end{lemma}
\section{Benchmark: when random assignment justifies iid bootstrap inference}
For row $n$, let
\[
\mathcal D_n=\sigma(D_{1,n},\ldots,D_{n,n}),
\qquad
\mathcal E_n=\sigma(\varepsilon_{1,n},\ldots,\varepsilon_{n,n}).
\]
We say that \emph{strong exogeneity} holds when the sigma-fields $\mathcal D_n$ and $\mathcal E_n$ are independent. Thus no regression error depends on any assignment draw. This is stronger than the assignment-dependent representation used later: in Section~\ref{sec:filter}, assignment is independent of the latent array $e$, but $\varepsilon_i=\lambda(D_i)'e_i$ may itself depend on $D_i$. Controls may remain dependent across units and correlated with the errors; assignment--control independence is already imposed in Section~\ref{sec:setup}.
\begin{lemma}[Scalar randomized-score LLN]\label{lem:scalar-diag}
Suppose the assignment structure in Section~\ref{sec:setup} and strong exogeneity hold and, for some $\delta\in(0,2]$,
\[
\sup_n E|D_{1,n}-\mu_{D,n}|^{2+\delta}<\infty,
\qquad
\frac1n\sum_{i=1}^n|\varepsilon_{i,n}|^{2+\delta}=O_p(1),
\]
and
\[
\frac1n\sum_{i=1}^n
\{\varepsilon_{i,n}^2-E(\varepsilon_{i,n}^2)\}
=o_p(1).
\]
Then
\[
\frac1n\sum_{i=1}^n
(D_{i,n}-\mu_{D,n})^2\varepsilon_{i,n}^2
-
\frac1n\sum_{i=1}^n
E[(D_{i,n}-\mu_{D,n})^2\varepsilon_{i,n}^2]
=o_p(1).
\]
\end{lemma}
Under strong exogeneity, for $i\ne j$,
\[
\mathbb E[(D_{i,n}-\mu_{D,n})(D_{j,n}-\mu_{D,n})\varepsilon_{i,n}\varepsilon_{j,n}]=0
\]
even if $\mathbb E[\varepsilon_i\varepsilon_j]\ne0$.
\begin{theorem}[Feasible iid multiplier score bootstrap under strong exogeneity]\label{thm:strong}
Suppose strong exogeneity holds, the conditions of Lemmas~\ref{lem:design}, \ref{lem:scalar-diag}, and \ref{lem:restricted}, and Assumption~\ref{ass:stable} hold. Let the standardized CHLS conclusion hold:
\[
\frac{\sqrt n(\widehat\beta-\beta)}{\sqrt{V_n}}\overset{d}{\to} N(0,1).
\]
Under strong exogeneity,
\[
V_n^\#=V_n.
\]
Under $H_0:\beta=\beta_0$,
\[
\widetilde V_n^\#(\beta_0)-V_n^\#=o_p(1),
\]
and hence
\[
\widetilde V_n^\#(\beta_0)-V_n=o_p(1),
\]
and
\[
\sup_{t\in\mathbb R}\left|
\Pr^*\{T_n^\#(\beta_0)\le t\}
-\Pr\{\sqrt n(\widehat\beta-\beta_0)\le t\}
\right|=o_p(1).
\]
The distributional conclusion does not require $V_n$ to converge to a fixed constant. Under Assumption~\ref{ass:stable}, however, $V_n=V_n^\#$ is bounded away from zero, so the absolute variance consistency displayed above is also relative consistency.
\end{theorem}
Because $\mathbb E^*T_n^\#=0$ exactly,
\[
\mathbb E^*[(T_n^\#)^2]=\widetilde V_n^\#.
\]
Thus, if $V_n\to V$ for some fixed $V\in(0,\infty)$,
\[
\mathbb E^*[(T_n^\#)^2]\overset{p}{\to} V.
\]
This does not follow from weak convergence alone; here it follows directly from variance consistency.
\section{When dependence survives random assignment: the covariance filter}\label{sec:filter}
Define
\[
\sigma_{\mu,0,n}^2=\frac1n\sum_i\operatorname{Var}(\mu_A'e_i),
\]
\[
\Delta_{A,n}=\frac1n\sum_{i\ne j}\operatorname{Cov}(\mu_A'e_i,\mu_A'e_j).
\]
Equivalently, with
\[
\Omega_{e,n}=\operatorname{Var}\left(n^{-1/2}\sum_i e_i\right),\qquad
\Omega_{e,0,n}=n^{-1}\sum_i\operatorname{Var}(e_i),
\]
we have
\[
\Delta_{A,n}=\mu_A'(\Omega_{e,n}-\Omega_{e,0,n})\mu_A.
\]
\begin{lemma}[Randomized diagonal-score LLN]\label{lem:diag}
Suppose the assignment structure in Section~\ref{sec:setup} holds and, for some $\delta\in(0,2]$,
\[
\sup_n E|D_{1,n}-\mu_{D,n}|^{2+\delta}<\infty,
\qquad
\frac1n\sum_{i=1}^n \|e_{i,n}\|^{2+\delta}=O_p(1),
\]
the assignment loading is uniformly bounded over the support of the randomized regressor,
\[
\sup_n\sup_{d\in\operatorname{supp}(D_{1,n})}\|\lambda_n(d)\|<\infty,
\]
and
\[
\left\|
\frac1n\sum_{i=1}^n
\{e_{i,n}e_{i,n}'-E[e_{i,n}e_{i,n}']\}
\right\|=o_p(1).
\]
Then
\[
\frac1n\sum_{i=1}^n\psi_{i,n}^2
-
\frac1n\sum_{i=1}^nE[\psi_{i,n}^2]
=o_p(1).
\]
\end{lemma}
\begin{theorem}[Oracle covariance-filter identity for the linearized score]\label{thm:filter}
Suppose the conditional-heteroskedastic representation in Section~\ref{sec:setup} holds, the iid assignment vector is independent of the full latent array $\{e_{i,n}\}_{i=1}^n$, $\sigma_{D,n}^2>0$ for all sufficiently large $n$, and, for every sufficiently large $n$ and $i\le n$,
\[
E\|A_{i,n}\|^2<\infty,
\qquad
E\|e_{i,n}\|^2<\infty.
\]
Because $\varepsilon_{i,n}=\lambda_n(D_{i,n})'e_{i,n}$, this assumption does \emph{not} require assignment to be independent of the regression-error array: the error may respond to $D_{i,n}$ through the loading $\lambda_n(D_{i,n})$. Independence is imposed on assignment and the latent array $e$, which is the condition used by the covariance decomposition below.
Define
\[
\sigma_{e,1,n}^2
=
\frac1n\sum_{i=1}^n
E[((A_{i,n}-\mu_{A,n})'e_{i,n})^2],
\]
\[
\sigma_{\mu,0,n}^2
=
\frac1n\sum_{i=1}^n
\operatorname{Var}(\mu_{A,n}'e_{i,n}).
\]
Then the generic diagonal target in Section~\ref{sec:setup} satisfies the following exact finite-$n$ identity at the level of the leading-score variance:
\[
V_n^\#
=
\sigma_{D,n}^{-4}
\left(
\sigma_{e,1,n}^2+\sigma_{\mu,0,n}^2
\right).
\]
Moreover, independence of iid assignment from the full latent array implies the exact orthogonal decomposition
\[
\sigma_{D\varepsilon,n}^2
=
\sigma_{e,1,n}^2
+
\operatorname{Var}\!\left(
n^{-1/2}\sum_{i=1}^n\mu_{A,n}'e_{i,n}
\right).
\]
Hence
\[
V_n
=
\sigma_{D,n}^{-4}
\left[
\sigma_{e,1,n}^2
+
\operatorname{Var}\!\left(
n^{-1/2}\sum_{i=1}^n\mu_{A,n}'e_{i,n}
\right)
\right],
\]
and therefore
\[
\boxed{
V_n^\#-V_n
=
-\sigma_{D,n}^{-4}\Delta_{A,n}
}
\]
with
\[
\Delta_{A,n}
=
\frac1n\sum_{i\neq j}
\operatorname{Cov}(\mu_{A,n}'e_{i,n},\mu_{A,n}'e_{j,n}).
\]
Equivalently,
\[
\Delta_{A,n}
=
\mu_{A,n}'(\Omega_{e,n}-\Omega_{e,0,n})\mu_{A,n}.
\]
These equalities are exact algebra for the deterministic variance sequences of the leading score; no claim is made here that $V_n$ equals $\operatorname{Var}\{\sqrt n(\widehat\beta-\beta)\}$ at finite $n$. That link is asymptotic and comes from the CHLS linearization \eqref{eq:chlsrep} together with the scale conditions used in the distributional results below.
If, in addition, the conditions of Lemma~\ref{lem:diag} hold, then the oracle diagonal estimator satisfies
\[
\boxed{
\sigma_{D,n}^{4}\{\widehat V_{n,o}^\#-V_n^\#\}=o_p(1).
}
\]
If also $\inf_n\sigma_{D,n}^2>0$, then
\[
\widehat V_{n,o}^\#-V_n^\#=o_p(1).
\]
\end{theorem}
\begin{remark}[From the score identity to OLS inference]\label{rem:score-vs-ols}
Theorem~\ref{thm:filter} is exact for the variance of the leading linearized score. Its implication for standardized OLS inference also requires the remainder in \eqref{eq:chlsrep} to be negligible at the relevant variance scale. If $V_n$ is bounded away from zero, the displayed $o_p(1)$ remainder is sufficient. If $V_n\to0$, one instead needs
\[
\sqrt n(\widehat\beta-\beta)
-
\sigma_{D,n}^{-2}n^{-1/2}\sum_i\psi_{i,n}
=
o_p(V_n^{1/2}),
\]
or a standardized CHLS conclusion stated directly. The scale-free distributional results below impose that standardized conclusion explicitly.
\end{remark}
The covariance identities in Theorem~\ref{thm:filter} are finite-$n$ consequences of the assignment structure and second moments. The additional conditions of Lemma~\ref{lem:diag} enter only for consistency of the oracle diagonal estimator $\widehat V_{n,o}^\#$.
\begin{remark}[Why only one covariance direction survives]\label{rem:filterorthogonality}
The identity in Theorem~\ref{thm:filter} does not require the latent array $\{e_{i,n}\}$ to be independent across $i$. What matters is that the iid assignment vector is independent of the full latent array. Writing $A_{i,n}=\mu_{A,n}+\widetilde A_{i,n}$, centering and iid assignment eliminate every cross-index covariance term containing $\widetilde A_{i,n}$. Dependence survives only through $\mu_{A,n}'e_{i,n}$, which is why the relevant off-diagonal covariance correction is the scalar
\[
\mu_{A,n}'(\Omega_{e,n}-\Omega_{e,0,n})\mu_{A,n}
\]
rather than the full latent covariance matrix.
\end{remark}
\begin{remark}[Why non-iid assignment requires a new filter]\label{rem:designboundary}
The iid assignment assumption matters. Under complete randomization with exactly $n_{1,n}$ treated units, put $p_n=n_{1,n}/n$. Then $\sum_i(D_{i,n}-p_n)=0$ identically, so the Bernoulli covariance filter need not survive as an $O(n^{-1})$ perturbation. For example, if $Y_i(0)=U$, $EU=0$, $E[U^2]>0$, and the treatment effect is zero, then
\[
n^{-1/2}\sum_i(D_{i,n}-p_n)U
\]
is identically zero under fixed-count assignment but has variance $p_n(1-p_n)E[U^2]$ under iid Bernoulli assignment with probability $p_n\in(0,1)$. Matched, covariate-adaptive, and cluster-level assignment likewise change the assignment covariance structure. Such designs require their own score filter rather than a generic correction. Regression adjustment can reduce residual heterogeneity when covariates explain common effect modifiers \citep{Lin2013}, but its adjusted score and surviving covariance projection must also be re-derived.
\end{remark}
Define the relevance index
\[
R_{A,n}=\frac{\Delta_{A,n}}{\sigma_{D\varepsilon,n}^2}.
\]
The exact oracle identity implies
\[
\frac{V_n^\#}{V_n}=1-R_{A,n}.
\]
Hence $R_{A,n}$ is the signed off-diagonal covariance contribution relative to the first-order sampling variance in the direction that survives random assignment. It is this dimensionless quantity---rather than dependence in the full outcome process---that governs dependence-ignorant bootstrap validity.
\begin{remark}[The phase function is well defined, and the failure is one-sided]\label{prop:rbound}
Suppose the covariance-filter conditions of Theorem~\ref{thm:filter} hold and $V_n>0$. Then $s_{\#,n}^2>0$ and therefore $R_{A,n}<1$ for every $n$ satisfying these conditions. Without a uniform lower bound on the diagonal score scale, however, a sequence $R_{A,n}$ may still approach one. If Assumption~\ref{ass:stable} also holds, then for all sufficiently large $n$,
\[
R_{A,n}\le1-\frac{c}{C}<1,
\]
where $c$ and $C$ are the constants in Assumption~\ref{ass:stable}.
\end{remark}
\begin{theorem}[Marginal indistinguishability and an information limit]\label{thm:marginalimpossibility}
Fix $p\in(0,1)$ and $\theta\in(0,1)$. Let
\[
D_i\stackrel{iid}{\sim}\operatorname{Bernoulli}(p),\qquad
u_i\stackrel{iid}{\sim}N(0,\sigma_u^2),\qquad
\eta_i\stackrel{iid}{\sim}N(0,1),
\]
with the three sequences mutually independent, and consider the following two binary-treatment designs
\[
Y_i^{\pm}(0)=u_i,
\qquad
\tau_i^{\pm}=\eta_i\pm\theta\eta_{i-1},
\qquad
Y_i^{\pm}=Y_i^{\pm}(0)+D_i\tau_i^{\pm}.
\]
Then:
\begin{enumerate}[label=(\roman*)]
\item the one-observation law of $(D_i,Y_i^+)$ is identical to that of $(D_i,Y_i^-)$;
\item the diagonal bootstrap target is identical under the two designs and equals
\[
V^{\#}
=
\frac{\sigma_u^2}{p(1-p)}+
\frac{1+\theta^2}{p};
\]
\item their aggregate treatment-effect covariance contributions satisfy
\[
\Delta_{\tau,n}^{\pm}
=
\pm 2\theta\left(1-\frac1n\right),
\]
so their leading-score sampling variances satisfy
\[
V_n^{\pm}
=
V^{\#}
\pm2\theta\left(1-\frac1n\right),
\qquad
V_n^+-V_n^-
=
4\theta\left(1-\frac1n\right).
\]
Under both designs the population treatment coefficient is zero. The leading score is stationary and one-dependent, and, in this intercept-only construction,
\[
\frac{\sqrt n\,\widehat\beta_n^{\pm}}{(V_n^{\pm})^{1/2}}\overset{d}{\to} N(0,1).
\]
\end{enumerate}
Call a variance estimator $\widehat V_n^{m}$ \emph{asymptotically marginal} on these two designs if there exists a deterministic sequence of functionals $\{\nu_n(\cdot)\}$ of the one-observation law such that, under either design,
\[
\widehat V_n^{m}
-
\nu_n\!\left(\mathcal L(D_i,Y_i)\right)
\overset{p}{\to}0.
\]
Allowing the functional to vary with $n$ makes the class broader than one with a fixed probability-limit functional.
No asymptotically marginal variance estimator can satisfy
\[
\widehat V_n^{m}/V_n^{+}\overset{p}{\to}1
\quad\text{under the $+$ design}
\qquad\text{and}\qquad
\widehat V_n^{m}/V_n^{-}\overset{p}{\to}1
\quad\text{under the $-$ design}.
\]
Consequently, no conditionally centered-Gaussian bootstrap whose conditional variance is asymptotically marginal can be first-order valid under both designs. More generally, if a bootstrap law is first-order valid and conditionally uniformly integrable in squared standardized units as in Theorem~\ref{thm:ui}, then its conditional second moment cannot be asymptotically marginal over this pair.
\end{theorem}
Theorem~\ref{thm:marginalimpossibility} concerns \emph{marginal-only} procedures; it does not rule out data-driven estimators that use joint dependence information. The two designs are indistinguishable in the one-observation law, but not in the joint law of the sample: their cross-unit covariance structures differ in sign. Joint-dependence information is therefore indispensable, but it need not arrive through a known ordering; a grouping, spatial metric, network, or other defensible dependence model can distinguish the designs. What no refinement of the one-observation resampling law can supply is the sign and magnitude of the surviving covariance.
The construction is intercept-only by design. If the analyst observes covariates that proxy the latent effect modifiers, the one-observation law of $(D_i,Y_i,W_i)$ need not remain common across the two designs, and the impossibility theorem does not apply to procedures that exploit those covariates. Observed effect modifiers therefore provide an important qualification. Treatment interactions or regression adjustment based on such covariates, in the spirit of \citet{Lin2013}, can reduce the residual heterogeneity whose covariance generates $\Delta_{\tau,n}$; they complement rather than replace a dependence correction for the remaining score.
The information limit also has a useful finite-population analogue. Classical Neyman variance calculations depend on treatment-effect heterogeneity that is not point identified from the two marginal potential-outcome distributions; \citet{AronowGreenLee2014} develop sharp variance bounds for that problem. Here the unidentified-from-marginals object is cross-unit treatment-effect covariance. Because it can have either sign, marginals alone do not tell the researcher whether to raise or lower the diagonal target: exact first-order correction requires joint-dependence information. This does not preclude conservative bounds under additional restrictions.
\begin{theorem}[Covariance quotient and scale-free minimal dependence correction]\label{thm:minimality}
Let $\mathbb S^L$ denote the ambient space of symmetric $L\times L$ matrices and define
\[
H_n=\Omega_{e,n}-\Omega_{e,0,n},
\qquad
\mathcal L_n(B)=\mu_{A,n}'B\mu_{A,n},
\]
with null space
\[
\mathcal N_n=\{B\in\mathbb S^L:\mathcal L_n(B)=0\}.
\]
Under the exact leading-score covariance-filter conditions of Theorem~\ref{thm:filter}, define the first-order variance correction
\[
c_n:=V_n-V_n^\#
=
\sigma_{D,n}^{-4}\mathcal L_n(H_n).
\]
Hence:
\begin{enumerate}[label=(\roman*)]
\item holding the assignment law, $\mu_{A,n}$, $\sigma_{D,n}^2$, and the diagonal bootstrap target $V_n^\#$ fixed, two \emph{model-compatible off-diagonal covariance operators} $H_n$ and $\widetilde H_n$ generate the same deterministic first-order variance target $V_n$ whenever
\[
H_n-\widetilde H_n\in\mathcal N_n;
\]
\item if $\mu_{A,n}\neq0$, then the ambient quotient $\mathbb S^L/\mathcal N_n$ is one-dimensional; model-compatible covariance operators inherit the equivalence relation by intersection with its cosets. If $\mu_{A,n}=0$, every model-compatible off-diagonal covariance direction is irrelevant for the first-order variance target $V_n$;
\item suppose $V_n>0$ and $V_n^\#>0$ for all sufficiently large $n$,
\[
\frac{\widetilde V_n^\#-V_n^\#}{V_n}\overset{p}{\to}0,
\]
and define
\[
\widehat V_n^{\,*}
=
\widetilde V_n^\#+\widehat c_n,
\qquad
P(\widehat V_n^{\,*}>0)\to1.
\]
On the event $\{\widehat V_n^{\,*}>0\}$, let $T_n^*$ be conditionally centered Gaussian with conditional variance $\widehat V_n^{\,*}$; on its complement, define $T_n^*$ arbitrarily. If
\[
\frac{\sqrt n(\widehat\beta-\beta_0)}{V_n^{1/2}}\overset{d}{\to} N(0,1),
\]
then
\[
\sup_x\left|
P^*\{T_n^*\le x\}
-
P\{\sqrt n(\widehat\beta-\beta_0)\le x\}
\right|\overset{p}{\to}0
\]
if and only if
\[
\boxed{
\frac{\widehat c_n-c_n}{V_n}
\overset{p}{\to}0.
}
\]
\end{enumerate}
In particular, if
\[
\widehat c_n
=
\sigma_{D,n}^{-4}\mathcal L_n(\widehat H_n),
\]
then first-order validity requires and, within the stated centered-Gaussian class, is implied by
\[
\boxed{
\frac{
\sigma_{D,n}^{-4}\mathcal L_n(\widehat H_n-H_n)
}{V_n}
\overset{p}{\to}0.
}
\]
Thus full-matrix covariance consistency is generally stronger than necessary. Errors in directions belonging to $\mathcal N_n$ are exactly irrelevant, while errors outside $\mathcal N_n$ need only be negligible at the first-order sampling-variance scale.
\end{theorem}
\begin{remark}[Representation invariance of the filtered covariance]\label{rem:representation}
The latent factorization $\varepsilon_i=\lambda(D_i)'e_i$ is a convenient representation, not an identified coordinate system. The filtered covariance correction is invariant to nonsingular changes of latent coordinates. In particular, for any nonsingular $L\times L$ matrix $C_n$, define
\[
e_{i,n}^{\star}=C_ne_{i,n},
\qquad
\lambda_n^{\star}(d)=C_n^{-\prime}\lambda_n(d).
\]
Then $\varepsilon_{i,n}=\lambda_n^{\star}(D_{i,n})'e_{i,n}^{\star}$,
\[
\mu_{A,n}^{\star}=C_n^{-\prime}\mu_{A,n},
\qquad
H_n^{\star}=C_nH_nC_n',
\]
and therefore
\[
(\mu_{A,n}^{\star})'H_n^{\star}\mu_{A,n}^{\star}
=
\mu_{A,n}'H_n\mu_{A,n}.
\]
The scalar correction $c_n$ and the one-dimensional quotient are therefore invariant under nonsingular reparameterization. The ambient null space itself changes coordinates, but the induced equivalence classes are mapped isomorphically. What matters is the scalar filtered covariance, not a particular choice of latent basis.
\end{remark}
\begin{lemma}[The covariance quotient is attained by model-compatible processes]\label{lem:sharpquotient}
Consider the binary specialization of Section~\ref{sec:binary} with fixed $p\in(0,1)$ and $q=p(1-p)$. Let $Y_i(0)=u_i$, where $\{u_i\}$ is iid with finite variance, and let
\[
\tau_i=\eta_i+\theta\eta_{i-1},\qquad \theta\in\mathbb R,
\]
with $\{\eta_i\}$ iid, mean zero, variance one, and independent of $\{u_i\}$ and assignment. Then the latent process is stationary and satisfies the maintained second-moment and assignment conditions, and
\[
\mathcal L_n(H_n)=q^2\Delta_{\tau,n}=2q^2\theta\left(1-\frac1n\right).
\]
Hence, for every $n\ge2$, varying $\theta$ moves the model-compatible covariance class through a continuum of distinct cosets of $\mathcal N_n$, including both signs. The one-dimensional quotient in Theorem~\ref{thm:minimality}(ii) is therefore attained rather than merely an ambient algebraic possibility.
\end{lemma}
\begin{remark}[Geometry versus asymptotic relevance]\label{rem:minimality}
Applying HAC or dependent multipliers to an asymptotically linear score is standard. Theorem~\ref{thm:minimality} identifies the additional content supplied by random assignment: before any HAC device is chosen, random assignment quotients out, in the ambient symmetric-matrix representation, all off-diagonal covariance directions in $\mathcal N_n$. When $\mu_{A,n}\neq0$ this quotient is algebraically one-dimensional, but its statistical relevance can still vanish along a triangular array. In particular,
\[
\frac{c_n}{V_n}\to0
\]
makes the iid correction asymptotically unnecessary even if $\mu_{A,n}\neq0$ for every $n$. Conversely, $\|\mu_{A,n}\|\to0$ alone does not imply $c_n/V_n\to0$ unless the latent covariance operator is controlled. The scale-free object is the surviving covariance contribution relative to $V_n$, not the norm of $\mu_{A,n}$ by itself.
\end{remark}
\begin{theorem}[Hahn--Liao moment bridge]\label{thm:ui}
Let $Z_n^*=T_n^*/V_n^{1/2}$, where $V_n>0$. Suppose
\[
\sup_x\left|P^*\{Z_n^*\le x\}-\Phi(x)\right|\overset{p}{\to}0
\]
and $\{(Z_n^*)^2\}$ is conditionally uniformly integrable in probability, meaning that for every $\varepsilon>0$,
\[
\lim_{M\to\infty}\limsup_{n\to\infty}
P\!\left(
E^*\!\left[(Z_n^*)^2\mathbf 1\{|Z_n^*|>M\}\right]>\varepsilon
\right)=0.
\]
Then
\[
E^*[(Z_n^*)^2]\overset{p}{\to}1,\qquad E^*[Z_n^*]\overset{p}{\to}0,
\qquad \frac{\operatorname{Var}^*(T_n^*)}{V_n}\overset{p}{\to}1.
\]
The second conclusion is needed because $T_n^*$ is not assumed conditionally centered in this theorem. Conditional uniform integrability of $(Z_n^*)^2$ implies conditional uniform integrability of $Z_n^*$; together with conditional weak convergence this yields $E^*[Z_n^*]\to_p0$. The variance conclusion therefore follows from the first two moments rather than from the second moment alone. If, in addition,
\[
E^*[(T_n^*)^2]=\widetilde V_n^\#+\widehat c_n,
\qquad
\frac{\widetilde V_n^\#-V_n^\#}{V_n}\overset{p}{\to}0,
\]
then weak validity necessarily implies
\[
\boxed{\frac{\widehat c_n-c_n}{V_n}\overset{p}{\to}0},
\qquad
c_n:=V_n-V_n^\#.
\]
Under the covariance-filter conditions of Theorem~\ref{thm:filter}, this correction is
\[
c_n=\sigma_{D,n}^{-4}\mathcal L_n(H_n).
\]
Thus, under moment control, an arbitrary bootstrap law must recover the same filtered second moment identified by Theorem~\ref{thm:minimality}. The converse is not asserted outside the Gaussian class.
\end{theorem}
A primitive sufficient condition for the uniform-integrability requirement is
\[
E^*|Z_n^*|^{2+\eta}=O_p(1)
\quad\text{for some }\eta>0.
\]
This condition is automatic for the Gaussian projected-dependence bootstrap whenever its conditional variance ratio is tight. Online Appendix Section~\ref{app:ui} also verifies it under the existing fourth-moment conditions for the regularized pairs bootstrap and, after conditioning on nondegenerate resamples, for the binary-treatment pairs bootstrap. It is \emph{not} automatic for unrestricted ordinary pairs OLS with controls: rare nearly singular resampled Hessians can preserve weak validity while destroying bootstrap moments. This is the failure mode emphasized by Hahn--Liao.
\begin{assumption}[Scale-free design and diagonal-score consistency]\label{ass:scalefree}
Under $H_0:\beta=\beta_0$, suppose $\sigma_{D,n}^2>0$ and $s_{\#,n}^2>0$. Restoring the row index in the design scale, assume
\[
\frac{\widehat Q_n}{\sigma_{D,n}^2}\overset{p}{\to}1,
\qquad
\frac{n^{-1}\sum_{i=1}^n\{\widetilde\psi_i(\beta_0)-\psi_i\}^2}{s_{\#,n}^2}=o_p(1),
\qquad
\frac{n^{-1}\sum_{i=1}^n\psi_{i,n}^2-s_{\#,n}^2}{s_{\#,n}^2}=o_p(1).
\]
\end{assumption}
Assumption~\ref{ass:scalefree} collects the three relative requirements needed when the treatment or diagonal-score scale can shrink: treatment-design concentration, feasible-score replacement, and oracle diagonal-score concentration. Under Assumption~\ref{ass:stable}, Lemma~\ref{lem:design} and Lemmas~\ref{lem:restricted}--\ref{lem:diag} imply these requirements from the primitive conditions used below. If $s_{\#,n}$ can shrink, the absolute score-replacement bound
\[
n^{-1}\sum_i
\{\widetilde\psi_i(\beta_0)-\psi_i\}^2=o_p(1)
\]
is not enough by itself; a rate relative to $s_{\#,n}^2$ is genuinely needed. Primitive sufficient rates for that case are given in Online Appendix Section~\ref{app:proofs-main}.
The relative formulation matters in triangular arrays with diminishing treatment variation, for example Bernoulli designs with $p_n\to0$ or $p_n\to1$, and in sequences where the diagonal score scale itself becomes small. In sparse-treatment designs $\sigma_{D,n}^2=p_n(1-p_n)$ vanishes and the OLS variance need not shrink---it can instead diverge. This is not weak identification in the instrumental-variables sense. Rather, an absolute $o_p(1)$ error may still be large relative to a moving design or score scale, which is why the scale-free conditions are stated in relative terms.
\begin{theorem}[Feasible variance target]\label{thm:feasible}
Under $H_0:\beta=\beta_0$, suppose the exact covariance-filter conditions of Theorem~\ref{thm:filter} hold. No lower bound on $V_n$ is needed for this variance-target result.
\begin{enumerate}[label=(\alph*)]
\item If Assumption~\ref{ass:scalefree} holds, then
\[
\frac{\widehat Q_n}{\sigma_{D,n}^2}\overset{p}{\to}1,
\qquad
\frac{\widetilde V_n^\#(\beta_0)}{V_n^\#}\overset{p}{\to}1.
\]
\item Suppose, in addition, the conditions of Lemmas~\ref{lem:diag} and \ref{lem:restricted} and Assumption~\ref{ass:stable} hold. Then Lemma~\ref{lem:design} and Lemmas~\ref{lem:diag} and \ref{lem:restricted} imply all three clauses of Assumption~\ref{ass:scalefree}, and
\[
\widetilde V_n^\#(\beta_0)-V_n^\#=o_p(1).
\]
Consequently,
\[
\widetilde V_n^\#(\beta_0)-V_n
=
-\sigma_{D,n}^{-4}\Delta_{A,n}+o_p(1).
\]
\end{enumerate}
\end{theorem}
\begin{theorem}[Bootstrap validity frontier and size phase diagram]\label{thm:phase}
Suppose the covariance-filter conditions of Theorem~\ref{thm:filter} hold, $V_n>0$ and $V_n^\#>0$ for all sufficiently large $n$, and
\[
\frac{\widetilde V_n^\#(\beta_0)}{V_n^\#}\overset{p}{\to}1,
\]
as guaranteed by Theorem~\ref{thm:feasible} under either its scale-free or stable-scale route, and suppose the CHLS standardized sampling law satisfies
\[
\frac{\sqrt n(\widehat\beta-\beta_0)}{V_n^{1/2}}\overset{d}{\to} N(0,1)
\]
under $H_0$. If
\[
R_{A,n}\to r<1,
\]
then
\[
\frac{\sqrt n(\widehat\beta-\beta_0)}
{\{\widetilde V_n^\#(\beta_0)\}^{1/2}}
\overset{d}{\to}
N\!\left(0,\frac1{1-r}\right).
\]
Consequently, the two-sided nominal level-$\alpha$ test based on the iid multiplier score-bootstrap standard error has limiting rejection probability
\[
\boxed{
\pi_\alpha(r)
=
2\Phi\!\left(
-z_{1-\alpha/2}\sqrt{1-r}
\right).
}
\]
Thus $r=0$ is the first-order validity frontier for the iid multiplier score-bootstrap standard error, $r>0$ yields anti-conservative inference, and $r<0$ yields conservative inference. Remark~\ref{prop:rbound} gives $R_{A,n}<1$ at each finite $n$ satisfying the stated conditions, but sequences may approach one. The maintained condition $r<1$ excludes that scale-free boundary, where $V_n^\#/V_n\to0$; Assumption~\ref{ass:stable} ensures uniform separation from one.
\end{theorem}
\begin{corollary}[Oracle equal-size benchmark and feasible local power]\label{cor:localpower}
Suppose the covariance-filter conditions of Theorem~\ref{thm:filter} hold, $V_n\to V\in(0,\infty)$, $V_n^\#>0$ for all sufficiently large $n$, and $R_{A,n}\to r<1$. Along the local alternatives
\[
\beta_n=\beta_0+\frac{c}{\sqrt n},
\]
suppose
\[
\frac{\sqrt n(\widehat\beta-\beta_n)}{V_n^{1/2}}\overset{d}{\to} N(0,1),
\qquad
\frac{\widetilde V_n^\#(\beta_0)}{V_n^\#}\overset{p}{\to}1.
\]
Let $\delta=c/V^{1/2}$ and $z=z_{1-\alpha/2}$. Then the unadjusted iid-variance test has limiting local rejection probability
\[
\Phi\!\left(-z\sqrt{1-r}-\delta\right)
+1-\Phi\!\left(z\sqrt{1-r}-\delta\right).
\]
Any first-order valid projected test whose variance estimator remains relatively consistent for $V_n$ along the same local alternatives has limiting local power
\[
\Phi(-z-\delta)+1-\Phi(z-\delta).
\]
As a theoretical equal-size benchmark, if the iid statistic is given the oracle critical value $z/\sqrt{1-r}$ so that its asymptotic null rejection probability is exactly $\alpha$, its local-power curve is also
\[
\Phi(-z-\delta)+1-\Phi(z-\delta).
\]
Thus the raw first-order power difference generated by the wrong iid variance is a calibration effect. The oracle iid statistic can attain the same first-order curve only by using the unknown relevance-index calibration $r$. A first-order valid projected procedure attains that curve feasibly by estimating the surviving covariance contribution. The correction restores calibration and feasibility; it does not create a first-order efficiency gain in this Gaussian local experiment.
\end{corollary}
Corollary~\ref{cor:localpower} separates calibration from efficiency. Positive surviving covariance can make the raw iid test appear more powerful only because it over-rejects under the null, while negative covariance can make it appear less powerful because it under-rejects. Once the two procedures are calibrated to the same asymptotic size, the first-order local-power difference disappears. Feasibility is what remains: the equal-size iid benchmark requires the unknown limiting relevance index, whereas the projected procedure estimates the covariance contribution that determines it. The projected correction should therefore be viewed as a feasible size correction, not as a source of first-order power gains.
\begin{figure}[!htbp]
\centering
\includegraphics[width=.72\textwidth]{figure_local_power.pdf}
\caption{Asymptotic local rejection probabilities as a function of the standardized local shift $|\delta|$. The raw iid curves use the limiting relevance indices of the positive and negative MA(1) benchmark designs. The projected curve coincides exactly with the iid curve after equal-size calibration, so the raw power differences are calibration effects rather than first-order efficiency gains.}
\label{fig:localpower}
\end{figure}
\begin{remark}[Separation from CHLS]\label{rem:chlsnovelty}
CHLS characterize the first-order OLS variance represented here by $V_n$. We take their sampling linearization as given; the corresponding multivariate expansion is developed in their setup and formalized in their main theory. The present paper characterizes the bootstrap target $V_n^\#$ induced by dependence-ignorant weighting and compares it with $V_n$. The covariance-filter identity, validity frontier, minimality characterization, and projected repair are therefore bootstrap results rather than alternative derivations of the CHLS variance formula.
\end{remark}
\section{Binary treatment and heterogeneous effects}\label{sec:binary}
For the cleanest potential-outcomes specialization, take $W_{i,n}=1$ and let
\[
D_{i,n}\stackrel{iid}{\sim}\operatorname{Bernoulli}(p_n),
\qquad
0<p_n<1,
\]
independently of the full potential-outcome array. Write
\[
Y_{i,n}
=
Y_{i,n}(0)+D_{i,n}\tau_{i,n},
\qquad
\tau_{i,n}=Y_{i,n}(1)-Y_{i,n}(0).
\]
In this binary representation, constant treatment effects recover the strong-exogeneity benchmark; heterogeneous $\tau_{i,n}$ is the economically important route through which assignment-dependent regression errors arise.
Assume that, within each row,
\[
E Y_{i,n}(0)=\alpha_n,
\qquad
E\tau_{i,n}=\beta_n
\]
for every $i$. Then the population coefficient in the regression of $Y_{i,n}$ on an intercept and $D_{i,n}$ is exactly $\beta_n$. Indeed, with
\[
q_n=p_n(1-p_n),
\]
independence of assignment and potential outcomes gives
\[
E[(D_{i,n}-p_n)Y_{i,n}]
=
q_n\beta_n.
\]
The associated regression error is
\[
\varepsilon_{i,n}
=
Y_{i,n}(0)-\alpha_n
+
D_{i,n}(\tau_{i,n}-\beta_n).
\]
Thus the CHLS representation can be written as
\[
e_{i,n}
=
\begin{pmatrix}
Y_{i,n}(0)-\alpha_n\\
\tau_{i,n}-\beta_n
\end{pmatrix},
\qquad
\lambda_n(D_{i,n})
=
\begin{pmatrix}
1\\D_{i,n}
\end{pmatrix}.
\]
Since
\[
A_{i,n}
=
(D_{i,n}-p_n)
\begin{pmatrix}
1\\D_{i,n}
\end{pmatrix},
\]
we have exactly
\[
\mu_{A,n}
=
\begin{pmatrix}
0\\q_n
\end{pmatrix},
\qquad
\mu_{A,n}'e_{i,n}
=
q_n(\tau_{i,n}-\beta_n).
\]
\begin{corollary}[Treatment-effect covariance discrepancy]\label{cor:te}
Under the exact covariance-filter conditions of Theorem~\ref{thm:filter} and the binary-treatment specialization above, define
\[
\Delta_{\tau,n}
=
\frac1n\sum_{i\ne j}
\operatorname{Cov}(\tau_{i,n},\tau_{j,n}).
\]
Then, exactly for every $n$ satisfying the stated conditions at the level of the leading-score variance sequences,
\[
\Delta_{A,n}=q_n^2\Delta_{\tau,n}
=\sigma_{D,n}^4\Delta_{\tau,n},
\]
and therefore
\[
\boxed{
V_n^\#-V_n=-\Delta_{\tau,n}.
}
\]
The displayed equality is not an exact finite-sample identity for $\operatorname{Var}\{\sqrt n(\widehat\beta-\beta)\}$; its implication for OLS inference uses the CHLS linearization.
Whenever $V_n>0$,
\[
\boxed{
R_{A,n}=\frac{\Delta_{\tau,n}}{V_n}.
}
\]
\end{corollary}
The sign classification concerns a non-negligible asymptotic contribution. Under the standardized sampling central limit theorem (CLT) in Theorem~\ref{thm:phase}, for example, if
\[
V_n\to V\in(0,\infty),
\qquad
\Delta_{\tau,n}\to\Delta_\tau,
\qquad
V-\Delta_\tau>0,
\]
then $\Delta_\tau>0$ implies anti-conservative dependence-ignorant inference, $\Delta_\tau=0$ implies first-order correct inference, and $\Delta_\tau<0$ implies conservative inference. More generally the relevant object is the relative discrepancy $\Delta_{\tau,n}/V_n$. Treatment-effect independence is sufficient but not necessary for first-order validity.
\begin{corollary}[Binary-treatment size distortion]\label{cor:size}
Under the conditions of Theorem~\ref{thm:phase}, in the binary-treatment specialization suppose
\[
\frac{\Delta_{\tau,n}}{V_n}\to\rho<1.
\]
Finite-$n$ strictness from Remark~\ref{prop:rbound} does not rule out the boundary limit $\rho=1$; Assumption~\ref{ass:stable} does. Under the maintained $\rho<1$, the two-sided nominal level-$\alpha$ iid multiplier score-bootstrap test has limiting rejection probability
\[
2\Phi\!\left(
-z_{1-\alpha/2}\sqrt{1-\rho}
\right).
\]
Thus positive aggregate treatment-effect covariance is anti-conservative, negative aggregate covariance is conservative, and asymptotically negligible aggregate covariance is first-order correct.
\end{corollary}
\subsection{Relation to design-based clustering and practical guidance}\label{sec:practice}
The binary identity is closely related to the treatment-effect-heterogeneity channel emphasized by \citet{AbadieEtAl2023} and to the randomization/clustering perspective of \citet{BarriosEtAl2012}, but it answers a different inferential question. Their design-based analysis asks when clustered variance estimation is warranted by the sampling and assignment design. Here assignment remains iid at the unit level and the object of interest is whether an independent resampling scheme reproduces the covariance that survives random assignment. The exact identity adds two features that are especially useful for practice: the correction is signed, and only the treatment-effect covariance contribution relevant to the scalar randomized-regressor score must be recovered.
Three practical implications follow for the randomized-regressor coefficient under the maintained iid unit-level assignment. First, outcome dependence alone is not a sufficient reason to cluster this coefficient: under strong exogeneity, Theorem~\ref{thm:strong} gives $V_n^\#=V_n$ even with broad cross-unit dependence in outcomes and controls. Second, when heterogeneous treatment effects co-move, the relevant dependence geometry is the geometry of the effect modifiers---shared teachers, local labor markets, institutions, spatial exposure, or another defensible grouping---rather than automatically the level at which treatment was assigned. Third, the correction is not a one-way conservative adjustment: negative treatment-effect covariance makes dependence-ignorant inference conservative, so a valid correction can narrow rather than widen confidence intervals.
In practice, we recommend reporting the diagonal score variance together with projected variances under substantively defensible, preferably prespecified, dependence specifications. If several are plausible, report all as sensitivity analysis rather than select among them ex post. Their gap is the primary object; the ratio $\widehat r_n$ is a descriptive normalization, not a sign pretest. A useful empirical placebo, when pre-treatment outcomes are available, is to apply the same projected-score calculation to a pre-treatment outcome: under the maintained mechanism there is then no treatment-effect covariance to survive random assignment, even if the outcome itself is strongly correlated across groups. This is a diagnostic implication, not a substitute for the iid-assignment requirement stated in the Introduction.
For a covariance-stationary treatment-effect process with autocovariance
\[
\varrho_{\tau,n}(h)=\operatorname{Cov}(\tau_{i,n},\tau_{i-h,n})
\]
that does not depend on $i$ within row $n$,
\[
\Delta_{\tau,n}
=
2\sum_{h=1}^{n-1}
\left(1-\frac hn\right)\varrho_{\tau,n}(h).
\]
If $\varrho_{\tau,n}(h)=\varrho_\tau(h)$ and
$\sum_{h\ge1}|\varrho_\tau(h)|<\infty$, then
\[
\Delta_{\tau,n}
\to
2\sum_{h\ge1}\varrho_\tau(h).
\]
\section{Implication for ordinary pairs bootstrap}\label{sec:pairs}
Let $\widehat\gamma$ and $\widehat\varepsilon_i=Y_i-D_i\widehat\beta-W_i'\widehat\gamma$ denote the unrestricted OLS nuisance estimate and residuals. Let
\[
\widehat v=M_WD,\qquad \widehat Q_D=n^{-1}\sum_i\widehat v_i^{\,2}=n^{-1}D'M_WD.
\]
The ordinary pairs bootstrap draws $n$ rows with replacement. Equivalently,
\[
(N_1^*,\ldots,N_n^*)\sim\operatorname{Multinomial}(n;n^{-1},\ldots,n^{-1})
\]
counts how often each observation is drawn. Its conditional moments are $E^*N_i^*=1$ and $\operatorname{Var}^*(N_i^*)=1-n^{-1}$, while $\operatorname{Cov}^*(N_i^*,N_j^*)=-n^{-1}$ for $i\ne j$.
Ordinary pairs resampling does not evade the covariance filter. Classical analyses of regression resampling include \citet{Freedman1981} and \citet{GoncalvesWhite2005}. Because row resampling perturbs both the score and the design matrix, its linearization requires leverage and empirical-Lindeberg regularity beyond the baseline CHLS conditions. Under those conditions, the first-order pairs statistic is
\[
\frac1{\sqrt n}\sum_{i=1}^n(N_i^*-1)\widehat\phi_i,
\qquad
\widehat\phi_i
=
\widehat Q_D^{-1}\widehat v_i\widehat\varepsilon_i,
\]
and its conditional variance is exactly the usual HC0 heteroskedasticity-robust coefficient variance,
\[
\widehat V_{\mathrm{HC0},\beta}
=
\widehat Q_D^{-2}\frac1n\sum_{i=1}^n
\widehat v_i^2\widehat\varepsilon_i^2.
\]
The equality is exact: the multinomial cross-terms contribute a term proportional to
\[
-\left(\sum_i\widehat v_i\widehat\varepsilon_i\right)^2,
\]
which vanishes identically because $\sum_i\widehat v_i\widehat\varepsilon_i=D'M_W\widehat\varepsilon=0$ by the OLS first-order conditions.
Theorem~\ref{thm:pairs} in the Online Appendix, under Assumption~\ref{ass:pairs}, proves the corresponding conditional central limit theorem and shows
\[
\widehat V_{\mathrm{HC0},\beta}
=
V_n^\#+o_p(1).
\]
Hence ordinary pairs resampling has the same first-order diagonal target as the iid multiplier score bootstrap. Analytical variance estimation is not inherently invalid here. HC0 and iid row resampling fail because both are diagonal and therefore discard surviving cross-unit score covariance. A dependence-aware HAC, cluster, or spatial variance estimator for the relevant residualized score can instead target $V_n$. Combining this result with Theorem~\ref{thm:filter},
\[
\widehat V_{\mathrm{HC0},\beta}-V_n
=
-\sigma_{D,n}^{-4}\Delta_{A,n}+o_p(1).
\]
For binary randomized treatment this reduces to
\[
\widehat V_{\mathrm{HC0},\beta}-V_n
=
-\Delta_{\tau,n}+o_p(1).
\]
Thus the pairs bootstrap inherits the same variance-target phase. Online Appendix Corollary~\ref{cor:pairsphase} formalizes the corresponding size result: under the sampling-law and relevance-index conditions of Theorem~\ref{thm:phase} with $R_{A,n}\to r<1$, a two-sided test using ordinary-pairs quantiles of $\sqrt n(\widehat\beta^*-\widehat\beta)$ has the same limiting rejection probability $\pi_\alpha(r)$ as the iid multiplier score-bootstrap test. It is therefore anti-conservative for $r>0$, first-order correct for $r=0$, and conservative for $r<0$. Without those sampling-law and relative-gap conditions, the displays above are variance-target statements rather than a rejection-probability theorem.
This first-order statement does not by itself imply convergence of the conditional variance of the nonlinear pairs-OLS estimator; that requires conditional uniform integrability or an equivalent tail condition. Online Appendix Section~\ref{app:pairs-full} treats this Hahn--Liao distinction, a regularized pairs construction with controls, and an unregularized binary-treatment procedure conditional on nondegenerate resamples whose bootstrap variance satisfies
\[
V_{\mathrm{pairs},H,n}^*-V_n
=
-\Delta_{\tau,n}+o_p(1).
\]
These results concern iid row resampling and do not cover fixed-treatment-count complete randomization, matched-pair, covariate-adaptive, or cluster-level assignment.
\section{Dependence correction on the projected score}\label{sec:repair}
The covariance-filter and minimality theorems suggest correcting only the dependence that survives random assignment. Theorem~\ref{thm:minimality} shows that, given a consistent diagonal component, the surviving covariance contribution measured relative to $V_n$ is necessary and sufficient for first-order conditionally centered-Gaussian bootstrap validity; the relevant first-order object is therefore the scalar randomized-regressor score, not the full vector of outcomes and controls.
Let $\widetilde\psi_i(\beta_0)$ be the feasible null-imposed score. Let
\[
\xi_n^*=(\xi_{1,n}^*,\ldots,\xi_{n,n}^*)'
\]
be conditionally mean-zero multipliers with
\[
E^*(\xi_n^*\xi_n^{*'})=K_n=(\kappa_{ij,n})_{i,j\le n},
\]
where $K_n$ is symmetric positive semidefinite. For Gaussian multipliers this condition is exactly what guarantees the existence of the stated conditional covariance law.
Define the \emph{projected-dependence multiplier implementation}
\[
T_{\mathrm{PD},n}^*(\beta_0)
=
\widehat Q_n^{-1}\frac1{\sqrt n}
\sum_{i=1}^n\xi_{i,n}^*\widetilde\psi_i(\beta_0),
\]
with conditional variance
\[
\widetilde V_{\mathrm{PD},n}(\beta_0)
=
\widehat Q_n^{-2}\frac1n
\sum_{i,j=1}^n
\kappa_{ij,n}\widetilde\psi_i(\beta_0)\widetilde\psi_j(\beta_0).
\]
The iid multiplier score bootstrap is the special case $K_n=I_n$. Block, cluster, spatial, or dependent-multiplier choices of $K_n$ alter only the covariance of the scalar score multipliers; they do not resample the full data vector or recompute OLS. The projection is a dimension reduction, not a model-free dependence estimator: $K_n$ must still be chosen so that the resulting scalar quadratic form consistently targets the first-order score variance. When Gaussian multipliers are used, the conditional uniform-integrability requirement behind the Hahn--Liao moment issue is automatic as soon as the projected variance ratio is tight.
\begin{theorem}[Projected-dependence bootstrap principle]\label{thm:pd}
Under $H_0$, suppose $V_n>0$ for all sufficiently large $n$,
\[
P\{\widetilde V_{\mathrm{PD},n}(\beta_0)>0\}\to1,
\]
and
\[
\frac{\sqrt n(\widehat\beta-\beta_0)}{V_n^{1/2}}\overset{d}{\to} N(0,1),
\qquad
\frac{\widetilde V_{\mathrm{PD},n}(\beta_0)}{V_n}\overset{p}{\to}1,
\]
and, on the event $\{\widetilde V_{\mathrm{PD},n}(\beta_0)>0\}$, conditionally on the data,
\[
\sup_x\left|
P^*\!\left(
\frac{T_{\mathrm{PD},n}^*(\beta_0)}
{\{\widetilde V_{\mathrm{PD},n}(\beta_0)\}^{1/2}}\le x
\right)-\Phi(x)
\right|\overset{p}{\to}0.
\]
On the complementary event the standardized statistic may be assigned any fixed bookkeeping value.
Then
\[
\sup_x\left|
P^*\{T_{\mathrm{PD},n}^*(\beta_0)\le x\}
-
P\{\sqrt n(\widehat\beta-\beta_0)\le x\}
\right|\overset{p}{\to}0.
\]
Thus, once the standardized conditional bootstrap law is asymptotically Gaussian, first-order validity can be achieved by consistently reproducing the covariance of the projected score only; reproducing the full joint dependence of $(Y_i,W_i)$ is unnecessary. Necessity within the centered-Gaussian class is supplied separately by Theorem~\ref{thm:minimality}.
\end{theorem}
\paragraph{Recentering the null-imposed covariance score.}
One may replace $\widetilde\psi_i(\beta_0)$ inside the covariance estimator by
\[
\widetilde\psi_i^{c}(\beta_0)
=
\widetilde\psi_i(\beta_0)-\overline{\widetilde\psi}(\beta_0),
\qquad
\overline{\widetilde\psi}(\beta_0)=n^{-1}\sum_{i=1}^n\widetilde\psi_i(\beta_0).
\]
This is harmless only under an appropriate bandwidth regime. If $K_n$ is the covariance kernel, $\overline{\widetilde\psi}=O_p(n^{-1/2})$, $\mathbf 1'K_n\widetilde\psi=O_p(m_n\sqrt n)$, and $\mathbf 1'K_n\mathbf 1=O(nm_n)$, then the centered and uncentered quadratic forms differ by $O_p(m_n/n)$. Hence they are first-order equivalent under the small-bandwidth routes when $m_n=o(n)$ and the target scale is bounded away from zero. At fixed bandwidth, $m_n/n\to b>0$, the centering term is generally first order and can change the Brownian-motion reference functional to a bridge-type functional. The fixed-bandwidth theorem below therefore uses the nonrecentered null-imposed score; a recentered fixed-bandwidth implementation requires its own joint limit. The exact algebra and the stated sufficient orders are recorded in Online Appendix Remark~\ref{rem:recenter}.
Primitive implementations are given in the Online Appendix. Theorem~\ref{thm:pdmixing} treats a stationary strong-mixing projected score, a positive-semidefinite lag kernel, and kernel denominator $m_n=o(n^{1/3})$; its proof separates kernel bias, quadratic-form concentration, feasible-score replacement, and the finite-sample/long-run-variance edge correction. Theorem~\ref{thm:pdcumulant} gives a complementary fourth-order-cumulant route with $m_n=o(n)$, while Proposition~\ref{prop:pdcluster} gives an independent-cluster route under a cluster-sum fourth-moment diffuseness condition. These are sufficient implementations of Theorem~\ref{thm:pd}; their role is to show how the one-dimensional covariance target can be estimated under familiar dependence structures. A useful feature of the cumulant route is that the nuisance-estimation error is low rank: exploiting that structure sharpens feasible replacement from a crude bandwidth-times-estimation-error bound to $O_p(n^{-1/2})+O_p(m_n/n)$, which is what permits $m_n=o(n)$. Online Appendix Proposition~\ref{prop:pduniform} records the corresponding uniform-size transfer when the sampling CLT, positivity, and projected-variance consistency hold uniformly over a class of data-generating processes. The HAC theorems use deterministic $m_n$; plug-in bandwidths require an additional validity argument.
\subsection{Fixed-bandwidth calibration for ordered data}\label{sec:fixedb}
The small-bandwidth theory above uses a kernel bandwidth that is negligible relative to $n$ and therefore a standard-normal reference law. When the kernel bandwidth is a non-negligible fraction of the sample size, that approximation can be inaccurate even when the projected covariance target is correct. Fixed-bandwidth asymptotics instead let the kernel denominator $m_n$ satisfy $m_n/n\to b\in(0,1]$ and let the reference law depend on the kernel and $b$ \citep{KieferVogelsang2005,SunPhillipsJin2008}.
Let $K_{n,b}$ have entries $k((i-j)/m_n)$, where $m_n/n\to b$, and let $\widetilde V_{\mathrm{PD},n}^{(b)}$ denote the corresponding projected quadratic form. In the Bartlett implementation used below, $\ell$ denotes the largest included positive lag and the weights are $1-h/(\ell+1)$, so $m_n=\ell+1$ and the reported finite-sample bandwidth fraction is $b_n=(\ell+1)/n$. For a continuous positive-semidefinite kernel, define the fixed-bandwidth Brownian functional
\[
Q_{k,b}(W)=\int_0^1\!\int_0^1 k\!\left(\frac{r-s}{b}\right)dW(r)dW(s),
\]
interpreted as the standard $L^2$ limit of the associated quadratic forms.
\begin{theorem}[Fixed-bandwidth calibration of the projected statistic]\label{thm:fixedb}
Suppose $b\in(0,1]$ and $P\{\widetilde V_{\mathrm{PD},n}^{(b)}(\beta_0)>0\}\to1$. For some positive score scale $\Lambda_n$, assume
\[
\left(
\frac{n^{-1/2}\sum_{i=1}^n\psi_{i,n}}{\Lambda_n^{1/2}},
\frac{n^{-1}\psi_n'K_{n,b}\psi_n}{\Lambda_n}
\right)
\overset{d}{\to}
\{W(1),Q_{k,b}(W)\},
\]
with $Q_{k,b}(W)>0$ almost surely. Suppose also that the OLS linearization and feasible-score replacement are negligible at this fixed-bandwidth scale:
\[
\frac{\sqrt n(\widehat\beta-\beta_0)-\sigma_{D,n}^{-2}n^{-1/2}\sum_i\psi_{i,n}}
{\sigma_{D,n}^{-2}\Lambda_n^{1/2}}\overset{p}{\to}0,
\qquad
\frac{\widehat Q_n}{\sigma_{D,n}^2}\overset{p}{\to}1,
\]
\[
\frac{n^{-1}\{\widetilde\psi'K_{n,b}\widetilde\psi-\psi'K_{n,b}\psi\}}{\Lambda_n}\overset{p}{\to}0.
\]
Then
\[
\frac{\sqrt n(\widehat\beta-\beta_0)}
{\{\widetilde V_{\mathrm{PD},n}^{(b)}(\beta_0)\}^{1/2}}
\overset{d}{\to}
\frac{W(1)}{\{Q_{k,b}(W)\}^{1/2}}.
\]
Hence, if $c_{1-\alpha}(b,k)$ is a continuity-point critical value satisfying
\[
P\!\left\{\frac{|W(1)|}{Q_{k,b}(W)^{1/2}}>c_{1-\alpha}(b,k)\right\}=\alpha,
\]
then the two-sided test that rejects when the absolute studentized statistic exceeds $c_{1-\alpha}(b,k)$ has asymptotic size $\alpha$.
\end{theorem}
The theorem concerns calibration; it does not claim that the fixed-bandwidth quadratic form consistently estimates $V_n$. It implements fixed-bandwidth inference directly through the projected HAC/HAR variance and the corresponding nonstandard critical value. A different route would derive the fixed-bandwidth conditional law of an unstudentized dependent-multiplier or dependent-wild-bootstrap statistic and then construct bootstrap critical values or adjusted $p$-values. That route requires a separate conditional joint-limit theorem and is not implied by the covariance calculation here. In particular, the Gaussian multiplier scheme in Theorem~\ref{thm:pd} does not automatically generate the nonstandard fixed-bandwidth law after re-studentization; that too requires a separate bootstrap-$t$ argument. Fixed-bandwidth calibration changes the reference law, not the covariance direction. A misspecified $K_{n,b}$ that omits relevant score covariance remains misspecified. The Brownian-motion functional is tied to the oracle null-imposed score and the explicit fixed-bandwidth feasible-score replacement condition; if estimation altered the score partial-sum process at first order, the maintained joint limit---and hence the reference functional---would have to be changed accordingly.
\begin{corollary}[Binary randomized treatment]\label{cor:pdbinary}
Under the binary-treatment conditions of Corollary~\ref{cor:te}, if the projected-dependence procedure satisfies Theorem~\ref{thm:pd}, then it is first-order valid. Thus the dependence correction need only recover the scalar treatment-effect covariance contribution identified in Corollary~\ref{cor:te} at the first-order sampling-variance scale; it need not reproduce the full dependence of the potential-outcome vector.
\end{corollary}
For Gaussian multipliers with the projected covariance kernel under the small-bandwidth theory of Theorem~\ref{thm:pd}, the conditional law is exactly normal, so the corresponding test is numerically identical to normal inference using the projected HAC or cluster variance estimator. The fixed-bandwidth calibration in Theorem~\ref{thm:fixedb} changes the reference law when the kernel-bandwidth fraction is not negligible. No bootstrap simulation is required. We use this representation because it makes the variance target transparent, not because Gaussian multipliers constitute new resampling technology. Non-Gaussian dependent multipliers may change higher-order behavior, as in the wild-bootstrap literature \citep{Mammen1993,DavidsonFlachaire2008,Shao2010,Hounyo2023}, but covariance matching alone does not make an arbitrary non-Gaussian multiplier array valid: the conditional standardized CLT in Theorem~\ref{thm:pd} must also hold. Higher-order refinement is outside the present first-order characterization. Online Appendix Sections~\ref{app:repair}--\ref{app:supplementary} give time-series and cluster implementations.
\paragraph{Descriptive relevance summary.}
On $\{\widetilde V_{\mathrm{PD},n}(\beta_0)>0\}$ define
\[
\widehat r_n=1-\frac{\widetilde V_n^{\#}(\beta_0)}{\widetilde V_{\mathrm{PD},n}(\beta_0)},
\]
with an arbitrary bookkeeping value on the complement. Online Appendix Proposition~\ref{prop:rhat} shows that $\widehat r_n$ consistently estimates $R_{A,n}$ when both variance estimators are relatively consistent. We report it only as a normalization of the gap between those two variance estimates. Its finite-sample sign is noisy in moderate samples, and a valid test of $R_{A,n}=0$ would require additional joint limit theory for an estimator of the aggregate off-diagonal covariance. In particular, simply summing all sample off-diagonal score products does not yield a root-$n$ estimator of that object. The diagnostic is therefore descriptive, not a pretest or a stand-alone classifier.
\section{Monte Carlo evidence}\label{sec:mc}
We use iid Bernoulli assignment and $Y_i(0)=u_i$, $Y_i(1)=u_i+\tau+b_i$, where $u_i$ is independent of treatment and of the stationary treatment-effect deviation $b_i$. The benchmark processes are
\[
b_i=\eta_i+\tfrac12\eta_{i-1},\qquad
b_i=\eta_i+\eta_{i-1}-\tfrac12\eta_{i-2},
\]
and $b_i=\eta_i-\tfrac12\eta_{i-1}$. Their limiting aggregate treatment-effect covariance contributions are, respectively, $1$, $0$, and $-1$; the finite-$n$ quantities include the usual edge corrections. We use $p\in\{0.25,0.5,0.75\}$, $n\in\{100,250,500,1000\}$, and 20,000 replications per cell. Monte Carlo standard errors are reported for rejection frequencies.
\begin{table}[!htbp]\centering
\caption{Monte Carlo results at $n=1000$ (20,000 replications per cell).}\label{tab:mc1000}
\scriptsize
\begin{tabular}{lrrrrrr}\toprule
Design & $p$ & $V_n^\#$ & $V_n$ & Avg. null variance & Rejection & MC s.e.\\\midrule
Constant effects & 0.25 & 5.333 & 5.333 & 5.344 & 4.72\% & 0.15 pp \\
Constant effects & 0.50 & 4.000 & 4.000 & 4.002 & 4.98\% & 0.15 pp \\
Constant effects & 0.75 & 5.333 & 5.333 & 5.342 & 5.07\% & 0.16 pp \\
Positive MA(1) & 0.25 & 10.333 & 11.332 & 10.342 & 6.12\% & 0.17 pp \\
Positive MA(1) & 0.50 & 6.500 & 7.499 & 6.498 & 6.76\% & 0.18 pp \\
Positive MA(1) & 0.75 & 7.000 & 7.999 & 7.011 & 6.88\% & 0.18 pp \\
Cancellation MA(2) & 0.25 & 14.333 & 14.334 & 14.342 & 4.81\% & 0.15 pp \\
Cancellation MA(2) & 0.50 & 8.500 & 8.501 & 8.504 & 5.04\% & 0.15 pp \\
Cancellation MA(2) & 0.75 & 8.333 & 8.334 & 8.335 & 4.92\% & 0.15 pp \\
Negative MA(1) & 0.25 & 10.333 & 9.334 & 10.354 & 3.92\% & 0.14 pp \\
Negative MA(1) & 0.50 & 6.500 & 5.501 & 6.500 & 3.27\% & 0.13 pp \\
Negative MA(1) & 0.75 & 7.000 & 6.001 & 7.010 & 3.38\% & 0.13 pp \\
\bottomrule\end{tabular}
\end{table}
At $p=0.5$ and $n=1000$, positive dependence gives $V_n^\#=6.5$ and $V_n=7.499$, so the dependence-ignorant target is about $13.3\%$ below the first-order sampling variance; in Table~\ref{tab:mc1000} the average null-imposed iid score variance is $6.498$ and rejection is $6.76\%$ (Monte Carlo standard error $0.18$ percentage points). With negative dependence, $V_n=5.501$ while the iid target remains $6.5$, and rejection is $3.27\%$ (standard error $0.13$ percentage points). In the cancellation design, the theoretical variance is $8.501$, the average null-imposed variance is $8.504$, and rejection is $5.04\%$ (standard error $0.15$ percentage points). These numbers come from the baseline grid reported in the table; the separate projected experiment below uses independent Monte Carlo streams. The rejection patterns track the variance-target gap, not dependence by itself.
\begin{table}[!htbp]\centering
\caption{Convergence at $p=0.5$ (20,000 replications per cell).}\label{tab:mcp05}
\scriptsize
\begin{tabular}{lrrrrrr}\toprule
Design & $n$ & $V_n$ & Emp. var. & Avg. null variance & Rejection & MC s.e.\\\midrule
Constant effects & 100 & 4.000 & 4.024 & 3.993 & 4.93\% & 0.15 pp \\
Constant effects & 250 & 4.000 & 3.995 & 4.003 & 5.04\% & 0.15 pp \\
Constant effects & 500 & 4.000 & 4.019 & 4.001 & 4.78\% & 0.15 pp \\
Constant effects & 1000 & 4.000 & 3.944 & 4.002 & 4.98\% & 0.15 pp \\
Positive MA(1) & 100 & 7.490 & 7.600 & 6.486 & 7.01\% & 0.18 pp \\
Positive MA(1) & 250 & 7.496 & 7.596 & 6.507 & 7.06\% & 0.18 pp \\
Positive MA(1) & 500 & 7.498 & 7.499 & 6.499 & 6.86\% & 0.18 pp \\
Positive MA(1) & 1000 & 7.499 & 7.495 & 6.498 & 6.76\% & 0.18 pp \\
Cancellation MA(2) & 100 & 8.510 & 8.601 & 8.486 & 5.31\% & 0.16 pp \\
Cancellation MA(2) & 250 & 8.504 & 8.618 & 8.504 & 5.03\% & 0.15 pp \\
Cancellation MA(2) & 500 & 8.502 & 8.657 & 8.495 & 5.10\% & 0.16 pp \\
Cancellation MA(2) & 1000 & 8.501 & 8.480 & 8.504 & 5.04\% & 0.15 pp \\
Negative MA(1) & 100 & 5.510 & 5.546 & 6.518 & 3.26\% & 0.13 pp \\
Negative MA(1) & 250 & 5.504 & 5.470 & 6.503 & 3.23\% & 0.13 pp \\
Negative MA(1) & 500 & 5.502 & 5.654 & 6.505 & 3.68\% & 0.13 pp \\
Negative MA(1) & 1000 & 5.501 & 5.466 & 6.500 & 3.27\% & 0.13 pp \\
\bottomrule\end{tabular}
\end{table}
Table~\ref{tab:mcp05} reports convergence over sample size at $p=0.5$. Additional baseline figures and remaining cells are reported in Online Appendix Section~\ref{app:supplementary}. Null imposition also improves finite-sample targeting in imbalanced-treatment cells, without changing the structural discrepancy $V_n^\#-V_n$.
\subsection{Finite-sample performance of projected dependence correction}\label{sec:mcprojected}
We begin with the cluster geometry because it maps directly to settings in which individually randomized units share a known grouping structure---for example, schools, firms, villages, or local labor markets. Table~\ref{tab:mccluster} keeps iid unit-level assignment but gives treatment-effect heterogeneity an independent cluster component. The groups used by the variance estimator are the true groups in the data-generating process. As cluster size grows from two to eight, iid rejection rises from $7.11\%$ to $18.54\%$ at a nominal $5\%$ level, whereas the projected cluster correction stays near $5\%$. The experiment changes the dependence geometry of treatment effects, not the assignment mechanism.
\begin{table}[!htbp]\centering
\caption{Projected cluster correction under iid unit-level assignment ($n=1000$, $p=0.5$, 20,000 common Monte Carlo samples). Parentheses report Monte Carlo standard errors in percentage points.}\label{tab:mccluster}
\resizebox{\textwidth}{!}{\begin{tabular}{rrrrrrrr}\toprule
Cluster size & Clusters & $V_n^\#$ & $V_n$ & IID reject. & Projected-cluster reject. & Oracle-cluster reject. & Mean $\widehat r_n$\\
\midrule
2 & 500 & 6.000 & 7.000 & 7.11\% (0.18) & 5.14\% (0.16) & 5.18\% (0.16) & 0.141 \\
4 & 250 & 6.000 & 9.000 & 11.10\% (0.22) & 5.04\% (0.15) & 5.02\% (0.15) & 0.329 \\
8 & 125 & 6.000 & 13.000 & 18.54\% (0.27) & 5.08\% (0.16) & 5.03\% (0.15) & 0.533 \\
\bottomrule\end{tabular}}\end{table}
We then return to the ordered time-series designs, where the sign of surviving covariance can be varied transparently. We compare iid multiplier-score inference with projected Bartlett-HAC inference applied only to the null-imposed scalar score and, in the baseline cell, an oracle-HAC benchmark. For this experiment we use $\ell=\lfloor n^{1/4}\rceil$ with Bartlett weights $1-h/(\ell+1)$. This deterministic rule is chosen for transparency, not as an optimal bandwidth prescription. Standard automatic HAC rules such as \citet{Andrews1991HAC} and \citet{NeweyWest1994}, and test-oriented criteria such as \citet{SunPhillipsJin2008}, can be used in applications on the projected score, but our deterministic-bandwidth theory does not itself validate plug-in selection. Because the relevance diagnostic is a ratio, a bandwidth optimized for long-run-variance estimation need not optimize the diagnostic itself. We therefore report explicit bandwidth sensitivity rather than presenting one rule as universally preferred.
\begin{table}[!htbp]\centering
\caption{Projected-HAC correction at $n=1000$, $p=0.5$ (20,000 common Monte Carlo samples within each design). Parentheses report Monte Carlo standard errors in percentage points.}\label{tab:mcprojected}
\scriptsize
\begin{tabular}{lrrrrrrr}\toprule
Design & $L$ & IID reject. & Projected reject. & $E[\widetilde V_n^\#]$ & $E[\widetilde V_{\mathrm{PD},n}]$ & $R_{A,n}$ & Mean $\widehat r_n$\\\midrule
Positive MA(1) & 6 & 6.93\% (0.18) & 5.07\% (0.16) & 6.502 & 7.353 & 0.133 & 0.109 \\
Cancellation MA(2) & 6 & 4.77\% (0.15) & 4.55\% (0.15) & 8.503 & 8.647 & 0.000 & 0.009 \\
Negative MA(1) & 6 & 3.40\% (0.13) & 4.78\% (0.15) & 6.497 & 5.649 & -0.182 & -0.159 \\
\bottomrule\end{tabular}\end{table}
\begin{figure}[!htbp]
\centering
\includegraphics[width=.85\textwidth]{figure_projected_rejection.pdf}
\caption{Rejection probabilities for iid, projected-HAC, and oracle-HAC score inference at nominal 5\% (20,000 common Monte Carlo samples within each design).}
\label{fig:mcprojected}
\end{figure}
The projected correction behaves in the directions predicted by the phase diagram. Positive surviving covariance makes iid inference anti-conservative and the projected variance moves rejection toward nominal size; negative surviving covariance makes iid inference conservative and the correction moves rejection upward; under covariance cancellation both procedures remain close. The projected test is not exact in finite samples, however. The Bartlett lag window creates the usual bias--variability tradeoff, and the direction of the lag-window bias interacts with the sign of the surviving autocovariance: in Table~\ref{tab:mcprojected} the mean projected variance is below $V_n$ in the positive design and above $V_n$ in the negative design. This helps explain why the projected rejection rates are not symmetric around $5\%$.
Bandwidth sensitivity is economically relevant rather than cosmetic. Online Appendix Table~\ref{tab:stressL} shows that at $n=1000$ the normal-calibrated projected rejection rate falls from $5.45\%$ to $4.49\%$ in the positive design as $\ell$ rises from 2 to 25, with analogous movement elsewhere. A new moderate-sample experiment at $n=250$ shows larger drift. Table~\ref{tab:fixedbmain} therefore compares the same projected statistic under the standard-normal and fixed-bandwidth reference laws at selected bandwidths. Over the grid examined, fixed-bandwidth calibration removes much of the residual under-rejection at large $\ell$, while leaving the covariance target unchanged; it is a calibration device, not a guarantee of uniformly smaller finite-sample size error.
The dependence specification also matters. Online Appendix Table~\ref{tab:groupmisspec} uses split, merged, and scrambled versions of the true cluster structure. Correct grouping restores size; splitting or scrambling the effect-modifier clusters leaves material covariance unaccounted for, while merging independent true clusters preserves the target on average but can increase variance-estimation noise. This is the finite-sample counterpart of the minimality theorem: only error in the surviving covariance direction matters, but a practitioner must still supply a defensible dependence model.
\begin{table}[!htbp]
\centering
\caption{Selected fixed-bandwidth calibration results. Here $\ell$ is the largest included positive Bartlett lag and $b=(\ell+1)/n$ is the kernel-bandwidth fraction. The projected statistic is identical across the last two rejection columns; only the reference critical value changes. Parentheses report Monte Carlo standard errors in percentage points.}
\label{tab:fixedbmain}
\small
\begin{tabular}{lrrrrr}
\toprule
Design & $n$ & $\ell$ & $b$ & Normal reject. & Fixed-$b$ reject.\\
\midrule
Positive MA(1) & 250 & 16 & 0.068 & 4.01\% (0.14) & 5.38\% (0.16) \\
Cancellation MA(2) & 250 & 16 & 0.068 & 3.74\% (0.13) & 5.10\% (0.16) \\
Negative MA(1) & 250 & 16 & 0.068 & 3.77\% (0.13) & 5.11\% (0.16) \\
Positive MA(1) & 1000 & 25 & 0.026 & 4.66\% (0.15) & 5.12\% (0.16) \\
Cancellation MA(2) & 1000 & 25 & 0.026 & 4.39\% (0.14) & 4.86\% (0.15) \\
Negative MA(1) & 1000 & 25 & 0.026 & 4.48\% (0.15) & 4.86\% (0.15) \\
\bottomrule
\end{tabular}
\end{table}
The ratio $\widehat r_n$ is reported in the tables as a descriptive normalization only. Table~\ref{tab:rhatdist} and Online Appendix Tables~\ref{tab:stressn}--\ref{tab:stresstheta} document that its finite-sample sign can be noisy and more bandwidth-sensitive than the projected test.
\begin{table}[!htbp]\centering
\caption{Finite-sample behavior of the relevance diagnostic ($p=0.5$, $L=\lfloor n^{1/4}\rceil$, 20,000 replications per row). ``Wrong sign'' is relative to the sign of $R_{A,n}$.}\label{tab:rhatdist}
\scriptsize
\begin{tabular}{lrrrrrrr}\toprule
Design & $n$ & $R_{A,n}$ & Mean $\widehat r_n$ & 5th pct. & 95th pct. & Wrong sign & Projected reject.\\\midrule
Positive MA(1) & 100 & 0.132 & 0.068 & -0.245 & 0.311 & 30.6\% & 5.00\% \\
Positive MA(1) & 250 & 0.133 & 0.091 & -0.132 & 0.275 & 21.7\% & 5.03\% \\
Positive MA(1) & 500 & 0.133 & 0.102 & -0.072 & 0.249 & 14.8\% & 5.16\% \\
Positive MA(1) & 1000 & 0.133 & 0.110 & -0.020 & 0.227 & 8.0\% & 5.32\% \\
Negative MA(1) & 100 & -0.180 & -0.164 & -0.576 & 0.160 & 24.7\% & 3.88\% \\
Negative MA(1) & 250 & -0.181 & -0.160 & -0.455 & 0.083 & 16.3\% & 4.61\% \\
Negative MA(1) & 500 & -0.181 & -0.160 & -0.391 & 0.037 & 10.1\% & 4.47\% \\
Negative MA(1) & 1000 & -0.182 & -0.160 & -0.332 & -0.002 & 4.7\% & 4.33\% \\
\bottomrule\end{tabular}\end{table}
\section{Conclusion}
Random assignment can simplify bootstrap inference even when the observed data remain dependent. What matters for the randomized-regressor coefficient is whether the bootstrap reproduces the part of that dependence that survives the assignment filter.
Under strong exogeneity, random assignment removes the relevant off-diagonal score covariance, so an iid multiplier bootstrap can be first-order valid despite dependence in outcomes and controls. When the regression error changes with assignment, a scalar covariance component can survive. An iid bootstrap deletes that component. With binary treatment, the resulting first-order variance gap is exactly minus the aggregate cross-observation covariance of heterogeneous treatment effects. Under the standardized Gaussian sampling limit, this gap makes iid inference over-reject, under-reject, or remain first-order correct according to the sign of the relative surviving covariance.
A valid centered-Gaussian procedure need not reproduce the full dependence structure; it needs only the covariance component that survives the randomized-score projection. The information-limit theorem shows why marginal information alone cannot recover that component in general. For first-order Gaussian inference, this means estimating dependence in the scalar randomized-regressor score rather than resampling the full observed-data vector. With a small bandwidth, the Gaussian implementation is a HAC or cluster-normal test on that score. When the bandwidth is non-negligible relative to sample size, the reference law instead requires fixed-bandwidth calibration. The ratio $\widehat r_n$ is best treated as a descriptive normalization of the diagonal--projected variance gap. The grouping experiments also make clear what projection cannot do: it reduces the dimension of the dependence problem but does not identify the appropriate ordering, clusters, or network for the researcher.
The local-power calculation leads to the same conclusion. The apparent power advantage or disadvantage of the raw iid test comes from its size distortion: after equal-size calibration, the iid and projected procedures have the same Gaussian first-order local-power curve. The oracle iid calibration, however, depends on the unknown relevance index, whereas the projected procedure estimates the needed covariance contribution. Its gain is calibration and feasibility, not first-order efficiency. In short, random assignment and resampling need not remove the same covariance. Once the surviving covariance component has been identified, inference can focus on that component instead of modeling the full dependence structure.
\begin{thebibliography}{99}
\small
\bibitem[Abadie et al.(2023)]{AbadieEtAl2023} Abadie, A., Athey, S., Imbens, G. W., and Wooldridge, J. M. (2023). When should you adjust standard errors for clustering? \emph{Quarterly Journal of Economics} 138(1), 1--35.
\bibitem[Abadie et al.(2020)]{AbadieAtheyImbensWooldridge2020} Abadie, A., Athey, S., Imbens, G. W., and Wooldridge, J. M. (2020). Sampling-based versus design-based uncertainty in regression analysis. \emph{Econometrica} 88(1), 265--296. doi:10.3982/ECTA12675.
\bibitem[Andrews(1991)]{Andrews1991HAC} Andrews, D. W. K. (1991). Heteroskedasticity and autocorrelation consistent covariance matrix estimation. \emph{Econometrica} 59(3), 817--858.
\bibitem[Aronow et al.(2014)]{AronowGreenLee2014} Aronow, P. M., Green, D. P., and Lee, D. K. K. (2014). Sharp bounds on the variance in randomized experiments. \emph{Annals of Statistics} 42(3), 850--871. doi:10.1214/13-AOS1200.
\bibitem[Barrios et al.(2012)]{BarriosEtAl2012} Barrios, T., Diamond, R., Imbens, G. W., and Kolesar, M. (2012). Clustering, spatial correlations, and randomization inference. \emph{Journal of the American Statistical Association} 107(498), 578--591.
\bibitem[Bradley(2005)]{Bradley2005} Bradley, R. C. (2005). Basic properties of strong mixing conditions. A survey and some open questions. \emph{Probability Surveys} 2, 107--144.
\bibitem[Bugni et al.(2018)]{BugniCanayShaikh2018} Bugni, F. A., Canay, I. A., and Shaikh, A. M. (2018). Inference under covariate-adaptive randomization. \emph{Journal of the American Statistical Association} 113(524), 1784--1796. \url{https://doi.org/10.1080/01621459.2017.1375934}.
\bibitem[Cameron et al.(2008)]{CameronEtAl2008} Cameron, A. C., Gelbach, J. B., and Miller, D. L. (2008). Bootstrap-based improvements for inference with clustered errors. \emph{Review of Economics and Statistics} 90(3), 414--427.
\bibitem[Canay et al.(2021)]{CanaySantosShaikh2021} Canay, I. A., Santos, A., and Shaikh, A. M. (2021). The wild bootstrap with a ``small'' number of ``large'' clusters. \emph{Review of Economics and Statistics} 103(2), 346--363. doi:10.1162/rest\_a\_00887.
\bibitem[Chetverikov et al.(2026)]{CHLS2026} Chetverikov, D., Hahn, J., Liao, Z., and Santos, A. (2026). Standard errors when a regressor is randomly assigned. \emph{Quantitative Economics}, forthcoming. Working-paper version: arXiv:2303.10306 (March 22, 2026 revision).
\bibitem[Davidson and Flachaire(2008)]{DavidsonFlachaire2008} Davidson, R. and Flachaire, E. (2008). The wild bootstrap, tamed at last. \emph{Journal of Econometrics} 146(1), 162--169.
\bibitem[de Jong(2000)]{deJong2000HAC} de Jong, R. M. (2000). A strong consistency proof for heteroskedasticity and autocorrelation consistent covariance matrix estimators. \emph{Econometric Theory} 16(2), 262--268.
\bibitem[Djogbenou et al.(2019)]{DjogbenouEtAl2019} Djogbenou, A. A., MacKinnon, J. G., and Nielsen, M. \O. (2019). Asymptotic theory and wild bootstrap inference with clustered errors. \emph{Journal of Econometrics} 212(2), 393--412. doi:10.1016/j.jeconom.2019.04.035.
\bibitem[Freedman(1981)]{Freedman1981} Freedman, D. A. (1981). Bootstrapping regression models. \emph{The Annals of Statistics} 9(6), 1218--1228.
\bibitem[Gon\c{c}alves and White(2005)]{GoncalvesWhite2005} Gon\c{c}alves, S. and White, H. (2005). Bootstrap standard error estimates for linear regression. \emph{Journal of the American Statistical Association} 100(471), 970--979.
\bibitem[Hahn and Liao(2021)]{HahnLiao2021} Hahn, J. and Liao, Z. (2021). Bootstrap standard error estimates and inference. \emph{Econometrica} 89(4), 1963--1977.
\bibitem[Hounyo(2023)]{Hounyo2023} Hounyo, U. (2023). A wild bootstrap for dependent data. \emph{Econometric Theory} 39(2), 264--289. \url{https://doi.org/10.1017/S0266466621000487}.
\bibitem[Ibragimov and M\"uller(2010)]{IbragimovMuller2010} Ibragimov, R. and M\"uller, U. K. (2010). t-Statistic based correlation and heterogeneity robust inference. \emph{Journal of Business \& Economic Statistics} 28(4), 453--468. doi:10.1198/jbes.2009.08046.
\bibitem[Imbens and Menzel(2021)]{ImbensMenzel2021} Imbens, G. W. and Menzel, K. (2021). A causal bootstrap. \emph{The Annals of Statistics} 49(3), 1460--1488. doi:10.1214/20-AOS2009.
\bibitem[Jiang et al.(2024)]{JiangEtAl2024} Jiang, L., Liu, X., Phillips, P. C. B., and Zhang, Y. (2024). Bootstrap inference for quantile treatment effects in randomized experiments with matched pairs. \emph{The Review of Economics and Statistics} 106(2), 542--556.
\bibitem[Kiefer and Vogelsang(2005)]{KieferVogelsang2005} Kiefer, N. M. and Vogelsang, T. J. (2005). A new asymptotic theory for heteroskedasticity-autocorrelation robust tests. \emph{Econometric Theory} 21(6), 1130--1164.
\bibitem[Kojevnikov et al.(2021)]{KojevnikovEtAl2021} Kojevnikov, D., Marmer, V., and Song, K. (2021). Limit theorems for network dependent random variables. \emph{Journal of Econometrics} 222(2), 882--908. doi:10.1016/j.jeconom.2020.05.019.
\bibitem[Kunsch(1989)]{Kunsch1989} Kunsch, H. R. (1989). The jackknife and the bootstrap for general stationary observations. \emph{Annals of Statistics} 17(3), 1217--1241.
\bibitem[Lin(2013)]{Lin2013} Lin, W. (2013). Agnostic notes on regression adjustments to experimental data: Reexamining Freedman's critique. \emph{Annals of Applied Statistics} 7(1), 295--318.
\bibitem[MacKinnon et al.(2023)]{MacKinnonNielsenWebb2023} MacKinnon, J. G., Nielsen, M. \O., and Webb, M. D. (2023). Cluster-robust inference: A guide to empirical practice. \emph{Journal of Econometrics} 232(2), 272--299. doi:10.1016/j.jeconom.2022.04.001.
\bibitem[Mammen(1993)]{Mammen1993} Mammen, E. (1993). Bootstrap and wild bootstrap for high-dimensional linear models. \emph{Annals of Statistics} 21(1), 255--285.
\bibitem[Newey and West(1987)]{NeweyWest1987} Newey, W. K. and West, K. D. (1987). A simple, positive semi-definite, heteroskedasticity and autocorrelation consistent covariance matrix. \emph{Econometrica} 55(3), 703--708.
\bibitem[Newey and West(1994)]{NeweyWest1994} Newey, W. K. and West, K. D. (1994). Automatic lag selection in covariance matrix estimation. \emph{Review of Economic Studies} 61(4), 631--653.
\bibitem[Politis and Romano(1994)]{PolitisRomano1994} Politis, D. N. and Romano, J. P. (1994). The stationary bootstrap. \emph{Journal of the American Statistical Association} 89(428), 1303--1313.
\bibitem[Shao(2010)]{Shao2010} Shao, X. (2010). The dependent wild bootstrap. \emph{Journal of the American Statistical Association} 105(489), 218--235.
\bibitem[Sun et al.(2008)]{SunPhillipsJin2008} Sun, Y., Phillips, P. C. B., and Jin, S. (2008). Optimal bandwidth selection in heteroskedasticity--autocorrelation robust testing. \emph{Econometrica} 76(1), 175--194. \url{https://doi.org/10.1111/j.0012-9682.2008.00822.x}.
\bibitem[Zhang et al.(2025)]{ZhangEtAl2025} Zhang, Z., Ding, P., Zhou, W., and Wang, H. (2025). With random regressors, least squares inference is robust to correlated errors with unknown correlation structure. \emph{Biometrika} 112(1), asae054. doi:10.1093/biomet/asae054.
\bibitem[Zhang and Zheng(2020)]{ZhangZheng2020} Zhang, Y. and Zheng, X. (2020). Quantile treatment effects and bootstrap inference under covariate-adaptive randomization. \emph{Quantitative Economics} 11, 957--982.
\end{thebibliography}
\clearpage
\thispagestyle{plain}
\begin{center}
{\Large\bfseries Online Appendix to\par}
\vspace{0.7em}
{\Large ``Bootstrap Inference with a Randomly Assigned Regressor:\\
Covariance Filtering and the Limits of Marginal Resampling''\par}
\vspace{1.5em}
{\large Ulrich Hounyo \qquad Jungbin Hwang\par}
\vspace{0.8em}
{\large October 2, 2026\par}
\end{center}
\vspace{2em}