Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
123,374 characters · 27 sections · 32 citation commands
Finite-Population Inference for Heterogeneity in Many-Group Synthetic Difference-in-Differences
\noindentKey Words: synthetic control; treatment effect heterogeneity; donor sharing; variance component; finite-population inference.
In the evaluation of policies and interventions, it is common for treatment to be assigned at the unit of a group (states, municipalities, school districts, hospitals, firms, regional markets, and so on) rather than individuals. In many such settings the treated groups share a common policy shock---in our applications, Medicaid expansion in 2014 and Clean Air Act nonattainment designation in 2005---forming a cohort across which one asks how the effect varies; we develop inference for this common-cohort case and defer staggered adoption to the conclusion. In this setting, what the researcher wants to know is often not only the average effect but the structure of heterogeneity---for which kinds of groups is the effect large (or small). Examples include whether the coverage effect of a health-insurance expansion depends on a state's pre-expansion coverage gap, or whether an environmental regulation reduces pollution more strongly in counties with higher baseline pollution. These are first-order concerns that bear directly on the external validity, generalizability, and optimal targeting of effects.
We target data with the following structure. Among groups $g=1,\dots,G$, both the treated-group set $\mathcal{T}$ ($|\mathcal{T}|=G_1$) and the control-group set $\mathcal{C}$ ($|\mathcal{C}|=G_0$) are numerous ($G_1,G_0$ large), and units $i=1,\dots,n_g$ belonging to each group $g$ are observed. Each unit carries unit covariates $X_{gi}$, and each group carries group covariates $W_g$. Treatment is assigned at the group level, and we write the group-specific ATT of group $g\in\mathcal{T}$ as $\tau_g$. Individual-level effects may vary within a group: writing $Y_{git}(1)=Y_{git}(0)+\Delta_{git}$, the group-specific ATT is the post-period average $\tau_g=|\mathcal T_1|^{-1}\sum_{t\in\mathcal T_1}\mathbb{E}(\Delta_{git}\mid g,t)$, so all finite-population objects below are functions of these group-level averages rather than of any constant individual effect. The interest lies in the dependence of $\tau_g$ on $W_g$ (first moment) and in the share of the dispersion of $\tau_g$ that can be explained by observed group characteristics (second moment).
This question faces two obstacles. First, because of group-level assignment, estimating each $\tau_g$ requires constructing the counterfactual for treated group $g$ from the control groups, so the estimation of synthetic-control weights and the residualization of covariates are unavoidable. When the estimation errors of these nuisances propagate into the second stage (projection and variance decomposition), inference is easily distorted. Second, when estimating the variance of $\tau_g$, the sampling error contained in each $\widehat\tau_g$ systematically biases the variance estimator upward (effects appear more dispersed than they are). This is a problem well known in the variance estimation of teacher value-added and establishment effects, and overestimation is unavoidable with naive plug-in estimation.
There is, moreover, a methodological tension: an aversion to random effects. Hierarchical models (random-effects models) that impose a prior or an exchangeable stochastic structure on $\tau_g$ are efficient through partial pooling, but in realistic situations where $\tau_g$ is correlated with $W_g$ or with treatment selection, they invite the criticism that the exchangeability assumption (unconditional independence) itself may drive the estimator. This paper responds to this tension head-on.
We adopt a finite-population standpoint that regards the group-effect sequence $\{\tau_g\}_{g\in\mathcal{T}}$ as fixed parameters, and we define two estimands without imposing distributional assumptions. One is the finite-population best linear predictor $\gamma^{*}=\operatorname*{arg\,min}_{\gamma}\sum_{g\in\mathcal{T}}(\tau_g-W_g^{\!\top}\gamma)^2$ of $\{\tau_g\}$ onto the group covariates $\{W_g\}$ (first moment); the other is the between-group variance $V$, the group covariance $C$, and the explained share $R^2=V_{\mathrm{expl}}/V$ by observed group characteristics (second moment).
The methodological core consists of three central contributions and the test theory that supports them.
We use a self-contained orthogonalized SDID first stage to obtain a linear representation of the group-specific estimates; orthogonality is conditional on a local bridge/balance condition, and online Appendix A gives primitive sufficient conditions and a basis-augmented variant that enforces it by design. The contribution of the paper begins from this representation and lies in the second-stage finite-population inference. Three contributions form a single moment system, without distributional assumptions. (i) Projected group-effect curve and variance decomposition. We define and estimate the projection $\gamma^{*}$ of the group effects onto observed group covariates and the induced projected group-effect curve $m_{\mathcal{T}}(w)=w^{\!\top}\gamma^{*}$---the effect predicted for a group with covariate value $w$---together with pointwise and finite-grid simultaneous confidence bands (Corollary (ref)), as well as the between-group variance $V$, the directed covariance $C$, and the explained share $R^2$, treating the group-effect sequence as fixed parameters. (ii) Joint covariance under shared donors. Because every treated group is synthesized from the same donors, the estimated effects are cross-sectionally dependent; we derive the joint covariance kernel that carries this dependence---and, more generally, block and spatial cross-dependence---into the second-stage inference, with donor sharing a special case. (iii) Noise correction and calibrated tests. The plug-in second moments are biased upward by first-stage estimation noise; we remove the bias by an analytic correction and a household $A/B$ leave-out, and we calibrate omnibus $H_0:V=0$ and directed $H_0:C=0$ tests that remain valid when counterfactual construction is imperfect. A directed test robust to the equilibration residual is the practically important instrument when the omnibus test is under-powered. A central message is methodological: existing group-wise DiD or SDID regressions may deliver similar point estimates for the projection slope and the projected group-effect curve $m_{\mathcal{T}}(w)$, but they do not deliver the joint donor-sharing covariance, the noise-corrected second moments, or calibrated inference for $V$, $C$, and the curve. The distinctive contribution is therefore not the point estimates but the valid uncertainty quantification for group-level heterogeneity---confidence bands and tests that carry the shared-donor dependence and remove the estimation-noise bias.
The projected group-effect curve is the object most directly interpretable to applied researchers: rather than a single average or an abstract slope, it answers “what effect should be expected for a group with this characteristic?” In the Medicaid application it answers what uninsured-rate reduction to expect for a state with a given pre-expansion uninsured rate; in the Clean Air Act application, how much PM$_{2.5}$ reduction is associated with a county's baseline pollution level. The novelty is not merely estimating the slope of this curve, but attaching a confidence band that accounts for the fact that all points on the curve are built from group effects sharing the same donors---a band that is wider than, and not recoverable from, the one an independent-outcome regression would report.
Monte Carlo experiments quantify the stakes: naive independence-based inference over-rejects the no-heterogeneity null severely and undercovers the projection intercept, and the joint covariance kernel together with the affine-calibrated placebo restores near-nominal size and coverage (Section (ref), Table (ref)). The Medicaid estimand is the incremental effect associated with expansion status relative to non-expansion states exposed to the other nationwide Affordable Care Act components (exchanges, subsidies, mandate), not the effect of the Act as a whole. The Medicaid application (Section (ref)) exhibits a large and robust bite gradient whose numerical explained share is scale-dependent; a household $A/B$ leave-out validates the variance decomposition against an independent analytic correction, and donor sharing inflates the average-effect standard error by a factor of $1.8$ while leaving the centered slope nearly unaffected. The Clean Air Act application (Section (ref)) illustrates the Type II-CD regime: county-wise estimates are individually noisy and the omnibus test is low-powered, yet the pre-specified directed projection on baseline PM$_{2.5}$ is sign-stable under state- and division-block covariance, survives pre-trend adjustment, and is corroborated by a placebo-calibrated excess and a baseline-matched county-donor design; an exploratory extension to county mortality is directionally consistent but lower-powered.
The foundation of this paper is synthetic difference-in-differences Arkhangelsky2021, which balances pre-treatment trends with donor and time weights and extends synthetic control Abadie2010,Abadie2015,BenMichael2021 and the difference-in-differences tradition Abadie2005,SantAnnaZhao2020. Estimation with many or disaggregated treated units is developed by penalized synthetic control AbadieLHour2021 and micro-level synthetic control Robbins2017, and inference for synthetic-control estimates has been developed mainly for a single treated unit or a pooled average---placebo permutations Abadie2010, conformal prediction ChernozhukovWuthrichZhu2021, prediction intervals CattaneoFengTitiunik2021, and subsampling for the average effect Li2020. None of these targets the joint distribution of the whole cross-section of group effects, its second moments, or the dependence induced when every treated group is synthesized from the same donor pool, which are the objects of this paper.
Orthogonal and debiased estimation of conditional average treatment effects (CATE) builds on the double machine learning of ChernozhukovDML2018; its low-dimensional projection (BLP) and group average effects (GATES) are treated by ChernozhukovGeneric2018, and orthogonal estimation and uniform inference of debiased CATE by SemenovaChernozhukov2021. Orthogonal and debiased estimation in the DiD setting is given by Chang2020. These are mainly at the individual level and cross-sectional, with the conditioning target being a function of continuous covariates. This paper, by contrast, targets a finite-population projection across groups and its variance decomposition, and is differentiated by integrating group-level assignment, two-stage nuisances, group-wise cross-fitting, and $G_1\to\infty$ double asymptotics.
Between-site variance estimation and empirical-Bayes inference (FIRC-type random-effects models, EB confidence intervals with frequentist guarantees, and Bayesian hierarchical and semiparametric approaches; Bloom2017,Armstrong2022,IgnatiadisWager2022), together with correlated-random-coefficient formulations, are reviewed in online Appendix G. The contrast that matters for this paper is invariant across that literature: those methods take the first-stage noise distribution as known and independent across groups, whereas our first stage generates heteroskedastic, donor-sharing-dependent noise whose covariance must itself be estimated. Closest to our second-moment construction, unbiased (leave-out) estimation of variance components in high-dimensional fixed-effects models was established by KSS2020, and variance estimation of teacher effects by ChettyFriedmanRockoff2014. We transplant this to the between-group variance of group effects estimated by SDID, but because the projection (hat) matrix in the weighted regression of SDID depends on the estimated weights, an extension of the leave-out algebra of quadratic forms is required. This extension, and the simultaneous handling of between-group dependence induced by donor sharing, are the theoretical core of this paper.
In short, this paper has its novelty in integrating (i) the projection of group effects (a group-level, finite-population version of the SemenovaChernozhukov2021 line), (ii) the variance decomposition of group effects (an SDID, between-group version of KSS2020), and (iii) joint inference under donor-sharing dependence (a refinement of the pooling of Dube2015) into a single moment system based on SDID, giving joint inference of $(\gamma^{*},V,C,R^2)$ without distributional assumptions.
Groups $g\in\{1,\dots,G\}$, treated-group set $\mathcal{T}$ ($|\mathcal{T}|=G_1$), control-group set $\mathcal{C}$ ($|\mathcal{C}|=G_0$), $G=G_0+G_1$. The notation used throughout is collected in online Appendix A. Units $i=1,\dots,n_g$ belong to group $g$, and the total number of units is $N=\sum_g n_g$. Among time points $t=1,\dots,T$, the pre-treatment period is $\mathcal T_0=\{1,\dots,T_0\}$ and the post-treatment period is $\mathcal T_1=\{T_0+1,\dots,T\}$ (we take simultaneous adoption by all treated groups as the baseline, and defer the extension to staggered adoption to the conclusion). We write the outcome of unit $(g,i)$ at time $t$ as $Y_{git}$ and the group-average trajectory as $\bar Y_{gt}=n_g^{-1}\sum_{i}Y_{git}$.
The treatment indicator is at the group level, $D_g=\mathbf{1}\{g\in\mathcal{T}\}$ (units of treated groups are all treated post-treatment). Unit covariates $X_{git}\in\mathbb{R}^{d_x}$ (possibly time-varying) are used for adjustment by orthogonalization (residualization), and group covariates $W_g\in\mathbb{R}^{p}$ ($W_g$ includes an intercept) are used for the projection of $\tau_g$.
Each unit has potential outcomes $\bigl(Y_{git}(1),Y_{git}(0)\bigr)$, and the observed outcome is \[ Y_{git}=D_g\,\mathbf{1}\{t\in\mathcal T_1\}\,Y_{git}(1)+\bigl(1-D_g\,\mathbf{1}\{t\in\mathcal T_1\}\bigr)\,Y_{git}(0) \] (pre-treatment, all units are untreated; no anticipation).
Our framework admits two data regimes. In the micro-replication regime (Type I), the outcomes $Y_{git}$ of multiple units within each group $g$ are observed, and the group-average trajectory $\bar Y_{gt}$ can be split into nearly independent subsamples. In this regime, the all-group $A/B$ leave-out variance decomposition of Section (ref) can be implemented directly. In the aggregate regime (Type II), only the group-average series $\bar Y_{gt}$ is observed. The first-stage SDID can be defined in the same way, but because $A/B$ leave-out cannot be executed, $V$ is replaced by analytic noise subtraction (or model-based covariance estimation). When units within a group are spatially or economically correlated (e.g., $i=$county), the $A/B$ split is done at the spatial-block level rather than for individual units, or an analytic correction based on a spatial-HAC-type covariance is used---that is, units are not treated as independent replicates. Because $\omega^{(g)}$ weights donor groups $h\in\mathcal{C}$ and not units $i$, we call these donor-group weights. In Section (ref), we present the Medicaid expansion application ($i=$household) as a full-fledged Type I application. In the spatially correlated version of Type I, where microdata are aggregated to correlated units such as counties, we use block-level splitting and covariance rather than individual units (our county aggregation layer corresponds to this).
In repeated-cross-section Type I data---such as the ACS microdata of Section (ref)---the primitive sampling units are $i\in\mathcal I_{gt}$, drawn independently within each group--time cell $(g,t)$, rather than persistent units observed for all $t$; we write $n_{gt}=|\mathcal I_{gt}|$ and $\bar Y_{gt}=n_{gt}^{-1}\sum_{i\in\mathcal I_{gt}}Y_{git}$. The all-group $A/B$ construction of Section (ref) then splits the sampling units within each cell into halves $\mathcal I^{A}_{gt}$ and $\mathcal I^{B}_{gt}$ and builds two group--time trajectories $\bar Y^{A}_{gt}$ and $\bar Y^{B}_{gt}$ that are independent conditional on the cell composition. All results below are stated at the level of the group-average trajectories and their half-sample versions, and therefore apply verbatim to persistent panels (splitting units $i$) and to repeated cross sections (splitting $\mathcal I_{gt}$ within each cell) alike.
Using unit covariates $X_{git}$ and unobserved factors, we posit an additively separable interactive fixed-effects model for the potential outcomes,
in which a common covariate function $h_0(X_{git})$ is removed by residualization and the group-time factor structure is spanned, approximately, by the donor groups. The group covariate $W_g$ enters only the effect-heterogeneity projection, not identification. The primitive factor-model formulation, the concentration of the group-average trajectory on its mean, the residualization projection, and the repeated-cross-section notation are given in online Appendix A.
Hereafter, let $\bar\tau=G_1^{-1}\sum_{g\in\mathcal{T}}\tau_g$, $\overline W=G_1^{-1}\sum_{g\in\mathcal{T}}W_g$, the centered group covariate $\widetilde{W}_g=W_g-\overline W$, and the (finite-population) variance matrix of the group covariates $\Sigma_W=G_1^{-1}\sum_{g\in\mathcal{T}}\widetilde{W}_g\widetilde{W}_g^{\!\top}$.
The counterfactual of each treated group $g\in\mathcal{T}$ is constructed from the donor pool $\mathcal{C}$. Let the pre-treatment period set be $\mathcal T_0$ and the post-treatment set be $\mathcal T_1$.
(The correspondence of Assumptions (ref)--(ref) to the standard SDID/synthetic-control identification conditions is discussed in online Appendix A.)
Next, we show the orthogonality of this estimand. We define the SDID estimand (population version) of treated group $g$ by the weighted double difference
where $\bar\mu_{g,\mathcal T_1}=|\mathcal T_1|^{-1}\sum_{t\in\mathcal T_1}\mu_{gt}$. Under Assumptions (ref)--(ref), $\alpha_g,\xi_t,\Gamma_g^{\!\top} F_t$ cancel from the right-hand side of (ref) and it coincides with $\tau_g$.
To adjust for observed covariates, we residualize the outcome as $Y^{\perp}_{git}=Y_{git}-h_0(X_{git})$. Here $h_0(\cdot)$ is the projection of $Y$ onto $X$ (the orthogonalization nuisance), which by Assumption (ref) coincides with the true covariate function $m$. The orthogonalized SDID moment is constructed by treating the weights $(\omega^{(g)},\lambda^{(g)})$ and the residualization $h_0$ as nuisance parameters, applying (ref) to the residualized outcome, and adding a correction derived from the first-order condition of the weight estimation.
online Appendix A provides the full primitive formulation---the outer-product Riesz representation, the mixed-remainder expansion, and the simplex-boundary argument---together with the proof; a sketch is reproduced there.
As explained in Section (ref), the first stage is any group-wise estimator of $\tau_g$ admitting the linear representation (ref); we use group-wise orthogonalized SDID. We residualize the outcome by a cross-fitted covariate function $\widehat h_0$, form the residualized group-time trajectory $\bar Y^{\perp}_{gt}$, and for each treated group solve for donor and time weights following Arkhangelsky2021,
where $\Delta$ is the simplex, $\zeta$ is a regularization parameter, and $\bar Y^{\perp}_{k,\mathcal T_1}=|\mathcal T_1|^{-1}\sum_{t\in\mathcal T_1}\bar Y^{\perp}_{kt}$. The group-ATT estimator is the weighted double difference
By the orthogonality of Lemma (ref), the estimation errors of $(\widehat\omega^{(g)},\widehat\lambda^{(g)},\widehat h_0)$ affect $\widehat\tau_g$ only at higher order.
In the second estimation stage, we project the group effects $\{\widehat\tau_g\}$ onto the group covariates. The projection coefficient is given by:
We perform the variance decomposition in the same second stage. In SDID, the sampling error of $\widehat\tau_g$ arises from both the treated side and the donor side, and the donor-side error is shared across all treated groups. Therefore the naive plug-in variance is biased not only through the variance of each $\widehat\tau_g$ but also through $\mathrm{Cov}(\widehat\tau_g,\widehat\tau_{g'})$. To remove this all at once, we randomly split the units of all groups (both treated and donor) into two halves $A,B$, and compute $\widehat\tau_g^{A},\widehat\tau_g^{B}$ by completing everything from residualization and weight estimation (ref)--(ref) to the double difference (ref) within each half-sample only. By construction, the vectors $(\widehat\tau_g^{A})_{g\in\mathcal{T}}$ and $(\widehat\tau_g^{B})_{g\in\mathcal{T}}$ are mutually independent conditional on the groups. Using these,
where $\widehat C^{A}=G_1^{-1}\sum_g\widetilde{W}_g\widehat\tau_g^{A}$, and similarly $\widehat C^{B}$. Since $C$ is linear in $\tau_g$, $\widehat C$ itself is unbiased, but $V$ and $V_{\mathrm{expl}}$ are quadratic functionals of $\tau_g$, and we use the cross product (leave-out) to remove the upward bias due to sampling error.
Proposition (ref) formalizes the elementary but crucial point: if the two half-sample effect vectors are conditionally independent and unbiased, their cross-product removes the noise bias in the quadratic functionals. The proof is in online Appendix B.
The premise of Proposition (ref) is stated at the level of the half-sample vectors and therefore covers both data regimes: in a persistent panel the split is over units, and in a repeated cross section it is performed within each group--time cell (Section (ref)), which is what delivers the conditional independence of $\widehat{\boldsymbol\tau}^{A}$ and $\widehat{\boldsymbol\tau}^{B}$ in the ACS application.
The all-group split (splitting every group, treated and donor, into halves) is necessary for the leave-out variance to be unbiased under the nonlinearity of the weights, and centering absorbs the common donor shock; a formal statement is in online Appendix B.
We state the asymptotic theory in a form that allows cross-sectional dependence among the primitive group-level shocks from the outset; the donor-sharing-only kernel of the baseline theory is the special case in which the primitive shocks are cross-sectionally independent, and the state- and division-block covariances used in the Type II application are alternative implementations of the same kernel. As the asymptotic framework, we take $G_1\to\infty$, $G_0\to\infty$, $T_0\to\infty$ ($|\mathcal T_1|$ may be fixed) as the baseline, and take the number of units $n_g$ in each group to be bounded below ($n_g\ge n_{\min}$). The primitive shock of series $k\in\mathcal{T}\cup\mathcal{C}$ is the vector of group-average errors $\varepsilon_k=(\bar\varepsilon_{kt})_t$ with $\bar\varepsilon_{kt}=n_k^{-1}\sum_i\varepsilon_{kit}$; weak dependence across time points within a series is allowed, and we write the cross-sectional covariance kernel as $\Omega_{k\ell}=\mathrm{Cov}(\varepsilon_k,\varepsilon_\ell\mid\mathcal F)$, with $\Omega_{k\ell}=0$ for $k\neq\ell$ under cross-sectional independence. We use the linear representation that pushes the estimation errors of the weights and residualization into the remainder,
where $\bar\varepsilon_{k,\mathcal T_1}=|\mathcal T_1|^{-1}\sum_{t\in\mathcal T_1}\bar\varepsilon_{kt}$ and $\sigma_g^2:=\mathrm{Var}(u_g)$. The first term is specific to treated group $g$, but the donor-side shock in the second term is shared across all treated groups, and here lies the source of between-group dependence. $R_g$ is the remainder arising from the weights, residualization, and balancing error, and is higher order by orthogonality (Lemma (ref)) and group-wise cross-fitting.
The joint covariance of the group effects follows from this representation by bilinearity, for any cross-sectional dependence structure.
For the baseline joint central limit theorem we now specialize to cross-sectional independence, $\Omega_{k\ell}=0$ for $k\neq\ell$; the dependence between the estimated group effects then arises solely through the shared donor coefficients $a_{gk}$, $k\in\mathcal{C}$, and $\Sigma^\tau$ takes the explicit donor-sharing form of Theorem (ref). Corollary (ref) below extends the central limit theorem to block-dependent shocks, the case used in the Type II application of Section (ref).
\noindentWhich regime applies. The main central limit theorem is a many-treated, many-donor approximation. When the donor pool is small we retain the donor shocks as finite-rank common components rather than averaging them away (Corollary (ref)). This is the regime used for the Clean Air application, where $G_0=23$ and the donor-leverage diagnostics of Section (ref) indicate that donor dilution is not a good approximation. The Medicaid application has only $G_1=25$ treated states; there we report finite-sample diagnostics, leave-one-state sensitivity, and design-based rather than super-population interpretations.
The proofs of Theorem (ref), Corollaries (ref) and (ref), and the theorem in online Appendix H are collected in online Appendix B.
Corollary (ref) is the regime of the Clean Air Act application, where the donor pool is $G_0=23$ state aggregates against $G_1=177$ treated counties. There the large-$G_0$ dilution condition of Assumption (ref) is not in force---the donor-leverage diagnostic of Section (ref) reports $G_1/G_0=7.70$ and $\max_h L_h/\sqrt{G_1}\approx 1.08$ with $L_h=\sum_c\widehat\omega_{ch}$---so we rely on this fixed-donor limit rather than on donor dilution, and the state- and division-block standard errors are exactly the implementation of the finite-rank common component in (ref) together with the regional correlation among the $v_h$.
\parNext, we separate the estimation-error component $\Lambda_W$ contained in the variance from the true heterogeneity component (design vs.\ sampling). Theorem (ref) is inference conditional on the groups, and $\Omega_\gamma$ contains the estimation-error component $\Lambda_W$ arising from the sampling on the treated and donor sides (including between-group dependence via shared donors) (fixed residuals do not contribute). This is the appropriate variance when the estimand is “$\gamma^{*}$ specific to the observed treated groups.”
If instead one adopts a super-population estimand $\gamma^\infty=\mathbb{E}[WW^{\!\top}]^{-1}\mathbb{E}[W\tau]$, the projection residual becomes stochastic and an Eicker--White between-group term is added to the variance, $\Omega_\gamma^{\text{super-pop}}=(M_W^\infty)^{-1}(\Lambda_W+\Psi_{\text{between}})(M_W^\infty)^{-1}$; this is the design-based versus sampling-based distinction AAIW2020,AAIW2023. We report the finite-population variance (using $\Lambda_W$ only) as primary---consistent with not imposing random effects---and give the super-population version in online Appendix F.
The null hypothesis $H_0:V=0$ (all $\tau_g$ equal, no group heterogeneity) is, in applications, the first gate---“is the between-group difference just noise?” The naive statistic
where $\widehat\Sigma^{H}$ is the plug-in covariance kernel of Section (ref) for half-sample $H\in\{A,B\}$ (the variance formula follows immediately from $\widehat V=G_1^{-1}(H\widehat\tau^{A})^{\!\top}(H\widehat\tau^{B})$ and $A\perp B$), and by the boundary $V\ge0$ we construct a one-sided test that rejects when $T_V$ is large. $\widehat{\mathrm{se}}_0$ permits both the heteroskedasticity of $\sigma_g^2$ and donor-sharing dependence.
However, the validity of (ref) depends critically on Assumption (ref)(b). Let us write the half-sample estimator as $\widehat\tau^{H}_g=\tau_g+b_g+e^{H}_g$. Here $b_g$ is the approximation error arising from the limits of weight estimation (regularization and finiteness of the donor pool), and enters commonly in both half-samples, conditional on the realized factor path. Unlike the sampling noise $e^H_g$, $b_g$ does not cancel in the cross product, so even under $H_0$, \[ \mathbb{E}\bigl[\widehat V\,\big|\,\text{design}\bigr]\;\approx\;\frac{1}{G_1}\|Hb\|^{2}\;\ge\;0 \] and since $b$ further varies with the realization of the factor path, the denominator (ref) that counts only sampling noise underestimates the variability of $\widehat V$. That is, the effective null of the naive test is “$V+\mathrm{Var}_g(b_g)=0$,” and it over-rejects when balance is imperfect (Section (ref): measured 28--34% against a nominal 5%). This distortion is not resolved no matter how much $\widehat{\mathrm{se}}_0$ is made robust.
We therefore take as the main test an affine-calibrated leave-out test that extends the placebo inference of synthetic control Abadie2010,Arkhangelsky2021 to the variance decomposition. For $b=1,\dots,B$, we randomly draw without replacement a set $J_b^{p}$ of size $G_p=\lfloor G_0/2\rfloor$ from the donor set $\mathcal{C}$ as a pseudo-treatment set, use the rest as the donor pool, and apply a pipeline exactly identical to the actual one (all-group $A/B$ splitting, group-wise orthogonalized SDID, (ref)) to obtain $\widehat V^{p}_b=\widehat V(J_b^p)$. For a general candidate set $J$ ($|J|=m$), with the centering matrix $M_m=I_m-m^{-1}\mathbf{1}\mathbf{1}^{\!\top}$ ($M_{G_1}=H$), we write $\widehat V(J)=m^{-1}(M_m\widehat\tau^{A}(J))^{\!\top}(M_m\widehat\tau^{B}(J))$. Since the true effect of the pseudo-treatment groups is identically zero, $\{\widehat V^{p}_b\}$ are draws from the $H_0$ distribution conditional on the realized factor path, automatically including the bias and variability arising from the equilibration residual $b_g$. The test statistic incorporates the imbalance in group counts, the finite-population correction of sampling without replacement, and the finite repetitions $B$:
where $\widehat V_{\mathcal{T}}=\widehat V(\mathcal{T})$. The correction factor $\kappa_G$ arises because the variances of the numerator and denominator are both of order $G^{-1}$ (online Appendix D), and for the baseline design $G_1\simeq G_0,\ G_p\simeq G_0/2$ it becomes $\kappa_G\simeq1$---the recentered statistic of the current Monte Carlo (Section (ref)) coincides with this special case ($\kappa_G=1$, excluding $B^{-1}$).
We emphasize that this is an empirical calibration of the null dispersion of $\widehat V$, not a randomization or permutation test: the donor pseudo-treatments are used to estimate the sampling variability of the second moment under $H_0$, and validity comes from the conditional-exchangeability condition rather than from exchangeability of treatment assignment.
The proof is in online Appendix D (a one-paragraph sketch opens the appendix).
The coefficient $\kappa_G=(G_p/G_1)/\{1-G_p/G_0\}$ converts the dispersion of placebo subsets (drawn without replacement from the donor pool) into the dispersion of the treated set, and the finite number of repetitions $B$ adds a $1/B$ term to the reference variance; the finite-$B$ behavior, the full-enumeration variant, and the resulting asymptotic pivotality of $T^{\mathrm{pc}}_V$ are developed in online Appendix D. A super-population version of the calibration, in which the donor pool is itself sampled, is given in online Appendix F.
Complementary to the omnibus test $H_0:V=0$, we define a heterogeneity test for a pre-specified direction. For the covariance $C=G_1^{-1}\sum_{g\in\mathcal{T}}\widetilde W_g(\tau_g-\bar\tau)$ with respect to a scalar-standardized group covariate $\widetilde W_g$, we test $H_0:C=0$ (equivalent to zero projection slope; also $C\neq0\Rightarrow V_{\mathrm{expl}}>0\Rightarrow V>0$) with the Wald statistic $T_C=\widehat C/\widehat{\mathrm{se}}(\widehat C)$ based on the joint covariance kernel. The variance is constructed from the $\widehat\Sigma^\tau$ of Theorem (ref) (including the shared-donor off-diagonal) through the linear representation $G_1^{-1}\sum_g\widetilde W_g\,\widehat e_g$ of $\widehat C$, superimposing state and regional blocks as needed. Because it is a centered contrast, the common shock due to shared donors cancels to first order (isomorphic to the slope SE of Section (ref)), and it is more robust than level quantities. The practical division of labor is clear: in noisy aggregated county data (Type II-CD), the omnibus test tends to be limited in information, but the pre-specified $T_C$ can detect the directional component stably---in the environmental application of Section (ref) we take this as the main test, reporting the calibrated test and the rank $p$-value of the donor-state placebo alongside. Note that when $T_C$ is significant and the omnibus is not, the correct reading is not “$V=0$ is true” but “the omnibus test lacks power in finite samples.” The omnibus $H_0:V=0$ and the directed $H_0:C=0$ are two distinct pre-specified hypotheses, not a search over many directions; when a single moderator $W$ is fixed in advance (as in both applications), no multiplicity correction is needed, whereas an analyst screening several moderators should control family-wise error via the max-$t$ statistic of Section (ref).
The directed test is valid for a single pre-specified moderator. When several moderators are screened, we control family-wise error with the max-$t$ statistic $T_{\max}=\max_j|\widehat C_j|/\widehat{\mathrm{se}}(\widehat C_j)$ whose null quantile is read from the joint covariance kernel; in both applications a single moderator is fixed in advance, so no adjustment is needed.
The $\Lambda_W$ of Theorem (ref) contains the cross term $\sum_{h}\omega_h^{(g)}\omega_h^{(g')}\upsilon_h^{(g,g')}$ via shared donors. We provide the following two implementations.
The joint covariance kernel can be estimated in two ways. In the first, the plug-in covariance kernel, for each group $k\in\mathcal{T}\cup\mathcal{C}$ we estimate the across-time covariance matrix of the within-group deviations $\widehat\varepsilon_{kit}=Y^{\perp}_{kit}-\bar Y^{\perp}_{kt}$ across units, and multiply by $n_k^{-1}$ to obtain $\widehat\Xi_k\approx\mathrm{Cov}\bigl((\bar\varepsilon_{kt})_{t}\bigr)$ (within-group serial correlation is automatically incorporated as across-time covariance). Using the time contrast $c^{(g)}$ ($|\mathcal T_1|^{-1}$ for $t\in\mathcal T_1$, $-\widehat\lambda_t^{(g)}$ for $t\in\mathcal T_0$), $\widehat\sigma_g^2=c^{(g)\top}\widehat\Xi_g c^{(g)}$, $\widehat\upsilon_h^{(g,g')}=c^{(g)\top}\widehat\Xi_h c^{(g')}$ we set these and substitute $\widehat\Sigma_{gg'}$ into the sandwich form to obtain $\widehat\Omega_\gamma,\widehat\Omega_V$, etc.
In the second, the multiplier bootstrap, the independent unit of the primitive shock is the group (treated group $g\in\mathcal{T}$ and donor $h\in\mathcal{C}$). By multiplying each group $k$ by an independent multiplier $\eta_k$ ($\mathbb{E}\eta_k=0$, $\mathrm{Var}\,\eta_k=1$) and reconstructing the influence representation (ref), we can approximate the joint distribution of $(\widehat\gamma,\widehat V,\widehat C,\widehat R^2)$ while preserving donor-sharing dependence.
In this section we verify the claims of the preceding sections in finite samples. The implementation is in Python (fixed random seeds, $2{,}990$ replications in total), and the code, raw data, and aggregates are included in the replication package.
The baseline design is $G_1=G_0=50$, $T_0=20$, $|\mathcal T_1|=5$. There are two latent factors (a linear-trend factor and a stationary AR factor), $Y_{git}(0)=\alpha_g+\xi_t+\Gamma_g^{\!\top} F_t+X_{git}^{\!\top}\beta+\varepsilon_{git}$ ($\varepsilon_{git}$ is within-group AR(1)), group covariates are $W_g\in\mathbb{R}^{2}$ plus an intercept, and $\tau_g=\gamma_0+W_g^{\!\top}\gamma+\nu_g$ (by default $V=\mathrm{Var}(\tau_g)=0.5$, $R^2=0.5$). Treatment is determined by group-level selection, the top $G_1$ groups by the latent score $a_1 z_{W,g}+a_2 z_{\Gamma,g}+(\text{logistic disturbance})$, and $a_2\neq0$ generates factor selection (parallel-trends violation). Main settings: FA0 (parallel trends hold: $a_2=0$, $\Gamma\perp W$, low noise $n_g=24$, $\sigma_\varepsilon=1.2$), FA2 (strong factor selection: $a_2=2$, $\mathrm{corr}(\Gamma,W)=0.5$), NB (high-noise baseline: $n_g=16$, $\sigma_\varepsilon=2.5$, $a_2=0$), NULL/NULL0 ($\tau_g\equiv$ constant, $V=0$; NB-type and FA0-type), POW1--3 ($V\in\{0.05,0.15,0.40\}$), \textbf{FSTR} (heteroskedastic $\sigma_{\varepsilon,g}\in\{1.2,3.5\}$ and $\mathrm{corr}(\nu_g,\text{group precision})\approx0.7$), \textbf{LEAK} (donor concentration: $G_0=30$, regularization $\times0.3$, effective number of donors $5.6$). $R=100$--$400$ replications per setting, and placebo calibration uses $B=59$. The comparison methods are DiD (equal weights), DR-DiD (group-level AIPW: outcome regression on $W$ plus propensity-score odds weighting), the proposal (orthogonalized SDID), and, as a random-effects benchmark for the second stage, FIRC-type precision-weighted GLS. An overview of the group-wise estimates in a representative replication is shown in the figure in online Appendix H.
As shown in online Appendix H, under FA0 where parallel trends hold, all three methods are unbiased, but SDID reduces the RMSE by about a factor of five because it matches the factor path group by group. Under FA2 with factor selection, DiD breaks down substantially in both level and slope, and DR-DiD corrects only the slope component explained by $W$ while the level remains (the loading gap orthogonal to $W$ cannot be removed by covariate conditioning). Orthogonalized SDID reduces both by about 90%. The residual (intercept $+0.083$) arises from the limit of balancing due to the regularization $\zeta$, and we make it explicit as a finite-sample implication of Assumption (ref)(b).
As shown in online Appendix H, the plug-in $\widehat V$ counts the estimation-error variance directly and is 54--75% too large relative to $V$, whereas leave-out removes 60--90% of this regardless of the splitting scheme. There are two additional findings. First, the linear donor-sharing leakage of treated-only splitting (the $G_1^{-1}\mathrm{tr}(HKH)$ of the remark in online Appendix B) was second-order at $0.002$--$0.008$ in the plug-in evaluation---because centering absorbs the common donor shock. Second, the remaining small positive bias is common to both splitting schemes and reduces to the between-group variance of the equilibration residual $b_g$ (the mean of $\widehat V$ in a $V=0$ design is $+0.024$ for the good design NULL0 and $+0.069$ for the high-noise NULL; the all-group split side is slightly larger because weight estimation degrades on half-samples). This is a direct quantification of the fact that Assumption (ref)(b) is binding in finite samples.
As shown in online Appendix H, for the intercept the independence (diagonal-kernel) SE gives 11--16pp undercoverage, and the joint kernel of Theorem (ref) recovers near the nominal level (NULL 94.3%, POW2 95.4%, LEAK 95.5%). For the slope, in this design where the weight profiles are similar across groups, the cross term cancels through the centering of $\widetilde W_g$, and the two SEs almost coincide---consistent with the theoretical implication that shared-donor dependence mainly distorts the inference of levels. The slope undercoverage in the NB-type settings (74--85%) arises from first-stage bias rather than variance, and in NULL0 where the bias vanishes both SEs achieve about 92%.
As shown in online Appendix H, the bare $T_V$ is severely distorted with size 28--34% (nominal 5%). As discussed in Section (ref), the cause is not the bias of $\widehat V$ (small at $+0.024$ in NULL0) but that factor mismatch generates variability not contained in the denominator (mean of the statistic $1.67$, standard deviation $3.4$ in the good design; the figure in online Appendix H). The affine-calibrated statistic (ref) recovers the size to nearly the nominal level at $6.5\%$ and $4.1\%$ (Figure (ref)), and this is retained even in the NB-type (PN) settings where selection shifts the loadings. Since this design has $G_1=G_0=50,G_p=25$, the correction factor is $\kappa_G=(25/50)/(1-25/50)=1$, and the measured values coincide with the special case of Theorem (ref) (excluding $B^{-1}=1/59$)---that is, the theorem shows that the calibration of the current MC is a consequence of the theoretical correction rather than working by chance. The permutation $p$-value version improves to $12.4\%$ and $10.6\%$, but an overshoot due to the group-count scale difference ($G_1=50$ vs.\ $G_p=25$) remains (the remark in online Appendix D). Under honest size, the power at $V=0.05/0.15/0.39$ is $12\%/52\%/86\%$.
The per-group caterpillar plot and the null-distribution diagnostic are in online Appendix H.
In FSTR ($\mathrm{corr}(\nu_g,\text{precision})\approx0.7$), the intercept bias of FIRC-type precision-weighted GLS was $+0.109$, whereas the proposal's equal-weight projection was $-0.002$ (the slopes are comparable: $+0.061$ vs.\ $+0.064$). When exchangeability breaks, precision weighting systematically distorts the level---a numerical corroboration of the stance of Section (ref).
To avoid over-fitting to the baseline design $(G_1,G_0,T_0)=(50,50,20)$, we vary each design axis a referee would probe---the group-count ratio $G_1/G_0$, the observation length $T_0$, the $R^2$ boundary, alternative $V$ estimators, calibration failure modes, and the number of placebo repetitions $B$. The qualitative conclusions are unchanged: the orthogonalized first stage removes first-stage bias, the analytic and leave-out corrections remove the upward bias of $\widehat V$, the donor-sharing kernel restores coverage of $\gamma$, and the affine-calibrated placebo restores the size of the $V=0$ test. Full tables with Monte Carlo standard errors appear in online Appendix H.
This is the paper's principal Type I application. The primary moderator is pre-specified: the pre-expansion uninsured rate for Medicaid and baseline fine-particulate pollution for Clean Air. The $m_{\mathcal{T}}(w)$ analysis is confirmatory for this primary moderator; any secondary moderators are exploratory, and multiple-moderator searches use a $\max$-$t$ adjustment reported in the online supplement. The 2014 ACA Medicaid expansion is a common-cohort policy shock whose average coverage effects are well documented Courtemanche2017,MillerJohnsonWherry2021; our object is the structure of the heterogeneity of those effects across states. Because the ACS provides household microdata, the household $A/B$ leave-out validates the variance decomposition against an independent analytic correction---the feature that distinguishes the Type I regime.
The outcome is the uninsured rate of low-income adults (ages 19--64, income below 138% FPL) constructed from the ACS 1-Year PUMS 2008--2019 ACSPUMS ($4.97$ million person--year records); the groups are states and the units are households, split within each state--year cell into $A/B$ halves. We adopt common-timing SDID ($\tau_0=2014$): the treated set is the $G_1=25$ states (with the District of Columbia) that expanded Medicaid in January 2014, and the donor pool is the $G_0=17$ states that had not expanded by the end of the window (expansion status as compiled by KFF2024). The estimand is not the effect of the ACA relative to no ACA---the exchanges, subsidies, and mandate took effect nationwide in 2014---but the incremental effect of Medicaid-expansion status relative to non-expansion states exposed to those components. The group covariate $W_g$ is the pre-expansion (2013) uninsured rate, the “bite” of the expansion. Early- and late-2014 through 2019 expanders outside this clean design are excluded; further detail is in online Appendix J.
The figure in online Appendix J shows the treated-group means and SDID synthetic controls for four representative states. The two series overlap well over the six pre-treatment years (the pre-treatment RMSPE across the 25 states has median 1.1pp and maximum 3.0pp), and after $\tau_0$ large divergences arise in states with high pre-expansion uninsured rates (KY, WV, CA). The figure in online Appendix J is a caterpillar plot of the group ATTs $\widehat\tau_g$, contrasting the joint-kernel SE with the independence (diagonal) approximation. $\widehat\tau_g$ ranges from $-18.1$pp in KY to small values (which can be positive near the floor) in states that already had low pre-expansion uninsured rates (MA, DC, VT), with a mean ATT of $-6.2$pp.
The central finding of this application is the bite slope (Figure (ref)). States with higher pre-treatment uninsured rates have larger reductions in the uninsured rate, and the projection coefficient onto standardized bite is $\widehat\gamma_{\text{bite}}=-5.86$pp (joint-kernel SE 0.23, $t=-25.2$)---a state with a pre-expansion uninsured rate one standard deviation higher has a reduction 5.9pp larger. The bite gradient is large and robust, and is strongest on the percentage-point scale, where the point estimate is $\widehat R^2=0.826$ (rounded to $0.83$). We report the delta-method interval $[0.75,0.90]$, computed from the noise-corrected variance components, as the primary confidence set for the explained share; because the denominator is well separated from zero ($t_V\approx13$), the Fieller/test-inversion interval agrees to two decimals. The explained share is a finite-population, percentage-point-scale descriptive summary, not a scale-invariant causal constant, and the inference below is based on $(V,V_{\mathrm{expl}},C)$. The explained share attenuates to $0.72$ on the relative-reduction scale and $0.40$ on the log-odds scale; that it persists after attenuation is evidence that it is not an artifact of mechanical room for improvement (Section (ref)). That the effect is nearly zero in states with low pre-expansion uninsured rates ($=$ small room) is not a failure of extrapolation but the reality of the bite (these states also have good pre-treatment fit; Section (ref)).
The effective number of donors per group is 10.4. For individual groups $\widehat\tau_g$, the joint-kernel SE exceeds the independence approximation (e.g., 0.44 vs.\ 0.29 in CA, $+52\%$), and for the average effect the difference is even more pronounced: the joint-kernel SE 0.44pp is 1.83 times the independence SE 0.26pp---the independence assumption overestimates precision by about 45%. This is because the shared-donor disturbance acts as a common factor, inflating the variance for level quantities like the mean and intercept while nearly cancelling for a mean-zero contrast (the bite slope) (the joint-kernel SE 0.23 of the slope is nearly equal to the diagonal 0.25). Restricting donors to the 12 always-non-expansion states strengthens the shared dependence further (SE ratio 1.92; Section (ref)).
The heterogeneity variance is estimated twice, independently (Figure (ref), left).
Against the plug-in $\widehat V_{\mathrm{plug}}=43.0$ (pp$^2\!\times\!10^{-4}$), an analytic noise subtraction using the household-cluster sampling variance and a household $A/B$ leave-out give nearly identical corrected values (about $34$--$35$), confirming that the between-state variation is real heterogeneity rather than estimation noise. The two routes rest on different assumptions---the analytic route on the survey variance, the leave-out route on within-cell independence (online Appendix C)---so their agreement is a strong internal check unavailable without micro-replication. On the same percentage-point scale the explained share is $\widehat R^2=0.83$, and donor sharing inflates the average-effect standard error to $1.8$ times its independence-based value while leaving the centered bite slope nearly unchanged.
We first conduct the test of no heterogeneity. Here the heterogeneity is real and large. In addition to the bare statistic $T_V=13.0$, the placebo-calibrated statistic is $T_V^{\mathrm{pc}}=24.7$ (correction factor $\kappa_G=0.60$) with rank $p$-value 0.005, strongly rejecting $V=0$ (Figure (ref), right). Furthermore, as a robustness check on the definition of the explained share, the uncorrected raw $R^2$ is $0.80$ with test-inversion Fieller set $[0.70,0.91]$ (online Appendix E; grid inversion, leave-one-donor jackknife variance) and the analytic-noise-corrected $R^2$ is $0.87$ with Fieller set $[0.76,1.00]$, so these bracket the noise-corrected headline $0.826$; the upper bound 0.91 $<1$ of the uncorrected set implies that residual heterogeneity not explained by bite is real (not sampling noise)---suggesting other factors such as implementation, prior Medicaid generosity, and take-up.
We address the donor-halving confound with a pool-size-matched placebo. In the placebo test, because the synthetic control of the pseudo-treatment units is constructed from fewer donors than the actual treatment, the dispersion of the equilibration residual can change mechanically (a confound pointed out in review). Against this, we implemented a pool-size-matched placebo: in each replication we draw $G_p=8$ pseudo-treatment donors and recompute the actual-treatment $\widehat V$ on the same residual donor sub-pool (size $m=9$) for comparison. The figure in online Appendix J is the result. The pseudo-treatment null distribution $\widehat V^p$ (mean 2.6) and the actual-treatment $\widehat V_T^m$ recomputed on the same-size sub-pool (mean 43.2; nearly unchanged from the full-donor-pool value 41.6) are completely separated ($T_{\mathrm{PM}}=634$, rank $p=0.001$). Even with matched pool sizes, the effect variance of the actual treatment overwhelms the pseudo, and the rejection of $V=0$ is not a product of donor halving. With $G_0=17,G_p=8$, all $\binom{17}{8}=24{,}310$ pseudo-assignments can be enumerated, and the rank $p$-value is based on that conditional reference distribution.
The projected group-effect curve $m_{\mathcal{T}}(w)=w^{\!\top}\gamma^{*}$, its confidence band, the between-state variance, and the explained share $R^2$ with a valid confidence interval are delivered only by the proposed joint donor-sharing kernel; two-way fixed effects, difference-in-differences, doubly robust DiD, and precision-weighted GLS may reproduce the average effect and even the point estimate of the bite slope, but none provides correct standard errors or calibrated heterogeneity tests for these objects, because they ignore the cross-sectional dependence induced by shared donors and the estimation-noise bias in the second moments. The comparison table and figure are in online Appendix J.
We check robustness to specification changes. The figure in online Appendix J shows the robustness of the bite slope and $R^2$. Under the core 12 donors, WI excluded, ridge $\zeta$ ($\times0.5$, $\times2$), pre-treatment window (2010--2013), long-term post (adding 2021--2024), and top-fit only (excluding the worst quartile of pre-treatment RMSPE), $\widehat\gamma_{\text{bite}}\in[-6.5,-5.6]$pp, $\widehat R^2\in[0.83,0.87]$, all stable with $t<-21$, and the agreement of leave-out and analytic $\widehat V$ is preserved. The only case that behaves differently is the level convex-hull restriction (excluding the 8 states whose pre-treatment uninsured rate falls below the donor level range), where $\widehat R^2$ drops to 0.46, but this is because excluding low-bite states (which anchor the low end of the slope) mechanically compresses the range of bite---the excluded states have good pre-treatment fit and their exclusion is unjustified on fit grounds---and even so the slope is significant with $t=-9.0$. In short, the bite slope is significant in all specifications, and its magnitude depends on whether low-bite states are included.
Finally, we verify that $R^2=0.83$ is not an artifact of mechanical explanation. We successively exclude, through five checks, the possibility that the bite slope is an artifact of common sampling error, scale dependence, pre-treatment contamination of early-expansion states, or mere “room for improvement” (the table in online Appendix J).
First, since $W_g$ (pre-expansion uninsured rate) is estimated from the same ACS sample as $\widehat\tau_g$, there is a concern that common sampling error inflates the negative slope. Against this, constructing $W_g$ from the A half-sample and $\widehat\tau_g$ from the B half-sample (and averaging over swaps) and re-estimating the slope under independent sampling error gives $\widehat\gamma_{\text{bite}}=-6.00$pp ($R^2=0.79$), nearly unchanged from the baseline ($-5.86$pp, 0.83), so the contribution of common sampling noise is negligible. Second, because the projection could be driven by a few high-coverage floor states with small positive $\widehat\tau_g$ (Massachusetts $+6.2$pp, DC $+4.1$, Vermont $+2.1$; the figure in online Appendix J), we re-estimate the slope excluding them: dropping Massachusetts gives $\widehat\gamma_{\text{bite}}=-5.4$pp and dropping Massachusetts, DC, and Vermont gives $-4.5$pp; the plug-in explained share moves from $0.80$ to $0.76$ and $0.68$ in the same unweighted projection (the headline $R^2=0.83$ is the noise-corrected $V_{\mathrm{expl}}/V$). The slope and the significant heterogeneity survive, and the attenuation quantifies how much of the descriptive share rides on the low-bite floor.
The remaining checks---split-sample stability, exclusion of the 2010--2013 early expanders, the secondary-outcome mechanism (Medicaid/public coverage gains mirror the uninsured reductions), and nonlinearity of the bite relation---leave the slope and the significant heterogeneity intact; the full six-point list with all numbers is reproduced in online Appendix J.
Finally, we report bias-aware intervals that allow the equilibration residual $b_g$ ($|b_g|\le\kappa\,\mathrm{RMSPE}_g$) representing the incompleteness of counterfactual construction. At $\kappa=1$, where the residual is bounded by the pre-treatment RMSPE, the 95% interval of $\widehat\gamma_{\text{bite}}$ is $[-7.28,-4.45]$, and even at the looser $\kappa=2$ it is $[-8.23,-3.49]$, excluding 0, and the bias-aware bracket of $V$ at $\kappa=2$ is $[12.8,86.7]$, excluding $V=0$. Even allowing the equilibration residual, the existence of the bite slope and the effect heterogeneity is robust.
In summary, this application exercises all three pillars and clarifies its unique value in comparison with existing methods. The object displayed in Figure (ref) and Table (ref) is the projected group-effect curve $m_{\mathcal{T}}(w)$, not merely the slope coefficient: it answers the policy question of what uninsured-rate reduction to expect for a state with a given pre-expansion uninsured rate. A high-bite state (75th percentile) is predicted to see a $7.8$ pp larger reduction than a low-bite state (25th percentile), with a $95\%$ confidence interval $[-8.4,-7.2]$ that reflects donor sharing; treating the state-specific estimates as independent would give a different uncertainty statement. We stress that $R^2=0.83$ is not a universal explanatory share: it is the percentage-point-scale explained share for this finite population of states, and it attenuates on the relative and log-odds scales. First, it estimated an interpretable and strong heterogeneity slope (bite; $\widehat\gamma_{\text{bite}}=-5.9$pp, $R^2=0.83$ on the percentage-point scale) as a projection. Second, it showed that donor-sharing dependence enlarges the SE of aggregate quantities to 1.8 times the independence assumption. Third, the heterogeneity variance from household $A/B$ leave-out agreed with the analytic estimate, independently verifying that the heterogeneity is real---this is possible only with Type I micro-replication and is absent from any of TWFE/DiD/standard SCM. Correct inference for the heterogeneity objects---valid $R^2$ confidence intervals, the projected group-effect curve and its confidence band, shared-dependence standard errors, and leave-out verification---is obtained only with the proposed method; that the point estimates happen to agree with existing regressions is incidental and does not confer valid uncertainty quantification on them.
A high-resolution version aggregating the individual-level analysis to counties (1{,}097 counties, SAHIE) is given in online Appendix C: the aggregation invariance that the state aggregation of county effects coincides with the state-level estimate at correlation 1.00 (the proposition in online Appendix C), the three-column standard errors of county-independent, state-block, and division-block (up to 1.7 times the independence assumption), the measurement-error propagation of the published MOE, and the auxiliary specifications (aggregation consistency, contextual effect, and between--within decomposition).
The second application is an implementation on aggregated county public data without micro-replication (Type II-CD). Here household $A/B$ leave-out is unavailable, and the basis of inference is placed on the integration of analytic noise correction, shared-donor dependence, and state/regional-block dependence, together with the pre-specified directed test $H_0:C=0$ of Section (ref). Environmental regulation has a clear prior direction---“it works more strongly where the standard is exceeded more”---making it a good place for a directed test. Because the county-level PM$_{2.5}$ effect is noisy, the omnibus $V=0$ test has limited power against diffuse heterogeneity---whereas the $H_0:C=0$ test along the pre-specified direction of the baseline pollution level has high power against a policy-relevant alternative. This is the design principle of this section. Note that the outcome is the PM$_{2.5}$ concentration itself, that is, the first-stage environmental response to regulation, and we present this section as an environmental-policy / environmental-exposure application (the bridge to health outcomes is described at the end of the section).
The treatment is the 2005 nonattainment designation under the 1997 PM$_{2.5}$ NAAQS EPA2005,EPAGreenBook; such designations have long served as quasi-experimental variation in environmental economics ChayGreenstone2005. The treated set is the $177$ counties designated whole-county nonattainment, and the donor pool is the $23$ states with no designated county, aggregated to state-level PM$_{2.5}$ series from satellite-derived county concentrations vanDonkelaar2019,Wu2020; the pre period is 2000--2004 and the post period 2006--2012. The group covariate $W_c$ is the county's baseline (pre-period) PM$_{2.5}$ level, an empirical proxy for the regulatory bite with the monotone property that higher $W_c$ means a larger reduction obligation. The county-wise orthogonalized SDID fits the pre-period well (median RMSPE $0.33\,\mu$g/m$^3$, about $2.6\%$ of the baseline level). Full institutional detail is in online Appendix I. Donor-leverage diagnostics confirm that this is not a donor-dilution regime: $G_1/G_0=7.70$ and $\max_h L_h/\sqrt{G_1}=1.08$, so inference is based on the fixed-donor/common-shock approximation (Corollary (ref)) rather than the many-donor dilution approximation.
Because the donor pool is small ($G_0=23$ state aggregates against $G_1=177$ treated counties), we first check which asymptotic regime governs the inference (online Appendix I). The donor-sharing ratio is $G_1/G_0=7.7$, and the maximum donor leverage $L_h=\sum_{c}\widehat\omega_{ch}$ satisfies $\max_h L_h/\sqrt{G_1}=1.08$---the large-$G_0$ dilution condition of Assumption (ref) is therefore not in force, and we rely on the fixed-donor finite-rank limit of Corollary (ref) rather than on donor dilution. At the same time the counterfactual support is broad: the effective donor count per county (inverse Herfindahl of the weights) is $11.6$, the average largest-donor share is only $0.16$, and just $0.6\%$ of counties place more than half their weight on a single donor. The common donor shock is thus a finite-rank object with well-spread loadings, exactly the case in which the state- and division-block covariances of Proposition (ref) estimate the non-vanishing component $\sum_{h,\ell}A_hA_\ell\Omega_{h\ell}$ of (ref). The directed conclusion is unchanged under the baseline-matched county-donor design (Table (ref)).
The central number of this subsection is the directed slope, and we first organize its estimand (Table (ref)). We do not interpret the uncalibrated raw slope as a county-level causal effect: the Clean Air application illustrates Type II-CD inference in aggregate public data, where directed contrasts must be read relative to placebo convergence benchmarks. The main-specification slope $-0.372$/SD is the total slope that includes the broad convergence of non-designated regions; the excess component obtained by calibrating and subtracting this convergence with a donor-state placebo is $-0.255$/SD, and the value that the baseline-matched county-level donor gives by constructively deducting the convergence is $-0.206$/SD. The directed component attributable to regulation is honestly read as a band of $-0.21$ to $-0.26$/SD, and the robustness analysis below (the table in online Appendix I) corroborates this band.
The timing of the directed effect is also informative (Figure (ref)). Tracing the event-time directed slope $\widehat C(e)=G_1^{-1}\sum_c(W_c-\bar W)\,\widehat\tau_c(e)$ across post years $e=2006,\dots,2012$, the effect is weak immediately after the 2005 designation ($-0.16$ in 2006, $t=-2.1$), strengthens as State Implementation Plans and attainment obligations bind ($-0.43$ to $-0.48$ over 2008--2011, $|t|$ up to $4.6$), and attenuates by 2012 ($-0.35$)---a profile consistent with a regulation that acts with a multi-year lag rather than an instantaneous jump, and whose seven-year average is the $-0.372$ reported above.
Figure (ref)(A) shows the starting point of this design: the county-wise $\widehat\tau_c$ scatter widely around a mean of $-1.13\,\mu$g/m$^3$ ($V=0.52$), and since the pointwise analytic noise is not small, it is hard to read “where it worked” from the list of group-wise estimates alone. In contrast, the directed projection of Figure (ref)(B) is clear: a county with baseline PM$_{2.5}$ one SD higher has a reduction $0.372\,\mu$g/m$^3$ larger.
Table (ref) summarizes the second-stage inference, and three points stand out. First, for level quantities such as the mean ATT the dependence structure is decisive: relative to the county-independent standard error, the state block is $2.7$ times and the division block $4.4$ times larger (Proposition (ref), Corollary (ref)). Second, the omnibus $H_0:V=0$ test is low-powered here---the group-wise estimates are individually noisy---so a non-rejection is uninformative. Third, and by contrast, the pre-specified directed statistic on baseline PM$_{2.5}$ is stable and informative, because it is a centered contrast in which the common donor shock cancels to first order (as in the Medicaid slope, Section (ref)). The directed slope is the scientifically meaningful quantity in this Type II-CD application. Equivalently, this directed slope traces a projected group-effect curve $m_{\mathcal{T}}(w)$ over baseline PM$_{2.5}$: the county-level estimates are too noisy to read one by one, but $m_{\mathcal{T}}(w)$ summarizes where the regulation was most effective. A county at the 75th percentile of baseline PM$_{2.5}$ is predicted to experience about $0.50\,\mu$g/m$^3$ larger reduction than a county at the 25th percentile (shared state-block $95\%$ CI $[-0.74,-0.26]$); the placebo-calibrated excess component of this contrast, net of regional convergence, is about $60$--$70\%$ of it. This curve, reported with state/division-block covariance and compared with donor-state placebo and matched county-donor benchmarks, is invisible to a conventional group-wise SDID that treats the county effects as independent.
The leave-one-donor-state placebo ($K=22$) gives a one-sided rank $p=0.087$, which is suggestive rather than decisive; the main evidence for the directed component comes from the sign-stable slope under state and division block covariance, the pre-trend adjustment, and the baseline-matched county-donor design. The finite-donor diagnostics (Section (ref)) confirm that the inference operates in the fixed-donor regime of Corollary (ref) rather than relying on donor dilution.
A natural concern is mean reversion: high-baseline counties might have declined anyway. Three pieces of evidence separate the regulatory effect from convergence. First, the placebo distribution of the directed slope is itself centered below zero, so we report the placebo-calibrated excess ($-0.26$/SD, Table (ref)) as the component attributable to regulation. Second, a baseline-matched county-donor design---restricting each treated county's donors to attainment counties with similar baseline PM$_{2.5}$---yields a similar directed slope ($-0.21$/SD). Third, the slope is essentially unchanged after adjusting for a linear pre-trend ($-0.365$/SD). The directed component is therefore sign-stable across state-block and division-block covariances and robust to the mean-reversion critique; the full audit is in online Appendix I.
PM$_{2.5}$ concentration is the first-stage environmental response the regulation targets; whether those reductions propagate to mortality is the downstream question. Applying the same Type II-CD machinery to IHME county age-standardized mortality IHME2017, exploratory estimates are directionally consistent for cardiovascular mortality---strengthening in lag-adjusted windows, in line with the event-time profile of the concentration response (Figure (ref))---weaker for ischemic heart disease, and near zero for cerebrovascular mortality. Because the IHME estimates are model-based and spatially smoothed (with smoothing that can correlate with baseline PM$_{2.5}$), we report block standard errors and treat these results as exploratory rather than confirmatory; full tables by cause and post-anchor window are in online Appendix I.
The table in online Appendix I contrasts what conventional analysis and the proposed method each give in this application---an organization isomorphic to the Medicaid comparison (online Appendix I): a conventional group-wise SDID reports a noisy list of county effects and a low-powered omnibus test, whereas the directed projection recovers a stable baseline-pollution gradient.
This paper proposed a framework for inferring the heterogeneity of group-level treatment effects as a single moment system with orthogonalized SDID as the first stage. The paper's central point is that many-group SDID estimates are not independent, because they share donors: once these estimates are used to study effect heterogeneity, the shared-donor covariance, the upward bias of second moments from estimation noise, and persistent mismatch must all be carried into the inference. Consequently the projection coefficients, the projected group-effect curve $m_{\mathcal{T}}(w)=w^{\!\top}\gamma^{*}$ and its confidence band, the between-group variance, and the explained share admit valid inference only through the joint-covariance framework developed here, and not through two-way fixed effects, difference-in-differences, or standard synthetic-control regressions that treat the group effects as independent. Feeding the same $\{\widehat\tau_g\}$ into both the projection (first moment) and the all-group leave-out variance decomposition (second moment), and placing the between-group dependence induced by shared donors at the center, we derived the joint asymptotic distribution of $(\widehat\gamma,\widehat V,\widehat C,\widehat R^2)$. To the aversion to random effects, we responded in a frequentist manner on three points: (i) defining the estimands as fixed quantities, (ii) explicitly decomposing the variance into the estimation-error component and the true heterogeneity component, and (iii) validly constructing the $V=0$ test under unknown, heteroskedastic, and between-group-dependent variance---and, for the equilibration residual, by placebo calibration. The Monte Carlo experiments (Section (ref)) quantitatively confirmed the robustness of the first stage under factor selection (about 90% bias reduction), the removal of variance bias by leave-out (60--90%), the recovery of intercept coverage by the joint covariance kernel (near the nominal level), and the recovery of test size by placebo calibration (28--34% $\to$ near 5%).
On the empirical side, we added an exploratory downstream health outcome (Application IIb) using IHME county mortality to the Clean Air Act application of Section (ref), but this is suggestive, not confirmatory---a systematic exploration of annual panels, finer causes of death, and lag structures, and a full development via a Poisson-offset-type rate contrast based on death counts and population, is a natural extension requiring restricted-use vital statistics. As limitations and extensions remaining on the methodological side, there are: (i) efficiency: a split-free analytic leave-out and data-fission constructions Leiner2025 that avoid the efficiency loss of sample splitting; (ii) weaker calibration assumptions: explicit corrections for systematically inferior treated fit beyond the fit-matched device, and uniformly valid inference in the weak-heterogeneity region where $V$ is near zero; and (iii) staggered adoption and alternative first stages: cohort-wise donor pools with clean comparisons, and matrix-completion, proximal, or Bayesian first stages plugged into the same second-stage moment system. These directions---together with the semiparametric efficiency bound and a cut/modular Bayesian extension---are discussed further in online Appendix G.
\noindentSupplementary Materials.\ The online supplement contains online Appendices A--J (included after the references in this file): the self-contained first-stage theory, the Section 3 proofs, the county aggregation layer, the complete proof of Theorem (ref), the Fieller-type $R^2$ inference, the super-population calibration, extended related work and discussion, the additional Monte Carlo experiments, and the supplementary application tables. The replication package provides a single-command workflow, staged public-data inputs, scripts for all figures and tables, and machine-readable result files used to populate the manuscript tables, together with the implementation of the analytic noise correction, the household $A/B$ leave-out, and the bias-aware intervals. A smoke-test workflow is provided for continuous verification, and every result is reproducible from public data (ACS PUMS, SAHIE, EPA Green Book, satellite-derived county PM$_{2.5}$, and IHME county mortality estimates).
\noindentAcknowledgments.\ We thank all those concerned for helpful discussions.
Funding.\ This research was supported by JSPS KAKENHI Grant Number 25K21651 and the RIKEN Center for Advanced Intelligence Project (AIP).
Disclosure Statement.\ The authors report there are no competing interests to declare.