Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
59,510 characters · 16 sections · 33 citation commands
Bayesian Robustness Values for Modern Causal Panel Estimators via Riesz Representations
{\it Keywords:} omitted variable bias; synthetic difference-in-differences; matrix completion; comparative case study; partial identification.
Causal panel estimators have become central tools in applied causal inference. Synthetic difference-in-differences (SDID) of ArkhangelskyEtAl2021 bridges synthetic control and difference-in-differences, delivering double-robustness against violations of either parallel trends or perfect synthetic-control match. Matrix completion (MC) of AtheyEtAl2021 imputes counterfactual outcomes via nuclear-norm regularization, exploiting low-rank structure in the panel. The group-time average treatment effect framework of CallawaySantAnna2021 and the imputation-based approach of BorusyakJaravelSpiess2024 provide flexible identification under staggered treatment adoption. These estimators are often motivated by latent-factor and interactive-fixed-effect failures of simple two-way fixed effects Bai2009,Xu2017. This paper therefore sits at the intersection of modern DiD heterogeneity work DeChaisemartinDHaultfoeuille2020,GoodmanBacon2021,RothSantAnnaBilinskiPoe2023, synthetic-control extensions DoudchenkoImbens2016,BenMichaelFellerRothstein2021, and sensitivity analysis for observational causal inference RosenbaumRubin1983,ImbensWooldridge2009,Manski1990,VanderWeeleDing2017,FermanPinto2021.
Despite these advances, sensitivity analysis for modern panel estimators remains uneven. SDID and matrix completion still lack an omitted-variable-bias-bound-based robustness-value analysis, while recent work on staggered $ATT(g,t)$ is frequentist rather than posterior predictive. Consider a concrete applied scenario. A researcher estimates the effect of California's tobacco-control program on cigarette consumption using SDID and reports $\widehattau\approx -15.6$ packs per capita, a large negative point estimate. A natural robustness question is whether this conclusion survives omitted state-level health-policy initiatives that correlate with both treatment timing and outcome trajectory. The researcher has no estimator-native answer. CinelliHazlett2020's OLS robustness-value framework does not extend directly to SDID; BachKlaassenKueckMattesSpindler2025's frequentist double-machine-learning extension covers $ATT(g,t)$ but not SDID; and LiuYamamoto2025's Bayesian framework is parameterized in latent-confounder space rather than in omitted-variable-bound and partial-$R^2$ robustness-value units.
This paper develops a sensitivity workflow for this setting. We combine three ingredients. First, the Riesz-representation omitted-variable-bias framework of ChernozhukovEtAl2026 provides a unified machinery for bounding omitted-variable bias of general causal estimands. Its use of fitted representers is connected to double/debiased machine learning and automatic Riesz-representer estimation ChernozhukovEtAl2018,ChernozhukovNeweySingh2022. Second, the partial-$R^2$ robustness-value framework of CinelliHazlett2020,CinelliHazlett2025 provides a scale-free and interpretable benchmarking device. This scale connects to selection-on-observables calibration Imbens2003,Oster2019,Frank2000, weighted-estimator sensitivity WainsteinHazlett2025, relative-correlation and partial-identification approaches Krauth2016,MastenPoirier2018,MastenPoirierZhang2024, endogenous-control sensitivity DiegertMastenPoirier2022, and local-misspecification sensitivity BonhommeWeidner2022. Third, Bayesian inference on the partial-$R^2$ sensitivity parameters provides a natural language for uncertainty quantification when those parameters are not themselves data-identifiable.
The third ingredient requires care. Priors placed directly on the sensitivity parameters are useful for transparent robustness profiling, which we call Route A, but they do not use observed covariate benchmarks as data. Our auxiliary benchmark formulation, Route B, treats observed-covariate partial-$R^2$ pairs as draws from a benchmark population and updates the resulting omitted-variable-bias bound only when diagnostic checks support that modeling step. When the benchmark set is too small, too coarse, poorly aligned with the estimator's Riesz representer, or unable to dominate the hidden-confounder strength distribution, Route B is demoted to an exploratory stress test and Route A remains the primary analysis.
Our contribution differs from the closest concurrent literatures in four ways. First, relative to cross-fitted DML work on $ATT(g,t)$, we add closed-form fixed-weight SDID and target-level MC Riesz diagnostics on the partial-$R^2$ omitted-variable-bias scale. The $ATT(g,t)$ representer is adapted from BachKlaassenKueckMattesSpindler2025; the SDID, MC, and BJS diagnostics and the Route A/B reporting layer are the paper's main methodological additions. Second, relative to weighted partial-$R^2$ OVB work, our fixed-weight SDID diagnostic is the panel-contrast analogue of a weighted sensitivity calculation when weights are treated as fixed, but the fitted-weight feedback is not hidden: the refit audit reports both direction and decision-scale magnitude. Third, relative to latent-confounder Bayesian panel sensitivity, the route diagnostic can demote an observed-benchmark update when the benchmark population is too small, too coarse, alpha-side degenerate, or dependent in a way that leaves little effective information. Fourth, relative to parallel-trends-prior Bayesian DiD approaches such as HanMitraHettingerOganisian2025, the output is a decision-scale translation of frequentist breakdown thresholds into Route A prior probabilities and, only when diagnostics pass, Route B posterior-predictive probabilities.
In applications, the minimal reporting set is the fitted effect, sampling uncertainty, Riesz scale, Route A nullification and significance robustness values, benchmark count, alpha-side support, auxiliary-model check, dominance rationale, and route status.
The paper makes three contributions. First, it gives a Route A/Route B sensitivity workflow: Route A reports direct robustness-value profiles and prior-probability translations, while Route B is a diagnostic updating layer that uses observed-covariate benchmarks only after benchmark-count, alpha-side, dominance, dependence, and model-check diagnostics support calibration. Second, it derives estimator-specific Riesz diagnostics for SDID, matrix completion, fixed-effect imputation, and staggered $ATT(g,t)$, always stating whether the diagnostic is conditional on fitted weights, a target-level functional, or a first-stage fill-in operator. Third, it combines contraction and coverage theory, simulation stress tests, and two empirical studies: a single-treated-unit tobacco-control panel and a county-level staggered-adoption panel. The empirical studies show how the route diagnostics prevent observed-covariate benchmarks from being interpreted as calibrated posterior evidence when the benchmark population is not aligned with the estimator-specific Riesz representer.
Sections (ref)--(ref) set up the omitted-variable-bias bound and the Bayesian routes. Section (ref) gives the panel-estimator Riesz diagnostics. Section (ref) gives theory and operating characteristics. Sections (ref)--(ref) give the two empirical studies. Section (ref) gives reporting recommendations and limitations.
Let $W=(Y,D,X,U)$ denote the data, where $Y$ is the outcome, $D$ the treatment indicator, $X$ a vector of observed pretreatment covariates, and $U$ an unobserved confounder. The causal estimand $\theta_{\mathrm{long}}$ is identifiable under the long regression specification
where $m$ is a linear functional of the conditional expectation $g_{\mathrm{long}}$. The short specification accessible to the researcher replaces $g_{\mathrm{long}}$ by
The omitted-variable-biased estimand $\theta_{\mathrm{short}}$ generally differs from $\theta_{\mathrm{long}}$ by an amount that depends on the strength of $U$'s correlation with $Y$ and with the Riesz representer of the functional $m$ at $g_{\mathrm{short}}$.
The fundamental bound is
where $alpha(W)$ is the Riesz representer of $m$ at $g_{\mathrm{short}}$ under the empirical-distribution inner product and $\sigma^2_{Y\mid D,X}$ is the residual variance of $Y$ given $(D,X)$. We write $R^2_{yu}$ and $R^2_{alpha u}$ for the two partial $R^2$ terms and abbreviate \[ M=\left\{E[alpha^2(W)\sigma^2_{Y\mid D,X}(W)]\right\}^{1/2}. \] The bound factorizes the source of omitted-variable bias into an intrinsic scale $M$ depending on the estimator and the data, and a sensitivity term governed by the two partial $R^2$ values of the confounder.
The bound targets omitted components that can be summarized by their partial association with the outcome residual and with the Riesz representer. This class includes scalar or vector covariates $U$ after residualization, and it also includes projected latent components such as an interactive-factor term $U_{it}=\lambda_i'f_t$ when the researcher is willing to summarize the remaining unbalanced factor component by $(R^2_{yu},R^2_{alpha u})$. Arbitrary unbalanced interactive fixed effects can change the counterfactual surface in ways that require a design-level factor argument or an explicit projection into this partial-$R^2$ confounding class.
This distinction matters for interpretation. A low nullification RV in the California tobacco-control study means that a small equal-strength partial-$R^2$ confounder, measured in the outcome-residual and fitted-Riesz directions, can move the fitted contrast. Its scope is the additive or projected omitted-component class summarized by the two partial-$R^2$ coordinates. Latent-factor imbalance, SUTVA violations, spillovers, and misspecified treatment timing require a design-level argument or a projection of the violation into that class. The point of the workflow is therefore to state the target confounding class, report the Riesz scale, and then use route diagnostics to avoid turning observed-covariate benchmarks into automatic posterior validation.
We work with the Riesz-scaled effect size
For inference, let $c_t$ denote the critical value of the estimator's sampling distribution, typically $z_{1-alpha/2}=1.96$. The corresponding critical value on the $K$ scale is
For OLS, $SE/M=1/\sqrt{df}$, recovering the familiar Cinelli-Hazlett scaling. For SDID and MC, $M$ is computed from the Riesz representers of Section (ref), so the formula does not require residual degrees of freedom for a regularized estimator.
The robustness value $RV$ is the smallest equal-strength partial-$R^2$ value $r=R^2_{yu}=R^2_{alpha u}\in[0,1]$ at which the OVB-adjusted effect size hits the threshold $q c$, where $q=1$ reaches the boundary of conventional significance and $q=0$ nullifies the point estimate. Under equal strength, the squared scaled bound is $r^2/(1-r)$. Solving $r^2/(1-r)=\widetilde K^2$ with $\widetilde K=\max(K-qc,0)$ gives
The interpretation is invariant across estimators: an unobserved confounder must explain at least an $RV$ fraction of residual variance in both $Y$ and the Riesz representer $alpha$ to alter the conclusion at the chosen benchmark.
\noindentConnection to partial identification. For any fixed equal-strength confounding scale $r$, the OVB bound defines a partial-identification set around the short-regression estimand, up to the sampling component used for inference. The nullification RV is the smallest $r$ at which zero enters this set; the significance RV is the corresponding threshold for a sampling-adjusted set. Route A therefore does not identify the long-regression effect. It places a transparent probability model over the radius of a partial-identification set, while Route B updates that radius only when the observed benchmark population is credible.
We develop two routes for placing a Bayesian framework on top of the Riesz-representation OVB machinery. Route A places direct priors on $(R^2_{yu},R^2_{alpha u})$ without using observed-covariate information. Route B places a model on the population distribution of covariate strengths and updates via observed-covariate benchmarks. Route A is the default reporting baseline because it is always defined; Route B is reported as calibrated only after benchmark diagnostics support the auxiliary modeling step.
\noindentRoute A direct prior profile. Route A places Beta priors
The prior parameters are set by the user. Observed covariates can still help choose a prior scale, for example through the heuristic that an unobserved confounder should be no stronger than the strongest observed covariate, but Route A should be interpreted as a prior-predictive sensitivity profile rather than a data-identified posterior analysis.
\noindentObserved-covariate calibration and Route B. For each observed covariate $X_j$, compute a bivariate partial-$R^2$ benchmark \[ m_{\mathrm{obs}}^j=(R^2_{Y,X_j\mid rest},R^2_{alpha,X_j\mid rest}) \] using the most granular covariate representation available. We use one pessimism map throughout. For $r\in(0,1)$ and $\kappa\geq1$, define the odds-scale multiplier
A Route A calibration sets the prior means to
and fixes the Beta concentration after this mean calibration. We use concentration 20 as the reference value and report concentration sensitivity.
Measurement error has two distinct effects. Attenuated benchmark strengths can make a conditional Route B update look too robust, whereas the route gate works in the opposite direction by demoting weak observed alpha-side support. The full-pipeline experiment regenerates noisy covariates, residualizes them, and recomputes the partial-$R^2$ pairs. When reliability falls from 1.0 to 0.5, the median observed-to-latent alpha benchmark ratio falls to 0.812, the alpha-support gate rate falls from 0.963 to 0.819, and conditional coverage among promoted replications falls from 0.983 to 0.945. These results quantify both attenuation and the protective, but incomplete, response of the gate.
Route B turns the benchmarking heuristic into an auxiliary population model. Let $T_0(\theta)$ denote a mean, tail-quantile, or upper-support functional of the benchmark population. The same odds multiplier is applied coordinatewise:
where the minimum is coordinatewise and $\eta_{\mathrm{cap}}>0$ keeps the OVB denominator away from zero. Independent Beta marginals provide the baseline working family, with logit-normal and Gaussian-copula alternatives used as sensitivity checks.
Observed benchmark pairs share outcomes, residualization steps, and often correlated covariates, so a literal product likelihood can overstate information. We therefore treat the auxiliary update as a generalized or power posterior in the sense of generalized Bayes and robust/coarsened Bayes BissiriHolmesWalker2016,GrunwaldVanOmmen2017,MillerDunson2019. The tempered working likelihood is
Here $p_{\mathrm{eff}}\leq p$ is an effective benchmark count, so the composite score has information of order $p_{\mathrm{eff}}$ rather than $p$. In an equicorrelation benchmark-score model with common correlation $\rho_b$, $p_{\mathrm{eff}}=p/\{1+(p-1)\rho_b\}$ follows from the variance of the average score. In non-equicorrelated designs we use the same mean-score formula, $p_{\mathrm{eff}}=p^2/({\bf 1}'\widehat R{\bf 1})$, where $\widehat R$ is the empirical benchmark-score correlation matrix after residualization; a spectral participation-ratio diagnostic, $(\operatorname{tr}\widehat R)^2/\operatorname{tr}(\widehat R^2)$, is reported as a secondary dimension-count check. These adjustments are modeling choices, not identification assumptions. If dependence leaves $p_{\mathrm{eff}}$ bounded, Route B does not contract as more redundant benchmarks are added. Updating with $L_c$ gives $\pi(\theta\mid m_{\mathrm{obs}})$, and the posterior of $(R^2_{yu},R^2_{alpha u})$ is the pushforward through $T(\cdot;\kappa)$.
\noindentPosterior summaries and route choice. For each posterior draw $(R^{2,(s)}_{yu},R^{2,(s)}_{alpha u})$, compute the OVB bound
and the worst-direction adjusted estimate
The posterior distribution of the robustness value is obtained by applying (ref) to each draw of the adjusted Riesz-scaled effect size.
Route choice is a reporting rule, not an estimator-selection rule. Route A should always be reported. Route B is promoted from exploratory to calibrated only when the effective number of benchmark covariates is large enough, alpha-side benchmarks are nondegenerate and computed at the same observational level as the Riesz representer, posterior-predictive checks do not reject the auxiliary benchmark model, and the predictive distribution is plausibly pessimistic enough to dominate the hidden-confounder strength. In the simulations and applications, rejection of the marginal benchmark-model check at the 10% level is used as the default M1 warning threshold; such a warning demotes Route B unless a design-specific justification is stated. A larger $\kappa$ is a sensitivity display rather than a calibration repair when the benchmark model is not supported.
\noindentBayesian interpretation. Route A probabilities are prior probabilities over sensitivity parameters, not posterior probabilities about an identified causal effect. Route B probabilities are posterior-predictive probabilities under an auxiliary benchmark-population model. They are reported as calibrated only when benchmark count, alpha-side alignment, model-check diagnostics, and predictive-dominance plausibility all pass. If any diagnostic fails, Route B remains an exploratory stress profile and Route A is primary.
Before giving formulas, we fix the scope of each diagnostic. Each diagnostic is conditional on a first-stage object. The SDID diagnostic is the fitted-contrast Riesz vector evaluated at the estimated unit and time weights. The matrix-completion diagnostic is the treated-cell target vector for a fixed treated-cell index set. The fixed-effect-imputation diagnostic is the imputation-operator Riesz vector conditional on the untreated-cell two-way projection. The group-time diagnostic is the fitted score contrast given the nuisance scores and comparison sets. These objects put several fitted causal-panel contrasts on a common omitted-variable-bias scale while keeping the first-stage convention visible.
\noindentFixed-weight SDID Riesz diagnostic. The SDID estimator ArkhangelskyEtAl2021 computes an ATT using unit weights $\widehat\omega_i$ and time weights $\widehat\lambda_t$. Let $g(i)\in\{C,T\}$ denote treatment-group membership and $p(t)\in\{pre,post\}$ denote period status. Conditional on the estimated weights being held fixed, the fitted SDID contrast is linear in the outcome surface.
The proof is immediate: apply the weighted double-difference to a generic outcome surface and match coefficients under the empirical cell inner product. The $NT$ prefactor cancels the empirical-mass normalization. Synthetic panels and the California tobacco-control panel verify the identity to machine precision.
If the weights were externally fixed, this diagnostic would be the panel-contrast analogue of a weighted partial-$R^2$ omitted-variable-bias calculation such as WainsteinHazlett2025. The difference is that SDID weights are estimated from the same outcome panel. We therefore interpret ((ref)) as a fitted-contrast diagnostic and explicitly report a refit finite-difference audit; the audit is the device that quantifies the first-stage weight feedback omitted by the fixed-weight representation. A fully orthogonalized SDID sensitivity diagnostic is conceptually possible, in the spirit of orthogonal-score DiD and DML constructions ChernozhukovEtAl2018,SantAnnaZhao2020,BachKlaassenKueckMattesSpindler2025. We do not use it as the baseline because the two-sided simplex constraints in SDID weights can generate nonstandard corner-solution asymptotics; the fixed-weight diagnostic plus refit audit is therefore a deliberate choice for interpretability and transparent measurement of weight feedback.
\noindentMatrix-completion target-level diagnostic. The MC estimator imputes the counterfactual outcome surface for treated cells via nuclear-norm regularization. The diagnostic isolates the treated-cell counterfactual-mean target, with the nuclear-norm training map treated as the first-stage procedure that defines the imputed surface.
This representer is constant on the treated post-treatment cells and zero elsewhere. It is intentionally target-level: it bounds omitted-variable bias in the treated-cell counterfactual target conditional on the imputed surface. This target-level quantity is an interpretable scale for comparing MC with concentrated SDID weights, while a fully differentiated nuclear-norm training-map diagnostic is a distinct estimator-level object.
\noindentGroup-time and fixed-effect-imputation representers. For staggered designs, we record the $ATT(g,t)$ representer for the doubly robust score of SantAnnaZhao2020, following the Riesz representation in BachKlaassenKueckMattesSpindler2025.
The fixed-effect imputation estimator of BorusyakJaravelSpiess2024 estimates unit and period effects on untreated cells and imputes treated-cell counterfactuals from the fitted two-way fixed-effect structure. Let $\mathcal U$ and $\mathcal W$ denote untreated and treated cells, $N_1=|\mathcal W|$, and $N_{\mathrm{cell}}=NT$. Let \[ \Pi=X_{\mathcal W}(X_{\mathcal U}'X_{\mathcal U})^{-}X_{\mathcal U}', \qquad \widehat Y_{\mathcal W}(0)=\Pi Y_{\mathcal U}. \] The BJS ATT is $\widehattau_{\mathrm{BJS}}=N_1^{-1}{\bf 1}_{\mathcal W}'(Y_{\mathcal W}-\Pi Y_{\mathcal U})$.
Table (ref) compares the Riesz second moments of MC, BJS, and fixed-weight SDID on synthetic and California tobacco-control panels. The MC second moment equals $NT/|\mathcal M|$ and therefore depends only on panel dimensions and treated target cells. The BJS and SDID moments also depend on the untreated-cell imputation operator and fitted SDID weights. The table reports $E[alpha^2]$ on the empirical-distribution scale; the Route A scale $M$ used in robustness values is $\{E[alpha^2\sigma^2_{Y\mid D,X}]\}^{1/2}$ and therefore also incorporates the residual-variance factor.
Figure (ref) plots the cell-level concentration behind this scale contrast. In the California tobacco-control panel, the fitted SDID weights are sufficiently concentrated that the fixed-weight SDID representer has a much larger second moment than either the MC target representer or the BJS fixed-effect-imputation representer. We call $1/\{E[alpha_{\mathrm{SDID}}^2]/E[alpha_{\mathrm{MC}}^2]\}$ the equivalent product-support share: the share of a uniform product support that would generate the same second-moment inflation. The California tobacco-control ratio $567.83/100.75=5.64$ implies an equivalent product-support share of 0.177. This is not a replacement for the Riesz calculation, but it gives an accessible diagnostic for why concentrated SDID weights amplify the omitted-variable-bias scale.
A finite-difference refit diagnostic quantifies what the fixed-weight convention omits. Across 80 random outcome perturbation directions in the California tobacco-control panel, refit and fixed-weight derivatives have correlation 0.941 with median symmetric relative difference 0.250. The same diagnostic translates to the downstream Route A decision scale: the exact fixed-weight, random-projection, and refit-projection nullification RVs are 0.054, 0.054, and 0.045. Refit derivatives raise the local Riesz scale from 281.98 to 339.68 and lower the nullification RV from 0.054 to 0.045. In this application, ignoring first-stage weight feedback is therefore mildly anti-conservative on the nullification scale. We do not assume that this sign is universal: the mechanism is that perturbing the outcome surface changes the pre-period balance problem, and the finite-difference audit measures the induced local movement in the fitted SDID contrast. The qualitative conclusion remains in the same low-single-digit robustness range.
\FloatBarrier
We state two asymptotic theorems for the calibrated Route B case and then report simulation operating characteristics. The theorems characterize the case in which the auxiliary benchmark model and dominance conditions are credible. When those diagnostics fail, the empirical workflow reverts to Route A as the primary analysis. Proofs are provided in Appendix A.
We work with the dependence-adjusted auxiliary formulation of Section (ref). The benchmark sequence may be weakly dependent, but its tempered composite score is assumed to satisfy a local asymptotic normal expansion with effective information proportional to $p_{\mathrm{eff}}$. The remaining conditions are identification of the working population family; plug-in benchmark-estimation consistency at rate $(\log p/n)^{1/2}$; differentiability of the odds-scale link; prior support around the truth; non-borderline robustness-value smoothness; and interior sensitivity coordinates. Regime A is auxiliary-information dominant, with $n$ growing faster than $p_{\mathrm{eff}}\log p$; Regime B is benchmark-estimation dominant. The contraction argument follows the finite-dimensional LAN template in VanDerVaart1998 and GhosalVanderVaart2017, with composite-score information and plug-in error carried separately. If $p_{\mathrm{eff}}$ does not diverge, adding redundant covariates cannot produce a shrinking Route B posterior.
The dominance requirement is therefore best read at the scale of the bound that enters the decision problem. Product-order dominance of $(R^2_{yu},R^2_{alpha u})$ remains a transparent sufficient condition; scalar dominance of $B(R)$ is weaker and can be defended directly by arguing that the selected link and $\kappa$ produce a predictive distribution no less pessimistic than the analyst's substantive upper envelope for hidden confounding. The M1 diagnostic does not test this premise. If the hidden-confounder scale cannot plausibly be bounded by the chosen link and $\kappa$, reporting reverts to Route A.
Theorem 2 covers regular sampling components. The single-treated-unit tobacco-control study uses a discrete donor-placebo distribution with 38 donor pseudo-treatments plus the realized California assignment, so the theorem is separate from the finite-donor rank calculation reported there. Because the baseline add-one rank is already $2/39=0.051$, the conventional-significance RV is mechanically near zero; the interpretable sensitivity quantity is the nullification RV.
All simulation and stress diagnostics use direct Monte Carlo draws rather than Markov-chain Monte Carlo. Each table states its replication count. The estimator-specific Riesz identities are checked algebraically and numerically, and the simulation designs are defined explicitly in the text and Appendix.
The numerical evidence is organized around distinct mechanisms of benchmark failure. The stress-test designs are clean exchangeability, dominance failure, coarse alpha-side benchmarks, selected mixtures, benchmark dependence, measurement error, and SDID weight concentration. Table (ref) reports representative operating characteristics recomputed under the unified odds-scale link. The clean design is calibrated. Dominance failure passes the route gate but undercovers because the hidden confounder violates joint dominance. The coarse-alpha design is never promoted, and the selected-mixture design is promoted selectively while retaining high conditional coverage.
The dominance surface confirms that $\kappa$ is a reporting profile rather than a universal constant. At $p=20$ and $\kappa=2.5$, coverage is 0.975 when the hidden odds multiplier is $\delta=2$, 0.907 at $\delta=3$, and 0.808 at $\delta=4$. The alpha-alignment experiment reaches the complementary conclusion: shrinking alpha-side odds to 0.03 of baseline lowers naive coverage to 0.564, while the full distributional-support gate demotes every replication.
Dependence among observed benchmarks changes the information count even when nominal $p$ is fixed. With $p=40$, increasing equicorrelation from 0 to 0.75 reduces the diagnostic effective count from 40.0 to 1.32. A nominal-$p$ working likelihood then lowers coverage from 0.974 to 0.785; the effective-count correction keeps coverage between 0.970 and 0.974. A non-equicorrelation stress check gives the same message: for four blocks of ten covariates with within-block correlation 0.75, naive 95th-percentile-link coverage is 0.636, while mean-score and spectral effective-count corrections give 0.996 and 0.988; for AR(1) correlation 0.90, the corresponding numbers are 0.545, 0.999, and 0.984. The mean-score correction is deliberately conservative under strong non-equicorrelated dependence, whereas the spectral correction is closer to nominal; when non-equicorrelation is suspected, both should be reported, with the spectral value serving as a less conservative calibration check. The full-pipeline measurement-error experiment separately regenerates noisy covariates and their partial-$R^2$ pairs. At reliability 0.50, the median observed alpha benchmark is 0.812 of its latent value, the gate promotes 0.819 of replications, and conditional coverage is 0.945. Appendix B reports the full grids.
\noindentCoverage diagnostics. Table (ref) gives the finite-sample counterpart of the posterior-predictive coverage theorem. Conditional coverage is computed only among replications classified as Route B calibrated. Clean exchangeability satisfies the auxiliary model and joint dominance premises; dominance failure keeps promotion near one but violates the substantive dominance condition.
\noindentDecision-scale reporting. The frequentist breakdown bound of RambachanRoth2023 and the Bayesian robustness value answer related but non-equivalent questions. A breakdown calculation reports a tipping point: the violation magnitude at which a conclusion changes. Route A adds a prior-probability reading of that same tipping point. Route B adds a second, data-updating layer when the observed benchmark population passes the route diagnostic. Tables (ref) and (ref) illustrate this separation for the California tobacco-control nullification threshold $RV=0.054$ and a calibrated full-update experiment.
For Route B, auxiliary-family sensitivity should be reported whenever the route is interpreted as calibrated. In this calibrated numerical update, Beta-marginal, logit-normal, and Gaussian-copula specifications give median posterior RVs of 0.128, 0.121, and 0.134, respectively, and $\Pr(RV<0.10)$ values from 0.188 to 0.226. Appendix B reports this result together with link-rule and prior-concentration sensitivity in a single panelized table: moderate concentration changes leave the predictive RV near 0.12, while moving from mean to maximum link rules shifts the interpretation as expected. Disagreement across plausible auxiliary families in an empirical application should be treated as evidence against calibrated Route B reporting.
\FloatBarrier
This empirical study evaluates the workflow in a canonical single-treated-unit tobacco-control panel. The study combines the SDID point estimate, finite-donor placebo uncertainty, the Riesz scale that converts partial-$R^2$ sensitivity parameters into effect units, and the observed-covariate diagnostics needed before any benchmark update is interpreted as calibrated.
The data are the state-level cigarette-consumption panel of AbadieDiamondHainmueller2010. California is treated beginning in 1989, and the donor pool consists of untreated states used in the standard application. The outcome is cigarette packs per capita. Available benchmark covariates are seven state-level summaries: log GDP per capita, pretreatment smoking mean, age 15--24 share, retail cigarette price, and cigarette consumption in 1980, 1975, and 1988.
Figure (ref) reports raw trajectories and the fitted SDID counterfactual. The SDID estimate is \[ \widehattau_{\mathrm{SDID}}=-15.60 \] packs per capita. The pre-treatment RMSE of the unit-weighted counterfactual path is 1.72, the mean post-treatment gap is $-16.11$, and the SDID time-weighted pre-gap adjustment is $-0.51$. The negative point estimate is not driven by an obvious pre-period mismatch in the fitted trajectory.
\noindentEstimator diagnostics. The one-treated-unit design makes finite-sample diagnostics essential. Figure (ref) reports SDID unit and time weights together with leave-one-donor-out influence. The most influential single donor is Nevada. Omitting one donor at a time gives SDID estimates between $-17.06$ and $-14.98$, so the sign and approximate magnitude are not driven by a single donor state.
Table (ref) compares the SDID estimate with simple alternatives. All estimates have a negative sign, although magnitudes vary with the identifying structure. We therefore read the empirical pattern as substantively stable but not estimator-invariant.
Figure (ref) gives the corrected leave-one-control-out placebo distribution. The placebo standard error is 9.49 and the add-one left-tail and absolute-rank placebo $p$-values are both $2/39=0.051$. The point estimate is close to the edge of the placebo distribution, but finite-sample uncertainty in this single-treated-unit design remains wide.
For the fitted SDID weights, the fixed-weight Riesz diagnostic verifies \[ \frac{1}{NT}\sum_{it}\widehatalpha_{it}m_{it}=\widehattau_{\mathrm{SDID}}(m) \] to numerical precision for random perturbation checks. The resulting Riesz scale is \[ M_{\mathrm{SDID}}=281.98, \qquad K=|\widehattau|/M_{\mathrm{SDID}}=0.055. \] Using the corrected placebo standard error, the equal-strength Route A robustness value for nullifying the point estimate is 0.054. The corresponding robustness value for preserving conventional 5% placebo significance is below 0.001, because corrected placebo inference is already borderline. The refit-derivative check moves the nullification RV to 0.045, leaving the qualitative conclusion unchanged. Table (ref) consolidates the fitted Route A scale and the observed-benchmark odds-multiplier diagnostics.
Panel B is deliberately diagnostic. Because the alpha-side observed benchmarks are at the numerical floor, even a 10-times odds-scale multiplier produces a small worst-direction bound relative to the point estimate. This reinforces the route decision rather than validating Route B calibration.
\noindentBenchmark diagnostics and route decision. The state-level covariates are useful for outcome-side benchmarking, but they are too coarse for the alpha side of this SDID sensitivity problem. The SDID Riesz representer is a cell-level object concentrated in specific state-year locations. Aggregating it to the state level leaves all observed alpha-side partial-$R^2$ benchmarks numerically near zero. This pattern identifies a benchmark-alignment limitation: the available observed covariates do not align with the sensitivity functional. Figure (ref) displays the seven outcome- and alpha-side benchmark pairs underlying this decision.
The effective benchmark count is seven, the maximum observed alpha-side benchmark is approximately $10^{-6}$, and the fitted Beta-marginal diagnostic rejects the alpha-side benchmark model with $p_alpha\approx 9\times10^{-6}$. Full-pipeline placebo diagnostics reinforce this decision: across 38 donor placebo assignments, the median nullification RV is 0.021 and the maximum is 0.086, while Route B is never classified as calibrated. Appendix B adds three route-stability checks: varying the minimum benchmark-count threshold, varying the alpha-support threshold, and dropping each observed benchmark one at a time all leave the application Route A-primary. It also reports a single-treated-unit $p=7$ stress design in which outcome-side benchmarks are nonzero but alpha-side benchmarks lie at the numerical floor; the route rule demotes the design in all replications.
As an exploratory Route B profile, we report the auxiliary benchmark calculation at $\kappa=2.5$ only as a stress profile. The median sign-truncated residual is $-6.37$ packs; 69.8% of draws leave a nonzero negative residual; and sampling plus OVB uncertainty gives a 95% interval $[-27.73,33.72]$. The California tobacco-control panel therefore combines a stable negative SDID point estimate with wide one-treated-unit placebo uncertainty and an observed benchmark set that classifies the design as Route A-primary.
\FloatBarrier
The second empirical study evaluates the workflow in a multi-cohort staggered-adoption panel using the mpdta minimum-wage panel considered by CallawaySantAnna2021. The panel has 500 counties from 2003 to 2007, log teen employment as the outcome, three treatment cohorts, and 309 never-treated counties. This design illustrates the group-time version of the Riesz diagnostic, its aggregation to an overall ATT summary, and the corresponding Route A robustness profile. A calibrated Route B analysis would require a credible $ATT(g,t)$-level benchmark population, so the available benchmark structure supports Route A-primary reporting in this empirical study.
For each treated cohort $g$ and post-treatment year $t\geq g$, we compute a nonparametric group-time contrast against never-treated counties using $g-1$ as the base period. Aggregating by cohort size yields $-0.040$ log points, with county-cluster bootstrap SE 0.012 and 95% interval $[-0.065,-0.016]$. The maximum pseudo-ATT is 0.034, the Route A nullification RV is 0.993, and the significance RV is 0.958. Table (ref) summarizes these estimates together with the group-time and fixed-effect-imputation Riesz norms. Thus the same workflow returns a high-RV multi-cohort case, in contrast to the low-single-digit California tobacco-control sensitivity scale. A leave-cohort-out check in Appendix B leaves the nullification RV above 0.98 in every deletion.
\FloatBarrier
This second application demonstrates that the $ATT(g,t)$ representer operates in a multi-cohort panel. The workflow reports sensitivity at the group-time level and aggregates only after making group-time weights explicit. Because no credible $ATT(g,t)$-level benchmark population is observed, the empirical study is classified as Route A-primary. The fixed-effect-imputation comparison is close in both effect and Riesz scale, providing a useful diagnostic cross-check.
\FloatBarrier
The workflow is designed to prevent sensitivity analysis from being read as automatic posterior validation. Route A is the default: it gives a deterministic robustness-value profile for the fitted Riesz scale and point estimate. Route B is promoted to calibrated status only when the benchmark set is large enough, alpha-side benchmarks are nondegenerate, auxiliary model checks are credible, and predictive dominance is substantively plausible. When the diagnostics do not support calibrated updating, increasing $\kappa$ is reported as sensitivity profiling rather than as a calibration repair, and Route A remains primary.
Applied users should report the fitted estimator, sampling uncertainty, Riesz scale, Route A RVs, benchmark count, alpha-side nondegeneracy, auxiliary-model check, dominance rationale, and route status before interpreting a Bayesian update. Aggregated covariates with degenerate alpha-side benchmarks imply Route A-primary reporting. Unit-time or state-year benchmarks with adequate count and credible diagnostics can justify a calibrated Route B update, but only if the hidden-confounder strength scale is plausibly dominated by the predictive distribution used for reporting.
The diagnostics are intentionally conditional. The SDID result conditions on fitted weights, the MC result is target-level, and the BJS result conditions on the untreated-cell fixed-effect fill-in operator. The California tobacco-control finite-difference check shows that refitting can add local movement; in that application, fixed weights give a slightly larger nullification RV than the refit projection, so the fixed-weight diagnostic is mildly anti-conservative on the decision scale. The target-level MC diagnostic is a robustness diagnostic for the treated counterfactual target rather than for the full nuclear-norm training map. Interactive fixed-effect imbalance, latent-factor violations, spillovers, and SUTVA failures require either a design-level argument or a deliberate projection into the partial-$R^2$ confounding class described in Section (ref). The donor-block exclusion diagnostic in Appendix B leaves the Route A conclusion in the same low-single-digit range, while interference remains a design-level concern.
The scope is short- or moderate-$T$ causal panels with ATT-type estimands. Long-$T$ macro panels with strong serial correlation would require either a HAC-adjusted Riesz norm or an explicit temporal-dependence model. Structural parameters such as demand elasticities would require a new Riesz representer for the structural estimand itself. These are useful extensions, but they are distinct from the ATT-type causal-panel problem studied here.