EconBase
← Back to paper

Learning about Treatment Effects in Panels under Unknown Interference

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

118,244 characters

Learning about Treatment Effects in Panels under Unknown Interference


\maketitle

\begin{abstract}
\setlength{\emergencystretch}{4em}
\sloppy
When comparison units may also respond to treatment, panel comparisons reflect
both the treatment effect and spillovers. If the interference pattern is unknown,
observed outcomes alone do not separate the two. I characterize what can
nevertheless be learned from panel outcomes under general restrictions,
without requiring an exposure mapping or prior classification of affected
donors. The framework scales validity bounds for every convex donor weight by
its fit before treatment and combines these bounds with prespecified
restrictions tailored to the application. The validity bounds constrain the
treatment effect relative to spillovers, while the additional restrictions
determine its possible values. Together these restrictions yield a sharp
identified set. When the additional restrictions have a finite linear
representation, checking whether a proposed treatment effect is compatible
with the model reduces exactly to asking whether a finite linear system has a
solution. Bootstrap calibration tests this condition. Inverting these tests
uniformly controls, in large samples, the probability of falsely excluding
each compatible value. In an application to the Legal Arizona Workers Act,
the resulting 95 percent inversion sets contain effects of both signs across
all reported specifications, leaving the sign of the treatment effect
unresolved.
\fussy
\end{abstract}



\section{Introduction}

Panel treatment effect methods use untreated units to construct counterfactual
outcomes. Their interpretation changes when the policy also affects the
comparison pool. The Legal Arizona Workers Act (LAWA) provides a useful
example. If Arizona's employment verification law deterred migrants from
Arizona or redirected them to other states, outcomes in comparison states may
also respond. Comparisons between Arizona and those states then reflect both
the treatment effect and spillovers. The same concern arises when policies
redirect activity across jurisdictions, providers, or locations or propagate
through social and economic networks, as in tax, environmental, health, and
place-based settings
\citep{agrawal2015taxgradient,lepissier2021climate,
alexander2023hospitalclosures,fischer2024obstetric,
banerjee2013diffusion,cai2015social}.

Applications often provide substantive restrictions on treatment and spillover
effects even when the spillover pattern remains unknown. Such restrictions
may come from outcome support, aggregate accounting, signs, orderings,
magnitude information, or a structural exposure model. I treat the full
vector of spillovers to comparison units as unknown and impose restrictions
directly on it, without requiring an exposure mapping or prior classification
of affected donors. Recent work formalizes the basic identification problem in
a difference-in-differences setting with two groups: under a parallel-trends
condition for no-policy outcomes and otherwise unknown interference, the
canonical estimand identifies the average total effect on treated units only
relative to the average spillover on controls
\citep{mealli2026difference}. This paper develops the corresponding problem
with many donors. Additional weights are not additional data.
They become informative when the same treatment effect and spillover vector
must satisfy a common validity rule in several comparison directions. Building
on \citet{liu2025synthetic}, I impose this common rule over the full simplex
generated by a prespecified donor pool.

The paper contributes a sharp identification result, an exact finite
representation of the restrictions over all donor weights, and a uniform
candidatewise inference result for the structured systems generated by that
representation. These three pieces separate the economic content of the model
from computation and sampling theory.

For each weight, the rule bounds its latent no-policy gap after treatment by a
prespecified multiple of its typical absolute discrepancy before treatment.
The distinctive
commitments are that the rule holds uniformly over the prespecified donor
simplex and that its baseline allowance has a zero floor when population
discrepancy before treatment vanishes. A condition on the span of factor
movements provides a sufficient condition for the full system. To make this
multiplier easier to interpret, I construct factor placebo indices by
leaving out one period at a time. These placebos put the multiplier on an
observed scale, following the calibration principle of
\citet{hsu2013calibrating}, but they benchmark rather than estimate the latent
threshold after treatment.

The analysis considers one candidate value of the treatment effect at a time.
A candidate is compatible if some unobserved no-policy outcomes and spillovers
satisfy all maintained restrictions. Comparison validity constrains the
treatment effect relative to spillovers, while restrictions tailored to the
application determine which levels of the treatment effect remain possible.
Together, they characterize the sharp identified set and show whether it has
an upper endpoint, a lower endpoint, or both.

Under the baseline finite representation, a candidate is compatible exactly
when a linear system of fixed dimension has a solution. An exact equivalence
replaces the continuum of comparison weights with finitely many inequalities
and shows when interior weights add information beyond individual donors.
Farkas' alternative provides a finite certificate of incompatibility. Building
on the normalized solvability statistic and bootstrap calibration of
\citet{goff2025inference}, I establish uniform validity for this model-generated
family of candidate systems and invert the resulting tests over candidate
values. The analysis exploits their common primitive sampling perturbation and
uses regularity conditions tailored to compatible right-hand sides and
structurally reachable certificate directions. In large samples, the procedure
uniformly controls the probability of falsely excluding each compatible
candidate, while candidates that remain a fixed normalized distance from
compatibility in the population are consistently excluded. These results are
conditional on the donor pool, chosen validity envelope, admissible rule, and
its finite representation.

I apply the method to the LAWA application of \citet{bohn2014lawa}. The
application combines state-level outcomes with outcome support and a gross
spillover budget scaled by population. The main specification bounds total
absolute responses across comparison units by twice the counterfactual
population of Arizona's target group; I also report bounds equal to one and
four times that population. This formulation allows interstate redistribution, changes in
the national stock, and responses outside the comparison pool without
requiring an exact relocation identity. Across the reported validity envelopes and
all three budget calibrations, neither zero nor the decline of 1.50 percentage
points reported by \citet{bohn2014lawa} is rejected. The original magnitude
therefore remains compatible with the data and restrictions, but the analysis
does not determine the sign of the effect after accounting for sampling
uncertainty.

The Monte Carlo design separates identification geometry from sampling
performance. When pre-treatment discrepancies are sign aligned, the
full simplex and donor vertices yield the same population identified set. When
convex aggregation produces near-exact fit, interior weights reduce the
width of that set by 86 percent. Across both geometries, false exclusion at
compatible endpoints shows no evidence of excessive rejection in the simulated
cells, while power at fixed distances rises with sampling precision.

The remainder of the paper proceeds as follows.
\hyperref[sec:setting]{Section~\ref*{sec:setting}} sets up the model,
\hyperref[sec:identification]{Section~\ref*{sec:identification}} develops the
exact finite representation of candidate compatibility and its projection
consequences, and
\hyperref[sec:inference]{Section~\ref*{sec:inference}} constructs compatibility
tests for fixed candidate values and their inversion.
\hyperref[sec:empirical_illustration]{Section~\ref*{sec:empirical_illustration}}
revisits LAWA, and \hyperref[sec:monte_carlo]{Section~\ref*{sec:monte_carlo}}
studies performance in finite samples. The appendix collects proofs,
implementation details, and additional results.


\subsection{Related Literature}\label{sec:related_literature}

The paper is closest to panel methods for treatment effects that allow an
intervention to contaminate comparison units. Some synthetic control
approaches specify a linear, spatial, or exposure structure, or identify which
units are affected, and
jointly recover direct and spillover effects
\citep{cao2019estimation,distefano2024inclusive,grossi2025direct,
sakaguchi2026bayesian}. \citet{melnychuk2024synthetic} compares strategies for
constructing counterfactuals for the treated unit when donors may be contaminated.
Others downweight, screen, or
exclude donors considered likely to be affected
\citep{oriordan2025spillover,fernandezmorales2026bayesian}. I instead treat the
full vector of spillovers to comparison units as unknown. Joint validity
across contaminated comparisons and prespecified restrictions on that vector
partially identify the treatment effect. This route requires neither a
maintained exposure mapping, a pure donor subset, nor point identification of
the spillover profile. \citet{celli2026identifying}
contrast counterfactuals based on controls and forecasts under pervasive
interference; the framework here retains panel comparisons based on controls but
replaces point identification with explicit compatibility restrictions.

The broader interference literature defines direct and spillover effects using
treatment assignment or exposure mappings and derives identification under
experimental, network, or neighborhood structures
\citep{manski2013identification,sobel2006randomized,hudgens2008toward,
aronow2017general,vazquezbare2023spillover,leung2020treatment,
leung2022causal,forastiere2021identification,huber2021framework}. Under
randomized assignment, \citet{savje2021average} define an average treatment
effect that remains meaningful under limited but otherwise unknown
interference. Extensions to
panel data adapt parallel trends restrictions to spatial or neighborhood
interference \citep{butts2021difference,xu2023difference}.
\citet{mealli2026difference} show that, under a parallel-trends condition for
no-policy outcomes and otherwise unknown interference, the canonical
difference-in-differences estimand identifies the average total effect on
treated units net of the average spillover on controls. The present paper starts from
the analogous accounting relation with many donors and derives sharp sets from
validity restrictions indexed by weights and joint restrictions on the spillover
vector. Work on misspecified exposure mappings is complementary: it studies
robustness to the mapping itself, whereas the admissible set here directly restricts the
resulting spillover profile and can incorporate implications of a maintained
exposure model \citep{savje2024causal,schroder2026causal}.

Synthetic control and related panel methods use pre-treatment outcomes to form
counterfactual comparisons
\citep{abadie2003economic,abadie2010synthetic,abadie2015comparative};
\citet{abadie2021using} surveys the design and its extensions. Like synthetic
control and synthetic difference-in-differences
\citep{arkhangelsky2021synthetic}, the framework here uses pre-treatment
outcomes to constrain post-treatment comparisons. Its distinction is that a
sensitivity envelope is imposed over every weight in a prespecified donor
simplex rather than attached to one estimated comparison. In all three
approaches, the connection between pre-treatment fit and post-treatment
validity ultimately rests on maintained restrictions on untreated outcomes.
Interactive factor models provide one common way to extrapolate from outcomes
before treatment to outcomes after treatment
\citep{bai2009panel,gobillon2016regional,xu2017generalized}.
Analyses of imperfect fit characterize its consequences, propose bias
corrections, or develop alternative synthetic control estimators
\citep{ferman2021synthetic,benmichael2021augmented,powell2026imperfect}.
Here the observed discrepancy instead enters directly as the scale of a
maintained validity envelope.

The closest connection is \citet{liu2025synthetic}, who studies the identifying
content of multiple comparison weights. In that paper, each weight that exactly
balances trends before treatment indexes a
possible counterfactual, and the identified set ranges over those alternatives.
Here the restrictions indexed by weights are intersected: every convex weight must
satisfy its own fit-scaled restriction for the same treatment effect and
spillover vector. Expanding the comparison domain therefore adds restrictions
to one joint feasible set. \citet{callaway2025beyond} likewise use sets of
close comparison groups, but target robustness across alternative panel
identification strategies. The fit scale is related to the relative-magnitude
sensitivity analysis of \citet{rambachan2023more}, which bounds post-treatment
departures using pre-treatment behavior. \citet{ferguson2021assessing}
calibrate misspecification in synthetic control using observable placebo
prediction errors. The calibration here instead benchmarks a common validity
envelope over all donor weights and combines it with restrictions on
spillovers to comparison units. It also adapts the calibration principle of
\citet{hsu2013calibrating}. Here, indices constructed by leaving out one period
at a time put specifications with fixed envelope values on an observed scale
without treating pre-treatment movements as estimates of the missing
post-treatment gap.

Finally, the paper relates to inference under partial identification and to
inference for optimization and problems involving linear systems. General
methods cover identified sets, moment inequalities, and low-dimensional projections
\citep{chernozhukov2007estimation,rosen2008confidence,
andrews2010inference,romano2010inference,romano2014practical,
chernozhukov2013intersection,bugni2017inference,kaido2019confidence,
gafarov2025simple}. Related work studies directionally differentiable maps,
linear moment models, and linear systems with known or estimated coefficients
\citep{fang2019inference,cho2023simple,andrews2023linear,fang2023large,
goff2025inference,bai2026linear}.
\citet{cox2025testing} develop tests for inequalities that are linear in
nuisance parameters, including applications in which target parameters are
bounded by linear programs.
Goff and Mbakop develop general tests for whether a linear system has a
solution. The inference result here is a structured specialization of their
procedure. I show that candidate compatibility in this model has an exact
finite representation with estimated coefficients, adapt their calibration to
the resulting candidate-indexed family, and establish validity under
model-tailored conditions on a uniform feasibility error bound and certificate
variance
before inverting the tests over candidate values.

\section{Setup and Notation}\label{sec:setting}
\noindent
I focus on settings where a binary treatment is implemented at an aggregate
level, such as states or countries. Denote each aggregate unit by
\(k\in\{1,\ldots,K\}\), where \(K\) is the total number of units, and let
\(k=1\) be the treated unit. Pre-treatment periods are indexed by
\(t=1,\ldots,T_0\), where \(T_0\) is the last pre-treatment period; \(T\)
labels one scalar post-treatment target rather than an additional date in this
sequence. In the LAWA application,
unit \(1\) is Arizona; Section~\ref{sec:empirical_illustration} gives the
corresponding sample, timing, and donor pool details.
Assume throughout that \(K\ge2\) and \(T_0\ge2\).
Throughout, \emph{prespecified} means fixed as part of the maintained
specification before estimating the corresponding candidate system and
evaluating its inversion; it does not denote formal preregistration.
I use one treated aggregate unit and one scalar post-treatment target for
notational simplicity. The accounting relation and finite-input construction
also apply when this scalar target is a prespecified linear aggregate of several
treated units or post-treatment outcomes. A simultaneous vector of
treatment-effect targets would additionally require definitions of the corresponding
spillover objects, admissible restrictions, and covariance input and is not
developed in the present notation.\footnote{The treated unit, comparison pool,
and treatment timing are taken as given; see
\citet{imbens2023identification} for an alternative framework in which aspects
of the design are viewed as random.}

For identification, each aggregate outcome is a population mean. Write
\(\mu_t^k\) for the observed mean of unit \(k\) in period \(t\). Let
\(\mathbf d^0=(0,\ldots,0)\) denote the baseline assignment without the focal
intervention and let
\(\mathbf d^1=(1,0,\ldots,0)\) denote the realized assignment in which unit
\(1\) receives the policy and the comparison units remain directly untreated.
Write \(\mu_t^k(\mathbf d)\) for the potential mean of unit \(k\) under
assignment \(\mathbf d\). In the LAWA example,
\(\mu_t^k(\mathbf d^0)\) is the population share in state \(k\) that would prevail
absent LAWA, while \(\mu_T^k(\mathbf d^1)\) is the share under the realized
LAWA assignment. The contrast changes only the focal intervention; background
policies remain part of the potential-outcome environment. Accordingly,
``no-policy'' below means without the focal intervention. I impose consistency
and no anticipation with respect to that intervention: for every \(k\),
\[
\mu_t^k=\mu_t^k(\mathbf d^0)=\mu_t^k(\mathbf d^1)\quad (t\le T_0),
\qquad
\mu_T^k=\mu_T^k(\mathbf d^1).
\]

The treatment and spillover effects compare these two complete assignments.
The target is the treatment effect on the treated aggregate unit,
\begin{equation}\label{eq:treatment}
\tau := \mu_T^1(\mathbf d^1)-\mu_T^1(\mathbf d^0),
\end{equation}
which changes unit \(1\)'s assignment while holding every comparison unit's
assignment at zero. The effect of the same assignment contrast on a
comparison unit is its spillover effect,
\begin{equation}\label{eq:spillover}
s_k := \mu_T^k(\mathbf d^1)-\mu_T^k(\mathbf d^0),
\qquad k=2,\ldots,K.
\end{equation}
Write \(s=(s_2,\ldots,s_K)^\top\). The benchmark \(s=0\) corresponds to zero
spillovers to comparison units. The analysis below treats \(s\) as unknown and
asks what contaminated comparisons and restrictions on \(s\) reveal about
\(\tau\).

The potential means are reduced-form equilibrium outcomes under each complete
assignment. Thus \(\tau\) and \(s\) can incorporate geographic propagation
through relocation, displacement, commuting, disease transmission, or regional
market adjustment, as well as network propagation through social, trade, or
input-output links \citep[e.g.,][]{miguel2004worms,greenstone2010agglomeration,
banerjee2013diffusion,cai2015social}. These channels may be nonlinear or
reciprocal; \(s_k\) records their net effect on comparison unit \(k\).
Throughout the paper, the treatment effect refers to the assignment contrast in
\eqref{eq:treatment}.

Identification of \(\tau\) therefore amounts to learning enough about
the missing post-treatment no-policy means
\((\mu_T^k(\mathbf d^0))_{k=1}^K\): the treated counterfactual
\(\mu_T^1(\mathbf d^0)\) determines the treatment effect, while the comparison
counterfactuals \((\mu_T^k(\mathbf d^0))_{k=2}^K\) determine the spillover
vector.

The aggregate means may be observed directly or estimated from micro data. In
the latter case, \(\mu_t^k\) denotes the population mean for individuals
attached to unit \(k\) in period \(t\), estimated by the corresponding
unit-period average from a panel or repeated cross section, with survey weights
when appropriate. Identification treats each \(\mu_t^k(\mathbf d)\) as a
population mean and restricts the
missing vector \((\mu_T^k(\mathbf d^0))_{k=1}^K\). Inference addresses sampling
uncertainty when the finite collection of aggregate inputs is estimated.

\noindent\textbf{Notation.} Throughout the paper,
\(\Delta^{K-2}:=\{w\in\mathbb R_+^{K-1}:\sum_{k=2}^K w_k=1\}\) denotes the
simplex of convex comparison weights over units \(k=2,\ldots,K\). The donor
pool is fixed before estimation using eligibility rules specific to the application,
and the baseline comparison domain is the full simplex.
For period changes, write
\(\Delta\mu_t^k(\mathbf d^0):=\mu_t^k(\mathbf d^0)-
\mu_{t-1}^k(\mathbf d^0)\) for \(t=2,\ldots,T_0\), and
\(\Delta\mu_T^k(\mathbf d):=\mu_T^k(\mathbf d)-
\mu_{T_0}^k(\mathbf d^0)\) for the post-treatment target contrast,
\(\mathbf d\in\{\mathbf d^0,\mathbf d^1\}\).
For a vector \(z\), \(\|z\|_1\), \(\|z\|_\infty\), and \(\|z\|\) denote the
\(\ell_1\), sup, and Euclidean norms. The vector \([z]_+\) has \(j\)th
coordinate \(\max\{z_j,0\}\). For a nonempty set \(\mathcal C\),
\(\operatorname{dist}\{z,\mathcal C\}:=\inf_{y\in\mathcal C}\|z-y\|\);
when \(\mathcal C\) is a polyhedron, \(\operatorname{ext}(\mathcal C)\)
denotes its set of extreme points. The symbol \(\mathbbm 1_m\) denotes the
\(m\)-vector of ones, with the subscript omitted when its dimension is clear.
For a matrix, \(\|\cdot\|\) denotes the Euclidean norm after vectorization.
Vector inequalities are interpreted componentwise.

\section{Identification}\label{sec:identification}
\noindent
The identifying mechanism is easiest to see with one treated unit and two
comparison units. Write \(s_2\) and \(s_3\) for the comparison unit spillovers.
As a benchmark with exact validity, suppose the untreated post-treatment change
of the treated unit equals the same weighted change among the comparisons. For
a weight \((\vartheta,1-\vartheta)\) with \(\vartheta\in[0,1]\), the observed
treated minus comparison change then satisfies
\begin{equation}\label{eq:exact_valid_comparison}
\Delta\mu_T^1(\mathbf d^1)
-\vartheta\Delta\mu_T^2(\mathbf d^1)
-(1-\vartheta)\Delta\mu_T^3(\mathbf d^1)
=\tau-\vartheta s_2-(1-\vartheta)s_3 .
\end{equation}
One equation fixes only one linear combination of
\((\tau,s_2,s_3)\). It cannot separate the treatment effect from the two
spillovers. A second valid weight
\((\vartheta',1-\vartheta')\), with \(\vartheta'\in[0,1]\) and
\(\vartheta'\ne\vartheta\), places different weights on the same spillovers
while leaving \((\tau,s_2,s_3)\) unchanged:
\begin{equation}\label{eq:exact_valid_system}
\begin{aligned}
\Delta\mu_T^1(\mathbf d^1)
-\vartheta\Delta\mu_T^2(\mathbf d^1)
-(1-\vartheta)\Delta\mu_T^3(\mathbf d^1)
&=\tau-\vartheta s_2-(1-\vartheta)s_3,\\
\Delta\mu_T^1(\mathbf d^1)
-\vartheta'\Delta\mu_T^2(\mathbf d^1)
-(1-\vartheta')\Delta\mu_T^3(\mathbf d^1)
&=\tau-\vartheta' s_2-(1-\vartheta')s_3 .
\end{aligned}
\end{equation}
The second equation is not new data; its identifying content comes from
maintaining exact validity for both weights. The equations are informative
jointly because they contain the same treatment effect and spillover vector.
Each additional linearly independent comparison further restricts differences
among spillovers. These comparisons cannot determine a common shift. For any
\(\delta\in\mathbb R\), replacing \((\tau,s)\) by
\((\tau+\delta,s+\delta\mathbbm 1_{K-1})\) leaves every expression
\(\tau-w^\top s\) unchanged because every convex weight sums to one. Appendix
Lemma~\ref{lem:exact_comparison_equivalence} formalizes this benchmark.

Exact validity clarifies the role of joint comparison restrictions, but it is
not the maintained empirical model. Weights that achieve exact validity may be
unavailable. For a generic \(w\in\Delta^{K-2}\), the population accounting
identity is
\begin{equation}\label{eq:contaminated_accounting}
\begin{aligned}
\Delta\mu_T^1(\mathbf d^1)-\sum_{k=2}^K w_k\Delta\mu_T^k(\mathbf d^1)
&=
\underbrace{
\Delta\mu_T^1(\mathbf d^0)-\sum_{k=2}^K w_k\Delta\mu_T^k(\mathbf d^0)
}_{\text{latent no-policy gap for }w}
\ +\ \tau-w^\top s .
\end{aligned}
\end{equation}
The first term on the right-hand side is the latent no-policy gap: it is
unobserved because it uses post-treatment means under \(\mathbf d^0\).
The framework first uses each weight's pre-treatment fit to bound this gap. It
then uses restrictions tailored to the application to determine which overall
levels of the treatment effect and spillovers remain possible.

\subsection{Comparison Validity and Relative Effects}
\label{subsec:comparison_validity_identification}

The first restriction converts the accounting identities into a system of
moment inequalities. Each convex weight in the donor simplex receives an
allowance determined by its own pre-treatment fit.

\begin{assumption}[Full-simplex fit-scaled comparison validity]
\label{ass:comparison_validity}
For a fixed envelope constant \(0\le L<\infty\), the untreated potential means
satisfy, for every \(w\in\Delta^{K-2}\),
\[
\left|
\Delta\mu_T^1(\mathbf d^0)
-\sum_{k=2}^K w_k\Delta\mu_T^k(\mathbf d^0)
\right|
\le
\frac{L}{T_0-1}
\sum_{t=2}^{T_0}
\left|
\Delta\mu_t^1(\mathbf d^0)
-\sum_{k=2}^K w_k\Delta\mu_t^k(\mathbf d^0)
\right|
.
\]
\end{assumption}

Once donor eligibility is fixed, the baseline applies the rule to every
nonnegative weight vector that sums to one. Every weight must satisfy its own
restriction for the same treatment effect and spillover vector, while a poorly
fitting weight receives a wider allowance than a well-fitting one. Like
synthetic control and synthetic difference-in-differences
\citep{abadie2010synthetic,arkhangelsky2021synthetic}, the assumption uses
pre-treatment behavior to constrain a post-treatment counterfactual through
maintained restrictions on untreated outcomes. Its distinctive commitment is
uniformity over the prespecified donor simplex. The baseline envelope is also
homogeneous and hence has a zero floor at exact population fit.
Section~\ref{subsec:comparison_validity_interpretation} shows that one
restriction on the factor span can imply validity for all weights and discusses
the implication of the zero floor.

For \(w\in\Delta^{K-2}\), write the observed post-treatment contrast and the
pre-treatment comparison gaps as
\[
\begin{aligned}
C_T^{\mathrm{obs}}(w)
&:=
\Delta\mu_T^1(\mathbf d^1)
-\sum_{k=2}^K w_k\Delta\mu_T^k(\mathbf d^1),\\
C_t(w)
&:=
\Delta\mu_t^1(\mathbf d^0)
-\sum_{k=2}^K w_k\Delta\mu_t^k(\mathbf d^0),
\qquad t=2,\ldots,T_0.
\end{aligned}
\]
Let \(\Pi_0\) denote the true finite-dimensional population input. It collects
the population means entering these contrasts, together with any other input
to the rule defining the admissible set that is estimated from the sampling data.
Prespecified design constants, such as
externally supplied exposure weights or weights based on reference size, are
held fixed and
absorbed into the rule itself rather than included in \(\Pi_0\).
For a treatment effect and spillover vector, set
\begin{equation}\label{eq:relative_effect_vector}
x=(x_2,\ldots,x_K)^\top:=\tau\mathbbm 1_{K-1}-s,
\qquad x_k=\tau-s_k .
\end{equation}
Assumption~\ref{ass:comparison_validity} restricts this vector of relative
effects to
\begin{equation}\label{eq:relative_effect_set}
\mathcal X_L(\Pi_0)
:=
\left\{
x\in\mathbb R^{K-1}:
\left|C_T^{\mathrm{obs}}(w)-w^\top x\right|
\le\frac{L}{T_0-1}\sum_{t=2}^{T_0}|C_t(w)|
\quad\text{for every }w\in\Delta^{K-2}
\right\}.
\end{equation}

The opening benchmark showed why multiple weight directions can be jointly
informative. Validity over the full simplex can also be more informative than
validity imposed only at the donor vertices. When donor pre-treatment gaps offset one
another, an interior weight can have a smaller fit allowance than the
corresponding convex combination of vertex allowances and therefore add a
restriction on \(x\). Under sign alignment this gain disappears. Appendix
Proposition~\ref{prop:comparison_multiplicity} gives the formal result, and
Section~\ref{sec:monte_carlo} illustrates its quantitative importance.

The defining inequalities in \eqref{eq:relative_effect_set} have a direct
geometric interpretation. Each weight defines a slab around the hyperplane for
exact validity, \(C_T^{\mathrm{obs}}(w)=w^\top x\), with tolerance determined by that
weight's pre-treatment fit. Their intersection is
\(\mathcal X_L(\Pi_0)\) in coordinates for the relative effects and,
equivalently, the region allowed by comparison validity in \((\tau,s)\)-space,
as illustrated
in Figure~\ref{fig:comparison_validity_geometry}.

\begin{figure}[!htbp]
\centering
\includegraphics[width=\textwidth]{theory_assumption31_feasible_region_mpl_better.pdf}
\caption{Comparison validity as inequalities scaled by pre-treatment fit. Each
comparison weight replaces exact equality with a region whose width depends on
that weight's pre-treatment fit. The panels successively impose one, two, and
many comparison restrictions. In every panel, blue denotes the intersection
that satisfies all displayed restrictions, gray marks the boundaries of the
individual restrictions, and the dashed line indicates the common-shift
direction. The open ends show that adding the same constant to the treatment
effect and all spillovers remains unrestricted.}
\label{fig:comparison_validity_geometry}
\end{figure}

\subsection{From Relative Effects to Treatment Effect Levels}
\label{subsec:sharp_projection}

Comparison validity constrains the treatment effect only relative to
spillovers. If \((\tau,s)\) satisfies all comparison restrictions, then, for
any \(\delta\in\mathbb R\), adding \(\delta\) to \(\tau\) and every component
of \(s\) leaves those restrictions unchanged. This common-shift ambiguity means
that the comparison model alone cannot determine the level of the treatment
effect.

Restricting the common shift does not require an exposure mapping or
unit-specific knowledge of the spillover pattern. To obtain level information,
it is enough to impose prespecified substantive restrictions that limit
movement in this direction. These restrictions may remain agnostic about which
comparison units are affected and how spillovers vary across them. Let \(\Pi\)
denote a generic value of the finite population input, and use
\(\Omega(\Pi)\) to collect these additional restrictions on \((\tau,s)\).

\begin{assumption}[Prespecified admissible rule]
\label{ass:admissible_rule}
The correspondence \(\Pi\mapsto\Omega(\Pi)\subseteq\mathbb R^K\) is fixed
before estimation. Membership in \(\Omega(\Pi)\) exactly summarizes the
additional maintained restrictions on \((\tau,s)\) that are not already part
of comparison validity. At the population input \(\Pi_0\), the set
\(\Omega(\Pi_0)\) is nonempty and closed and contains the population treatment
effect and spillover vector.
\end{assumption}

Formally, let \(\tau^c\in\mathbb R\) be a candidate value of \(\tau\), where
the superscript \(c\) labels the candidate. We call it \emph{compatible},
written \(\mathsf{Comp}(\Pi_0,\tau^c)=1\), when the missing post-treatment
no-policy means and spillovers can be completed under comparison validity and
the prespecified admissible rule with \(\tau=\tau^c\).

\begin{proposition}[Treatment effect projection]
\label{prop:treatment_effect_projection}
Suppose Assumptions~\ref{ass:comparison_validity} and
\ref{ass:admissible_rule} hold. Then, for every \(\tau^c\in\mathbb R\),
\begin{equation}
\mathsf{Comp}(\Pi_0,\tau^c)=1
\quad\Longleftrightarrow\quad
\exists x\in\mathcal X_L(\Pi_0):
\ (\tau^c,\tau^c\mathbbm 1_{K-1}-x)\in\Omega(\Pi_0).
\label{eq:sharp_completion_general}
\end{equation}
Consequently, the identified set for the treatment effect is
\begin{equation}\label{eq:sharp_anchor_union}
\begin{aligned}
\Theta_\tau(\Pi_0;L)
&=\{\tau^c\in\mathbb R:\mathsf{Comp}(\Pi_0,\tau^c)=1\}\\
&=
\bigcup_{x\in\mathcal X_L(\Pi_0)}
\left\{
\tau^c\in\mathbb R:
(\tau^c,\tau^c\mathbbm 1_{K-1}-x)\in\Omega(\Pi_0)
\right\}.
\end{aligned}
\end{equation}
Every value in the displayed set is attained by a completion of the
maintained model of potential means.
\end{proposition}

\begin{proof}
See \hyperref[proof:treatment_effect_projection]{Appendix~\ref*{app:candidate_system_representation_proofs}}.
\end{proof}

The result separates the roles of the two parts of the model.
\(\mathcal X_L(\Pi_0)\) contains the vectors of relative effects allowed by the
comparison restrictions, while \(\Omega(\Pi_0)\) determines which treatment
effect levels can accompany each vector. If \(\Omega(\Pi_0)\) leaves the
common-shift direction unrestricted, the treatment effect remains unbounded in
that direction. Ruling out movement in one or both directions produces one-
or two-sided bounds.

Because every value in the projection is attained,
\(\Theta_\tau(\Pi_0;L)\) is sharp relative to the maintained model in
Assumptions~\ref{ass:comparison_validity} and \ref{ass:admissible_rule}. In
particular, \(\Omega(\Pi_0)\) is the complete collection of additional
restrictions on \((\tau,s)\) beyond comparison validity. When the dependence
on a data-generating process (DGP) \(P\) is explicit, write \(\Pi(P)\) for
its population input and
\(\Theta_\tau(P;L):=\Theta_\tau(\Pi(P);L)\).

\subsection{Finite Representation and Treatment Effect Bounds}
\label{subsec:candidate_solvability}

Proposition~\ref{prop:treatment_effect_projection} characterizes candidate
compatibility in terms of \(\mathcal X_L(\Pi_0)\) and \(\Omega(\Pi_0)\), but
that characterization is not yet a finite computational problem:
\(\mathcal X_L(\Pi_0)\) contains a continuum of restrictions, one for each
weight,
while \(\Omega(\Pi_0)\) has so far been left abstract. The next lemma resolves
the first issue. It replaces the two-sided restriction for every \(w\) in the
donor simplex exactly by \(2(K-1)\) linear donor inequalities and box
constraints on two \((T_0-1)\)-dimensional auxiliary vectors.
For \(k=2,\ldots,K\), let \(e_k\in\Delta^{K-2}\) denote the simplex vertex
placing all weight on donor \(k\); hence \(C_t(e_k)\) and
\(C_T^{\mathrm{obs}}(e_k)\) are the corresponding contrasts for individual donors.

\begin{lemma}[Exact finite representation of validity over the full simplex]
\label{lem:finite_comparison_bridge}
For \(L\ge0\), \(x\in\mathcal X_L(\Pi_0)\) if and only if there exist
\[
v^\pm=(v_2^\pm,\ldots,v_{T_0}^\pm)^\top
\in
\left[-\frac{L}{T_0-1},\frac{L}{T_0-1}\right]^{T_0-1}
\]
such that
\begin{equation}\label{eq:finite_comparison_bridge}
\sum_{t=2}^{T_0}v_t^-C_t(e_k)
\le C_T^{\mathrm{obs}}(e_k)-x_k
\le \sum_{t=2}^{T_0}v_t^+C_t(e_k),
\qquad k=2,\ldots,K.
\end{equation}
\end{lemma}

\begin{proof}
See \hyperref[proof:finite_comparison_bridge]{Appendix~\ref*{app:candidate_system_representation_proofs}}.
\end{proof}

The reduction follows from the identity for the support function
\[
\frac{L}{T_0-1}\sum_{t=2}^{T_0}|C_t(w)|
=
\max_{\substack{v\in\mathbb R^{T_0-1}\\
\lVert v\rVert_\infty\le L/(T_0-1)}}
\sum_{t=2}^{T_0}v_tC_t(w).
\]
This identity replaces the fit allowance involving absolute values by a linear
expression in bounded auxiliary coordinates. Because the donor simplex and
the coefficient box are compact and convex and the criterion is bilinear,
Sion's minimax theorem permits the order of optimization to be interchanged.
This interchange yields one auxiliary vector that certifies the restriction
simultaneously for all donor weights. Conditional on that vector, both the
residual after treatment and the lifted allowance are linear in \(w\), so their
maximum over the simplex is attained at a donor vertex \(e_k\). The two signs
of the original restriction involving absolute values require the separate
certificates \(v^+\) and \(v^-\). These vectors are not economic parameters: together with
the finite donor rows in \eqref{eq:finite_comparison_bridge}, they certify all
restrictions for interior weights simultaneously.

Lemma~\ref{lem:finite_comparison_bridge} makes the comparison validity
component finite, but the admissible set
\(\Omega(\Pi)\) remains abstract under
Assumption~\ref{ass:admissible_rule}, which requires neither polyhedrality nor
a particular representation. To turn the entire compatibility problem into a
finite linear system, impose the following additional structure.

\begin{assumption}[Finite linear representation of admissible restrictions]
\label{ass:finite_admissible_representation}
The admissible rule in Assumption~\ref{ass:admissible_rule} has the
representation
\[
\Omega(\Pi)=
\left\{(\tau,s):\exists u\ \text{such that}\quad
A_\tau(\Pi)\tau+A_s(\Pi)s+A_u u\le a(\Pi)
\right\},
\]
where \(A_\tau(\Pi)\), \(A_s(\Pi)\), and \(a(\Pi)\) are conformable affine
functions of \(\Pi\), \(A_u\) is fixed, and \(g\) is a fixed vector satisfying
\begin{equation}
A_\tau(\Pi)+A_s(\Pi)\mathbbm 1_{K-1}=g
\quad\text{for every admissible }\Pi.
\label{eq:fixed_common_shift_loading}
\end{equation}
The auxiliary vector \(u\) has fixed dimension.
An exact equality is represented by two prespecified opposite inequality
rows. All dimensions, coordinates, and units are fixed before estimation.
\end{assumption}

When no auxiliary lift is needed, omit \(u\) and use the direct half-space
representation. The vector \(u\) otherwise accommodates linear lifts, such as
the auxiliary coordinates used to represent restrictions involving absolute
values.

Assumption~\ref{ass:finite_admissible_representation} does two things needed
for computation. Its finite inequality representation turns membership in
\(\Omega(\Pi)\) into finitely many rows. Its common-shift condition ensures
that the candidate \(\tau^c\) enters those rows through the fixed loading
\(g\). Indeed, substituting \(s=\tau^c\mathbbm 1_{K-1}-x\) gives
\begin{equation}
-A_s(\Pi)x+A_u u
\le a(\Pi)-g\tau^c,
\label{eq:common_shift_candidate_rows}
\end{equation}
so the coefficients on the relative effects may depend on estimated population
inputs, but the coefficient on \(\tau^c\) does not. Consequently, estimation
error enters the system in the same way for every candidate value.\footnote{This
structure permits, for example, estimated normalized
exposure weights whose row sums are fixed.} The projection result itself
depends only on the set \(\Omega(\Pi_0)\); the conditions here provide a finite
representation suited to computation and inference.

At a fixed candidate value \(\tau^c\), the preceding two steps now yield one
finite system. Lemma~\ref{lem:finite_comparison_bridge} supplies the comparison
and box rows, while
\eqref{eq:common_shift_candidate_rows} supplies the admissibility rows. The
coefficients \(C_t(e_k)\) and \(C_T^{\mathrm{obs}}(e_k)\) are affine functions
of the population input \(\Pi\) and are evaluated at that input below. Stack
the unknown coordinates as
\[
\eta=(x^\top,u^\top,(v^-)^\top,(v^+)^\top)^\top,
\]
omitting \(u\) when it is unnecessary. The candidate rows are
\begin{equation}\label{eq:bridge_candidate_master_system}
\begin{aligned}
\sum_{t=2}^{T_0}v_t^-C_t(e_k)+x_k
&\le C_T^{\mathrm{obs}}(e_k),
\qquad k=2,\ldots,K,\\
-\sum_{t=2}^{T_0}v_t^+C_t(e_k)-x_k
&\le-C_T^{\mathrm{obs}}(e_k),
\qquad k=2,\ldots,K,\\
-\frac{L}{T_0-1}\mathbbm 1_{T_0-1}\le v^-
&\le\frac{L}{T_0-1}\mathbbm 1_{T_0-1},
\qquad
-\frac{L}{T_0-1}\mathbbm 1_{T_0-1}\le v^+
\le\frac{L}{T_0-1}\mathbbm 1_{T_0-1},\\
-A_s(\Pi)x+A_u u
&\le a(\Pi)-g\tau^c.
\end{aligned}
\end{equation}
Let \(A(\Pi)\) be the coefficient matrix and \(b(\Pi,\tau^c)\) the offset
obtained by stacking the displayed inequalities in the form
\(b(\Pi,\tau^c)-A(\Pi)\eta\le0\). With the displayed coordinates and row order,
write the solution set as
\begin{equation}
\mathcal F(\Pi,\tau^c)
:=
\{\eta:b(\Pi,\tau^c)-A(\Pi)\eta\le0\}.
\label{eq:identification_candidate_feasible_set}
\end{equation}

The following theorem completes the reduction: a candidate is compatible with
the original model of potential means exactly when this linear system has a
solution.

\begin{theorem}[Exact finite candidate representation]
\label{prop:candidate_tau_linear_system}
Suppose Assumptions~\ref{ass:comparison_validity} and
\ref{ass:admissible_rule}--\ref{ass:finite_admissible_representation} hold.
Then, for every \(\tau^c\in\mathbb R\),
\begin{equation}
\mathsf{Comp}(\Pi_0,\tau^c)=1
\quad\Longleftrightarrow\quad
\mathcal F(\Pi_0,\tau^c)\ne\varnothing.
\label{eq:sharp_completion_finite_solvability}
\end{equation}
The finite system introduces neither discretization nor relaxation of the
comparison domain. Every feasible solution can be lifted to a completion of
the maintained model of potential means.
\end{theorem}

\begin{proof}
See \hyperref[proof:candidate_tau_linear_system]{Appendix~\ref*{app:candidate_system_representation_proofs}}.
\end{proof}

Under Assumption~\ref{ass:finite_admissible_representation}, the identified
set in Proposition~\ref{prop:treatment_effect_projection} is also the scalar
projection
\[
\Theta_\tau(\Pi_0;L)
=
\{\tau^c\in\mathbb R:\mathcal F(\Pi_0,\tau^c)\ne\varnothing\}
\]
of the finite candidate system.

\begin{corollary}[Shape of the identified set]
\label{cor:identified_set_shape}
Under the assumptions of
Theorem~\ref{prop:candidate_tau_linear_system}, the population identified set
\(\Theta_\tau(\Pi_0;L)\) is a nonempty polyhedron in \(\mathbb R\). Hence it is
a closed interval, possibly a singleton, a closed half-line, or all of
\(\mathbb R\).
\end{corollary}

\begin{proof}
The joint set
\[
\{(\tau^c,\eta):b(\Pi_0,\tau^c)-A(\Pi_0)\eta\le0\}
\]
is a polyhedron. It is nonempty because the population treatment effect is
compatible under the maintained assumptions. Its projection onto the scalar
\(\tau^c\)-coordinate is therefore a polyhedron in \(\mathbb R\), and every
nonempty polyhedron in \(\mathbb R\) has one of the stated forms.
\end{proof}

The corollary characterizes the shape of the population identified set but
does not determine whether it is bounded above or below. Comparison validity
is invariant to a common shift in the treatment effect and all spillovers, so
boundedness depends on whether the admissible set blocks movement in that
direction. The finite representation in
Assumption~\ref{ass:finite_admissible_representation} does not itself block
this movement; it makes the resulting compatibility problem computationally
tractable. The following proposition gives the exact criterion.

\begin{proposition}[Treatment effect anchoring by admissible restrictions]
\label{prop:exposure_anchoring}
Suppose Assumptions~\ref{ass:comparison_validity} and
\ref{ass:admissible_rule}--\ref{ass:finite_admissible_representation} hold.
Write \(\operatorname{rec}\{\Omega(\Pi_0)\}\) for the recession cone of
\(\Omega(\Pi_0)\), and let
\(\iota=(1,\mathbbm 1_{K-1}^\top)^\top\) denote the direction of a common
shift. The treatment effect
projection is bounded above if and only if
\(\iota\notin\operatorname{rec}\{\Omega(\Pi_0)\}\), bounded below if and only if
\(-\iota\notin\operatorname{rec}\{\Omega(\Pi_0)\}\), and bounded on both sides if
and only if
\[
\operatorname{rec}\{\Omega(\Pi_0)\}\cap\operatorname{span}\{\iota\}=\{0\}.
\]
\end{proposition}

\begin{proof}
See \hyperref[proof:exposure_anchoring]{Appendix~\ref*{app:candidate_system_representation_proofs}}.
\end{proof}

For a direct admissible row,
the corresponding component of
\(g=A_\tau(\Pi)+A_s(\Pi)\mathbbm 1_{K-1}\) is its loading on this direction.
A zero loading means that the row restricts only relative effects; a one-sided
row with nonzero loading may block one ray, and two opposing rows may block
both.\footnote{With auxiliary variables, a recession direction in the lift can
offset an individual row's direct loading, so the recession cone of the
complete admissible set remains the relevant criterion.}

These population results do not impose the same shape on the finite-sample
test-inversion set because the candidate-specific bootstrap critical value may
vary with \(\tau^c\). Compatibility at any fixed candidate is a feasibility
problem for a linear program (LP). Across candidates, the coefficient matrix
remains fixed and \(\tau^c\) changes only the affine right-hand side, which
permits computational reuse and ensures that estimation error enters in the
same way across candidates. The next subsection turns to the interpretation
and calibration of \(L\). Section~\ref{sec:inference} then introduces the
corresponding Farkas certificate and uses it to test fixed candidates.

\subsection{Interpretation and Calibration of \texorpdfstring{\(L\)}{L}}
\label{subsec:comparison_validity_interpretation}

The preceding results characterize identification at a fixed envelope value.
Once \(L\) is fixed, Assumption~\ref{ass:comparison_validity} converts each
weight's pre-treatment discrepancy into an allowance for its missing
post-treatment no-policy gap. The smallest allowance required by the missing
no-policy outcomes is not identified from observed data. I therefore treat
\(L\) as a prespecified sensitivity parameter. Each fixed value defines a
separate maintained model, and the sensitivity path records how the compatible
treatment effects change as more post-treatment discrepancy is allowed.

This creates an interpretation problem: a sensitivity path is useful only if
values such as \(L=1\) or \(L=2\) have a meaningful scale. I use
leave-one-period-out pre-treatment exercises for this purpose. Each observed
pre-treatment period is treated in turn as a pseudo-post period, and the
smallest envelope needed to accommodate that movement is computed. These
placebo values provide empirical benchmarks for the sensitivity scale; they do
not select the post-treatment envelope.

To formalize this distinction, recall the pre-treatment comparison gaps
\(C_t(w)\), and define the latent
post-treatment gap
\[
C_T^0(w)
:=
\Delta\mu_T^1(\mathbf d^0)
-\sum_{k=2}^K w_k\Delta\mu_T^k(\mathbf d^0).
\]
Assumption~\ref{ass:comparison_validity} links fit before treatment to the
latent gap after treatment:
\[
|C_T^0(w)|
\le
\frac{L}{T_0-1}\sum_{t=2}^{T_0}|C_t(w)|
\qquad\text{for every }w\in\Delta^{K-2}.
\]

\begin{definition}[Latent threshold and calibration using observed periods]
\label{def:observed_period_calibration}
Suppose \(T_0\ge3\).
The smallest envelope value required by the latent post-treatment gaps is
\[
L_T^\star
:=
\inf\left\{
\bar L\ge0:
|C_T^0(w)|
\le
\frac{\bar L}{T_0-1}
\sum_{t=2}^{T_0}|C_t(w)|
\quad\text{for every }w\in\Delta^{K-2}
\right\}.
\]
For a pre-treatment period \(\ell\in\{2,\ldots,T_0\}\), define its placebo
calibration index by leaving that period out:
\[
L_\ell^{\mathrm{pl}}
:=
\inf\left\{
\bar L\ge0:
|C_\ell(w)|
\le
\frac{\bar L}{T_0-2}
\sum_{\substack{t=2\\t\ne\ell}}^{T_0}|C_t(w)|
\quad\text{for every }w\in\Delta^{K-2}
\right\}.
\]
The infimum is \(+\infty\) if the corresponding set is empty.
\end{definition}

Assumption~\ref{ass:comparison_validity} at a fixed \(L\) holds if and only if
\(L\ge L_T^\star\). The threshold \(L_T^\star\) depends on missing no-policy
means and is not identified, whereas each \(L_\ell^{\mathrm{pl}}\) depends only
on pre-treatment population means. Following \citet{hsu2013calibrating}, the
placebo indices express candidate values of \(L\) in units of observed
pre-treatment instability. They neither estimate nor automatically bound
\(L_T^\star\), and choosing the sensitivity range remains a substantive
decision.

A low-rank factor model provides one primitive interpretation of the common
envelope.
Suppose untreated first differences satisfy
\[
\Delta\mu_t^k(\mathbf d^0)=\alpha_t+f_t^\top h_k,
\qquad t=2,\ldots,T_0,T,\quad k=1,\ldots,K,
\]
where \(\alpha_t\) is a common increment, \(f_t\) is a factor vector, and
\(h_k\) is unit \(k\)'s loading vector. Write
\(F_{\mathrm{pre}}=[f_2,\ldots,f_{T_0}]\). If
\(f_T=F_{\mathrm{pre}}\gamma\) for some
\(\gamma\in\mathbb R^{T_0-1}\), then the common increment \(\alpha_t\) cancels
and, simultaneously for every \(w\),
\[
|C_T^0(w)|
\le
\|\gamma\|_\infty
\sum_{t=2}^{T_0}|C_t(w)|.
\]
Thus Assumption~\ref{ass:comparison_validity} holds for all weights whenever
\(L\ge(T_0-1)\|\gamma\|_\infty\). A single restriction on factor movements
can imply validity for every weight. Applying the same argument with an
observed pre-treatment period held out gives the placebo counterpart. The
factor structure can rationalize a common envelope across weights, but it does
not identify the envelope required after treatment. Appendix
Example~\ref{ex:factor_comparison_validity} gives both derivations.

The envelope is positively homogeneous, so a population weight with exact fit
receives zero allowance for its latent post-treatment gap. The
relative-magnitude restriction of \citet{rambachan2023more}, although it bounds
changes in parallel trends violations rather than latent gaps for individual
weights, has the same implication. Under the condition on the factor span,
zero factor gaps before treatment force the gap after treatment to zero.
Appendix Remark~\ref{rem:additive_floor_extension} develops a relaxation
with an additive floor.

The baseline score averages absolute annual pre-treatment discrepancies, so
\(L\) bounds a weight's latent post-treatment gap relative to its typical
annual discrepancy. When the post-treatment target spans a different horizon,
as in the LAWA application, \(L\) also accounts for extrapolation between the
different horizons. The placebo indices based on smoothed factors in the application provide
annual benchmarks on this scale.\footnote{An application may
substitute another prespecified fit score with a clear interpretation.
Appendix Remark~\ref{rem:additive_floor_extension} gives an exact finite
representation for a nonnegative fit-independent floor. Holding the remaining
specification fixed, increasing this floor weakly enlarges the population
compatibility set.}

The finite representation depends on the prespecified collection of pre-treatment gap
coordinates \(C_2(w),\ldots,C_{T_0}(w)\), not on annual first differences per
se.\footnote{The coordinates may instead be constructed from levels, long
differences, residualized gaps, or horizon-matched contrasts.} Changing them
changes the economic content and scale of \(L\), but not the finite representation or the
geometry of relative effects. The baseline based on annual first differences provides eight
transparent pre-treatment movements for both the validity score and its
placebo calibration.

This scalar envelope belongs to the broader literature on partial identification,
which weakens point-identifying benchmark restrictions in explicit and
interpretable ways \citep[e.g.,][]{manski2000monotone,manski2018right,
rambachan2023more}. Here the restriction bounds a latent gap for each weight in
equations that also contain the unknown spillover vector \(s\).

\section{Inference}\label{sec:inference}

Both identification and inference ask whether a given value \(\tau^c\) is
compatible with the model. The null does not compare \(\tau^c\) with an
estimated endpoint. It asks whether some values of the unobserved no-policy
means and spillovers satisfy all assumptions when \(\tau=\tau^c\).
Theorem~\ref{prop:candidate_tau_linear_system} turns this into checking whether
a finite linear system has a solution, and Farkas' alternative provides a
certificate when it does not. I use the normalized solvability statistic and
bootstrap calibration of \citet{goff2025inference} for this structured family,
establish uniform validity under model-tailored regularity conditions, and
invert the tests over candidate values. The systems share a candidate-invariant
primitive sampling perturbation. A uniform feasibility error bound is imposed
only over model-generated compatible systems, and the variance condition is
restricted to structurally reachable certificate directions that can determine
rejection.

The donor pool, envelope value \(L\), admissible rule, and finite row
representation are held fixed throughout. Section~\ref{sec:candidate_lp_anatomy}
constructs a test for one candidate by turning sample infeasibility into a
normalized Farkas score and calibrating it with the bootstrap.
Section~\ref{sec:test_inversion_over_candidates} then inverts these tests, and
Section~\ref{subsec:uniform_validity} states their candidatewise coverage.

\subsection{Compatibility Test for a Fixed Candidate}
\label{sec:candidate_lp_anatomy}

For a fixed candidate \(\tau^c\), inference asks whether the sample evidence of
incompatibility is larger than estimation error can explain. Recall that
\(\eta\) collects the relative effects and auxiliary variables in the finite
representation. By Theorem~\ref{prop:candidate_tau_linear_system}, the
candidate is compatible at a DGP \(P\) exactly when
\(b(\Pi(P),\tau^c)-A(\Pi(P))\eta\le0\) has a solution. For sample size \(n\),
let \(\hat\Pi_n\) be the sample analogue of \(\Pi(P)\), and write
\(\hat A_n:=A(\hat\Pi_n)\) and
\(\hat b_n(\tau^c):=b(\hat\Pi_n,\tau^c)\).

Farkas' alternative turns failure of this system into a certificate. To see
why, suppose the population system were feasible. Multiplying its rows by any
\(\lambda\ge0\) and summing gives
\(b(\Pi(P),\tau^c)^\top\lambda
-\eta^\top A(\Pi(P))^\top\lambda\le0\). If
\(A(\Pi(P))^\top\lambda=0\), feasibility therefore requires
\(b(\Pi(P),\tau^c)^\top\lambda\le0\). Farkas' alternative states that the
converse also holds:
\begin{equation}
\mathcal F(\Pi(P),\tau^c)=\varnothing
\quad\Longleftrightarrow\quad
\exists\lambda\ge0:\
A(\Pi(P))^\top\lambda=0,\quad
b(\Pi(P),\tau^c)^\top\lambda>0.
\label{eq:farkas_candidate_alternative}
\end{equation}
Thus \(\lambda\) is a nonnegative combination of the system rows that cancels
the unknown \(\eta\) and leaves a contradiction.

The sample system replaces \(\Pi(P)\) by \(\hat\Pi_n\). Even when the
population system is feasible, it may lie on the boundary, where arbitrarily
small estimation error can make the sample system infeasible. Sample
infeasibility alone is therefore not valid evidence against compatibility.
Moreover, a Farkas certificate is scale-free: multiplying \(\lambda\) by a
positive constant gives the same contradiction. An unweighted normalization
would make the score depend on the arbitrary units used to write individual
inequality rows. Instead, let \(\widehat{\mathsf S}_n\) scale each row by the
largest bootstrap-estimated standard deviation among its estimated entries;
fixed rows receive zero scale. Weighting by these row scales makes the score
invariant to positive rescaling of a row and expresses a contradiction relative
to its sampling uncertainty. The normalized sample score is
\begin{equation}\label{eq:sample_uniform_infeasibility}
\hat Q_n(\tau^c)
:=
\sup_{\substack{\lambda\ge0:\
\hat A_n^\top\lambda=0\\
\mathbbm 1^\top\widehat{\mathsf S}_n\lambda\le1}}
\hat b_n(\tau^c)^\top\lambda.
\end{equation}
It is zero when the sample system is feasible, positive when a normalized
certificate separates it, and may be infinite when fixed rows alone certify
incompatibility.

The bootstrap cannot simply treat the estimated system as the null system. A
compatible population system may lie on the boundary of feasibility, so its
sample analogue can be infeasible with nonvanishing probability. The procedure
therefore recenters the bootstrap perturbation of the estimated system at a
near-feasible completion and evaluates it over certificate directions that are
nearly optimal for the candidate. Its bootstrap distribution measures how
large the normalized infeasibility score can be when the candidate remains
compatible.

Fix a nominal test level \(\alpha\in(0,1/2)\). The bootstrap estimates how
large the score can be from sampling error at a compatible boundary. The
following procedure summarizes the test; Appendix equations
\eqref{eq:uniform_eta_hat}--\eqref{eq:uniform_bootstrap_stat} give the exact
recentering and certificate calculations.

\begin{procedure}[Compatibility test for a fixed candidate]
\label{proc:fixed_candidate_compatibility}
Within the fixed specification, choose \(\tau^c\in\mathbb R\).
\begin{enumerate}[label=\arabic*.]
\item Construct the sample system
\(\hat b_n(\tau^c)-\hat A_n\eta\le0\) and compute
\(\hat Q_n(\tau^c)\).
\item Construct the near-feasible completion, perturb the estimated system
with the bootstrap, and maximize over the normalized certificates that are
nearly optimal for this candidate. Let \(\hat c_n(\tau^c,1-\alpha)\) be the
resulting conditional critical value.
\item Do not reject compatibility when
\[
\hat Q_n(\tau^c)<\infty
\quad\text{and}\quad
\sqrt n\,\hat Q_n(\tau^c)\le\hat c_n(\tau^c,1-\alpha).
\]
\end{enumerate}
\end{procedure}

\subsection{Test Inversion over Candidate Values}
\label{sec:test_inversion_over_candidates}

Let \(\mathcal T\subset\mathbb R\) be a compact reporting domain. It limits
the range over which the test is inverted; it does not enter the maintained
model or the population identified set. Inverting
Procedure~\ref{proc:fixed_candidate_compatibility} over this domain gives
the inversion set
\begin{equation}\label{eq:uniform_inverted_cs}
\mathcal I_n
:=
\left\{
\tau^c\in\mathcal T:
\hat Q_n(\tau^c)<\infty,\quad
\sqrt n\,\hat Q_n(\tau^c)
\le \hat c_n(\tau^c,1-\alpha)
\right\}.
\end{equation}

This is direct inversion of compatibility tests rather than inference based on
plug-in estimates of projection endpoints. Its coverage interpretation is
candidatewise: each compatible structural value is protected against false
exclusion by its own level-\(\alpha\) test. Section~\ref{subsec:uniform_validity}
states the corresponding uniform result.

The model structure also allows
substantial computational reuse within a fixed specification. The matrix from
the finite representation, row scales, and underlying bootstrap perturbations are constructed
once. Since \(\tau^c\) changes only the affine right-hand side, the feasible region
of the sample Farkas LP is common across candidates and only its objective
changes. At a given candidate, the bootstrap certificate LP has a common
feasible region across draws and only its objective changes.
The critical value nevertheless remains candidate-specific because the
near-feasible completion and near-optimal certificate set depend on
\(\tau^c\).

The set in \eqref{eq:uniform_inverted_cs} is defined by continuum inversion. In
practice, the application approximates it on an adaptive grid.\footnote{The
coverage result applies to each compatible candidate that is actually tested;
interpolation between tested values is a numerical display, not an additional
coverage claim. Appendix~\ref{app:arizona_policy_coding} records the grid,
tolerance, solver, search anchors, and audit using a denser grid.}

\subsection{Candidatewise Validity}\label{subsec:uniform_validity}

Let \(\mathcal P^U\) denote the class of DGPs over which the tests are required
to be uniformly valid. The attainment result in
Proposition~\ref{prop:treatment_effect_projection} implies that every compatible
candidate value can be the structural effect in a completion of the maintained
population model.
Accordingly, define the null index class
\[
\mathcal Q_0
:=
\{(P,\tau^c):P\in\mathcal P^U,\ \tau^c\in\mathcal T,\
\ \mathcal F(\Pi(P),\tau^c)\ne\varnothing\}.
\]
Write \(\mathbb{P}_P\) for probability under \(P\). The inferential target is
\begin{equation}
\limsup_{n\to\infty}
\sup_{(P,\tau^c)\in\mathcal Q_0}
\mathbb{P}_P\{\tau^c\notin\mathcal I_n\}
\le\alpha.
\label{eq:robust_candidate_coverage_target}
\end{equation}
This controls false exclusion whichever compatible value is the structural
truth.

\begin{assumption}[First-order sampling of aggregate inputs]
\label{ass:sampling}
The dimension of \(\Pi(P)\) is fixed and \(\Pi(P)\) lies in a fixed compact
set. Let \(\hat\Pi_n^*\) denote the input recomputed under the bootstrap, and
define
\[
Z_{n,P}:=\sqrt n\{\hat\Pi_n-\Pi(P)\},
\qquad
Z_n^*:=\sqrt n\{\hat\Pi_n^*-\hat\Pi_n\}.
\]
Uniformly over \(P\in\mathcal P^U\), \(Z_{n,P}\) admits a bounded-Lipschitz
Gaussian approximation with a uniformly bounded, possibly singular covariance
matrix \(\Sigma_\Pi(P)\). Conditionally on the data, \(Z_n^*\) consistently
reproduces this law and covariance, uniformly in probability.
\end{assumption}

Appendix Lemma~\ref{lem:cluster_ratio_sampling_sufficient} gives primitive
sufficient conditions for weighted aggregate means formed from independent
sampling clusters. Inputs observed without sampling error enter the candidate
system as fixed components.

The finite row representation is affine in \(\Pi\), while \(\tau^c\) has a
fixed loading on its right-hand side. The population, sample, and bootstrap
systems therefore share the same first-order sampling perturbation at every
candidate; indexing the test by \(\tau^c\) introduces no additional stochastic
equicontinuity requirement.

For each system row \(r\), let \(\sigma_r(P)\) be the largest asymptotic
standard deviation of the \(\sqrt n\)-scaled estimation errors among the
data-estimated entries in row \(r\) of
\([\hat A_n\ \ \hat b_n(\tau^c)]\). Set \(\sigma_r(P)=0\) when the row
contains only fixed entries, and define
\[
\mathsf S(P):=\operatorname{diag}\{\sigma_r(P)\}_r.
\]
These row scales do not depend on \(\tau^c\). Because the system has finitely
many entries with fixed loadings on \(\Pi\), the covariance bound in
Assumption~\ref{ass:sampling} implies a common finite upper bound \(\bar d\)
on their asymptotic standard deviations. Fixed entries have zero scale by
construction. The population analogues of the normalized dual region and
sample score are
\[
\begin{aligned}
\mathcal D(P)
&:=
\{\lambda\ge0:A(\Pi(P))^\top\lambda=0,
\ \mathbbm 1^\top\mathsf S(P)\lambda\le1\},\\
Q(P,\tau^c)
&:=\sup_{\lambda\in\mathcal D(P)}
b(\Pi(P),\tau^c)^\top\lambda.
\end{aligned}
\]
At a compatible candidate, \(Q(P,\tau^c)=0\); a nonzero maximizer is a
certificate direction that is potentially binding at the feasibility
boundary.

To describe sampling variation along these directions, let
\(\eta^0(P,\tau^c)\) be the minimum-norm feasible nuisance vector and write
\(D_\Pi\) for differentiation with respect to \(\Pi\). Define the first-order
sensitivity of the row residuals at this completion by
\[
\Gamma(P,\tau^c)
:=
\left.
D_\Pi\!\left[
b(\Pi,\tau^c)-A(\Pi)\eta^0(P,\tau^c)
\right]
\right|_{\Pi=\Pi(P)}.
\]
\begin{assumption}[Regular candidate systems]
\label{ass:regular_candidate_systems}
There exist constants
\(\underline d,c_\Gamma,\delta_\Gamma,\varepsilon_\Gamma>0\) and
\(\bar H<\infty\), common to the relevant DGP and null classes, with
\(\delta_\Gamma\le\min\{\underline d/2,\bar d\}\), such that:
\begin{enumerate}[label=(\roman*)]
\item every system entry classified as estimated has asymptotic standard
deviation at least \(\underline d\);
\item for every \((P,\tau^c)\in\mathcal Q_0\) and every \(\eta\),
\[
\operatorname{dist}\{\eta,\mathcal F(\Pi(P),\tau^c)\}
\le
\bar H\left\|[b(\Pi(P),\tau^c)-A(\Pi(P))\eta]_+\right\|.
\]
\item for every \((P,\tau^c)\in\mathcal Q_0\), if the population program
defining \(Q(P,\tau^c)\) has a nonzero
maximizer, at least one such maximizer \(\lambda\) satisfies
\[
\lambda^\top\Gamma(P,\tau^c)\Sigma_\Pi(P)
\Gamma(P,\tau^c)^\top\lambda\ge c_\Gamma.
\]
For each \(\widetilde\Pi\) in the input domain such that
\(\widetilde\Pi-\Pi(P)\) is supported on the coordinates estimated from the
sampling data, and each diagonal \(\widetilde{\mathsf S}\) having the same
zero pattern as \(\mathsf S(P)\), define
\[
\widetilde A:=A(\widetilde\Pi),
\qquad
\widetilde b:=b(\widetilde\Pi,\tau^c),
\qquad
\widetilde{\mathcal D}
:=
\{\lambda\ge0:\widetilde A^\top\lambda=0,
\ \mathbbm 1^\top\widetilde{\mathsf S}\lambda\le1\}.
\]
If zero is the unique maximizer of the population program, then, for every
such \((\widetilde\Pi,\widetilde{\mathsf S})\) satisfying
\[
\|\widetilde A-A(\Pi(P))\|
+\|\widetilde b-b(\Pi(P),\tau^c)\|
+\|\widetilde{\mathsf S}-\mathsf S(P)\|
\le\delta_\Gamma,
\]
every
\(\lambda\in\operatorname{ext}(\widetilde{\mathcal D})\setminus\{0\}\)
satisfying \(\widetilde b^\top\lambda\ge-\varepsilon_\Gamma\) obeys
\[
\lambda^\top\Gamma(P,\tau^c)\Sigma_\Pi(P)
\Gamma(P,\tau^c)^\top\lambda\ge c_\Gamma.
\]
\end{enumerate}
\end{assumption}

Part (i) rules out vanishing first-order variation in entries classified as
estimated; the restriction on \(\delta_\Gamma\) preserves this nondegeneracy
for nearby normalizers. Part (ii) is a uniform feasibility error bound over
model-generated compatible systems: row violations control distance to the
latent feasible completion. Part (iii) requires sampling variation in
certificate directions that can determine rejection; its second clause covers
nearby directions when zero is the unique population maximizer.

\begin{theorem}[Uniform candidatewise coverage]
\label{thm:uniform_candidate_test_validity}
Suppose Assumptions~\ref{ass:comparison_validity} and
\ref{ass:admissible_rule}--\ref{ass:finite_admissible_representation},
\ref{ass:sampling}, and
\ref{ass:regular_candidate_systems} hold uniformly over \(\mathcal P^U\).
Then the set in
\eqref{eq:uniform_inverted_cs} satisfies, for every
\(\alpha\in(0,1/2)\),
\[
\limsup_{n\to\infty}
\sup_{P\in\mathcal P^U}
\sup_{\tau^c\in\Theta_\tau(P;L)\cap\mathcal T}
\mathbb{P}_P\{\tau^c\notin\mathcal I_n\}
\le \alpha .
\]
\end{theorem}

\begin{proof}
See \hyperref[proof:uniform_candidate_test_validity]{Appendix~\ref*{app:estimated_system_regularity}}.
\end{proof}

The theorem is uniform over DGPs and compatible candidates, but its event is
candidatewise: it controls exclusion of one compatible structural value at a
time. It does not assert simultaneous inclusion of the complete identified
set. Inverting the candidate tests therefore requires no multiplicity
adjustment for this candidatewise guarantee.

Additional theoretical results are collected in the
Appendix.\footnote{Proposition~\ref{prop:separated_incompatibility_power}
establishes power against separated incompatibility. Appendix
Remark~\ref{rem:conservative_calibration} gives a conservative calibration that
does not require part (iii) of
Assumption~\ref{ass:regular_candidate_systems}.}

The next section specifies the finite mean input and admissible restriction for
LAWA and reports inversions for fixed values of \(L\) over a prespecified
sensitivity grid.

\section{Application: The Legal Arizona Workers Act}\label{sec:empirical_illustration}

I apply the framework to the Legal Arizona Workers Act (LAWA), studied by
\citet{bohn2014lawa} as a policy that may have reduced Arizona's likely
unauthorized immigrant population. LAWA mandated E-Verify for new hires and
took effect on January 1, 2008. Because legal status is not observed in
standard household surveys, \citet{bohn2014lawa} use population shares based
on Hispanic noncitizens as proxies for the policy target. Their synthetic
control estimates indicate a sizable decline in Arizona after LAWA.

Migration makes this a natural setting in which comparison states may also be
affected. A
policy that changes the attractiveness of Arizona can redirect prospective
migrants, induce current residents to move, and alter migration into or out of
the United States. \citet{ellis2014migrationresponse} study interstate
out-migration from Arizona, while \citet{amuedodorantes2019interstate} examine
the destination states of Arizona out-migrants. \citet{orrenius2016everify}
document diversion of new arrivals under E-Verify laws. I use this evidence
to motivate possible spillovers, but I do
not require the Arizona effect to be offset exactly within the observed
comparison pool. Instead, I place an interpretable upper bound on total
absolute spillovers across comparison states.

\subsection{Data and Benchmark Construction}

The empirical implementation uses Version 13.0 of the IPUMS Current
Population Survey (CPS) \citep{flood2025ipumscps} to construct annual
state-level series for 1998--2009. I aggregate Basic Monthly records to
state-year means using the final person-level weight \texttt{WTFINL}.
The outcome is the share of
Hispanic noncitizens in the civilian noninstitutional population represented
in the CPS, one of the measures used by \citet{bohn2014lawa}. The sample
contains \(1{,}216{,}961\) household
trajectories linked by the CPS household identifier (CPSID). For state \(k\)
and year \(t\), the empirical cell mean is
\[
\hat\mu_t^k
=
\frac{\sum_{i\in(k,t)} \omega_i Y_i}{\sum_{i\in(k,t)} \omega_i},
\]
where \(Y_i\) is the outcome indicator and \(\omega_i\) is \texttt{WTFINL}.

\begin{samepage}
Arizona is treated. Following \citet{bohn2014lawa}, I exclude Mississippi,
Rhode Island, South Carolina, and Utah from the donor pool.\footnote{The
excluded states had broadly applied restrictions on the employment of
undocumented immigrants.} The pool contains 46 units: 45 states and the
District of Columbia. Let
\(\mathcal K_{\mathrm{cmp}}\) denote this set. The
comparison domain is its full simplex. I use 1998--2006 as the pre-treatment
window, omit 2007 as a transition year, and define the post-treatment target
from the 2008--2009 average.
Inference reflects household-cluster sampling variation conditional on the
policy timing, donor pool, CPS \texttt{WTFINL} weights, and external population
constants; Appendix~\ref{app:arizona_policy_coding} gives the bootstrap design.
\end{samepage}

Relative to \citet{bohn2014lawa}, the main design change is how I construct the
outcome changes before and after treatment. The eight changes from 1999 through
2006 summarize the pre-treatment period, and the 2008--2009 average relative
to the 2006 level is the post-treatment outcome change.
Appendix~\ref{app:arizona_policy_coding} describes the external population
scaling weights and the alternative outcome for the working-age population
with low education.

Figure~\ref{fig:lawa_raw_outcome_trends} displays the annual cell means before
the identifying restrictions are imposed. It shows the entire donor pool
rather than a selected synthetic series because the framework does not
privilege a single comparison weight. Arizona's post-2007 decline is therefore
a feature of the observed data, not by itself an estimate of the policy effect.

\begin{figure}[!htbp]
\centering
\includegraphics[width=0.94\textwidth]{lawa_raw_outcome_trends.pdf}
\caption{Observed LAWA Outcome Paths in Arizona and the Donor Pool}
\label{fig:lawa_raw_outcome_trends}
\begin{minipage}{0.94\textwidth}
\footnotesize
\textit{Notes:} The outcome is the share of Hispanic noncitizens in the
civilian noninstitutional population. The red line is Arizona, the thin gray
lines are the 46 donor units, and the dashed blue line is their unweighted
median across states. All shares use \texttt{WTFINL}.
The gray band marks the omitted 2007
transition year; the blue band marks the 2008--2009 post-treatment target
window. The figure is descriptive: it neither selects a comparison weight nor
imposes the paper's identifying restrictions.
\end{minipage}
\end{figure}

\subsection{Empirical Restrictions}

Comparison validity restricts differences between the treatment effect and
spillovers but does not determine their overall level. I therefore impose a
substantive bound on total absolute spillovers after scaling each state's share
response by population. Because LAWA is an Arizona policy, the total response
across comparison states should not be arbitrarily large relative to the
population potentially affected in Arizona. Specifically, I bound it by a
prespecified multiplier \(\rho>0\) times the size of Arizona's target
population without LAWA. This
restriction does not require spillovers to share a sign or to offset Arizona's
treatment effect.

A natural reference value is \(\rho=1\). A one-person-equivalent response in
one comparison state contributes one unit to the gross-response scale. Using
the 2006 population scaling defined below, \(\rho=1\) therefore permits a total
resident-population-equivalent response as large as Arizona's entire target
population without LAWA.
This benchmark is already broad because it uses the full population of
Hispanic noncitizens rather than only new hires or workers directly covered by
LAWA. The main specification sets \(\rho=2\), allowing twice this benchmark to
accommodate diverted entry, further relocation, and offsetting responses across
comparison states. I report
\(\rho=1\) and \(\rho=4\) as tighter and looser sensitivity specifications.
These values are prespecified calibrations, not estimated parameters.

To put the share responses on a common population scale, let
\(N_k^{\mathrm{ref}}\) denote comparison unit \(k\)'s externally measured 2006
resident population for \(k\in\mathcal K_{\mathrm{cmp}}\), and define
\[
q_k=\frac{N_k^{\mathrm{ref}}}{N_{\mathrm{AZ}}^{\mathrm{ref}}},
\]
where \(N_{\mathrm{AZ}}^{\mathrm{ref}}\) is the corresponding Arizona reference population.
I construct these constants from the U.S. Census Bureau's
\texttt{ST-EST00INT-AGESEX} state intercensal population file
\citep{uscensus2012intercensal}. They are fixed before constructing the CPS
outcome moments and held fixed in the bootstrap. Multiplying a share effect by
its reference population expresses the change on a
resident-population-equivalent scale using the 2006 population base.\footnote{
This is a standardization of the effect, not an assertion that population
denominators after treatment remain equal to their 2006 values. The Census
resident-population base is not identical to the CPS civilian
noninstitutional outcome universe, so the resulting quantities are
resident-population-equivalent responses rather than literal counts of movers
or affected people.}

Let
\(
\mu_T^{\mathrm{AZ}}(\mathbf d^0)
=\mu_T^{\mathrm{AZ}}(\mathbf d^1)-\tau
\)
denote Arizona's no-policy post-treatment share. For a prespecified
\(\rho>0\), the application imposes the gross spillover budget
\begin{equation}\label{eq:empirical_gross_spillover_budget}
\sum_{k\in\mathcal K_{\mathrm{cmp}}}q_k|s_k|
\le
\rho\,\mu_T^{\mathrm{AZ}}(\mathbf d^0).
\end{equation}
In units based on the 2006 populations, this restriction is
\[
\sum_{k\in\mathcal K_{\mathrm{cmp}}}N_k^{\mathrm{ref}}|s_k|
\le
\rho\,N_{\mathrm{AZ}}^{\mathrm{ref}}\mu_T^{\mathrm{AZ}}(\mathbf d^0).
\]
The left-hand side is the total absolute response across comparison units, and
the right-hand side is \(\rho\) times Arizona's target population without LAWA.
Scaling by the population without LAWA rather than the observed population
after LAWA keeps the allowance tied to the population that could have been
affected in the absence of LAWA. This is a scale restriction rather than a conservation equation: it
does not equate Arizona's treatment effect with the net response across
comparison states, and adjustment through international movement or excluded states
need not appear on the left-hand side. At
\(\tau=\tau^c\), the normalized budget on the right-hand side of
\eqref{eq:empirical_gross_spillover_budget} equals
\(\rho\{\mu_T^{\mathrm{AZ}}(\mathbf d^1)-\tau^c\}\). A more negative candidate
implies a larger target population without LAWA and therefore a larger
admissible total response across comparison units. The budget is consequently asymmetric in
\(\tau^c\), which can contribute to different behavior at the lower and upper
endpoints of the compatibility set.

The application also imposes the known support of the share outcome:
\[
0\le \mu_T^{\mathrm{AZ}}(\mathbf d^1)-\tau\le1,
\qquad
0\le \mu_T^k(\mathbf d^1)-s_k\le1,
\quad k\in\mathcal K_{\mathrm{cmp}}.
\]
These inequalities rule out counterfactual shares outside \([0,1]\) and ensure
that the no-policy Arizona stock on the right-hand side of
\eqref{eq:empirical_gross_spillover_budget} is nonnegative.

The gross spillover budget can be written exactly as a finite linear system.
Introduce \(z_k\ge0\) satisfying
\(-z_k\le s_k\le z_k\). Then
\[
\sum_{k\in\mathcal K_{\mathrm{cmp}}}q_k z_k+\rho\tau
\le
\rho\,\mu_T^{\mathrm{AZ}}(\mathbf d^1).
\]
Together with outcome support, these rows define
\(\Omega^{\mathrm{emp}}_\rho(\Pi_0)\) and satisfy
Assumption~\ref{ass:finite_admissible_representation}.\footnote{Bootstrap
draws recompute these post-treatment means.}
The three reported calibrations are nested:
\[
\Omega^{\mathrm{emp}}_1(\Pi_0)
\subseteq
\Omega^{\mathrm{emp}}_2(\Pi_0)
\subseteq
\Omega^{\mathrm{emp}}_4(\Pi_0).
\]

\subsection{Calibration of the Validity Scale Using Observed Periods}
\label{subsec:empirical_l_calibration}

Because \(L\) compares an unobserved gap after treatment with observed fit
before treatment, its numerical scale is difficult to interpret. I therefore
use the placebo indices in
Definition~\ref{def:observed_period_calibration}. The 1998--2006 pre-treatment
window contains eight annual first differences. For each factor dimension
\(d_f\in\{1,2,3\}\), I hold out one difference at a time, estimate the factor
approximation on the other seven, and compute the smallest envelope over the
full simplex that covers the held-out fitted movement. Denote the resulting
envelope by \(\widehat L_{\ell,d_f}^{\mathrm{fac}}\).
Appendix Procedure~\ref{proc:factor_placebo_lbr} gives the calculation.

Let
\[
\widehat B_{d_f}
:=
\max_{\ell=2,\ldots,T_0}
\widehat L_{\ell,d_f}^{\mathrm{fac}}
\]
denote the most demanding held-out movement for factor dimension \(d_f\).
For the main outcome,
\[
(\widehat B_1,\widehat B_2,\widehat B_3)
=
(2.332,\ 2.684,\ 1.526).
\]
The reported path \(L\in[1,3]\) therefore begins below all three
benchmarks and ends above all of them. At \(L=2\), the envelope exceeds the
benchmark based on three factors but not those based on one or two. The
alternative subgroup's corresponding maxima are
\((1.800,\,1.770,\,1.466)\) and serve only as a scale check. Appendix
Table~\ref{tab:arizona_lbr_placebo_full} reports all period-specific indices.

Following \citet{hsu2013calibrating}, this exercise expresses an abstract
sensitivity parameter in units of observed variation. It also follows the idea
in \citet{rambachan2023more} of using movements before treatment to interpret a
departure after treatment. The benchmarks neither select \(L\) nor enter test
construction.

\subsection{Test Inversion Results}

For each \(L\in\{1,1.25,\ldots,3\}\) and \(\rho\in\{1,2,4\}\), I construct a
95 percent inversion set by inverting the candidate compatibility tests over
the full donor simplex. All specifications impose outcome support and use the
homogeneous comparison envelope with a zero floor. The main specification
sets \(\rho=2\); the \(\rho=1\) and \(\rho=4\) specifications vary only the
gross spillover budget.

For reference, using nine pre-intervention years, Table 3 of
\citet{bohn2014lawa} reports a decline of 1.50 percentage points for the share
of all residents who are Hispanic noncitizens, with a one-sided placebo
\(p\)-value of \(0.021\). The inversion sets below show whether zero or the
original estimate is rejected and characterize the other treatment effect
values not rejected under the maintained comparison validity envelopes and
gross spillover budgets.

Figure~\ref{fig:arizona_gm_inversion_paths} reports the lower and upper
boundaries of the 95 percent inversion sets obtained from the compatibility
tests and the adaptive candidate grid. These endpoints numerically approximate
the inversion over all candidate values in
Theorem~\ref{thm:uniform_candidate_test_validity}. In
the audit using a denser grid reported in Appendix~\ref{app:arizona_policy_coding},
every endpoint changes by at most \(1.2\times10^{-4}\).

\begin{figure}[!htbp]
\centering
\includegraphics[width=0.94\textwidth]{arizona_gm_inversion_rho_paths.pdf}
\caption{Inversion of LAWA Compatibility Tests Across \(L\) and \(\rho\)}
\label{fig:arizona_gm_inversion_paths}
\begin{minipage}{0.94\textwidth}
\footnotesize
\textit{Notes:} Each pair of lines with the same style gives the numerically
approximated lower and upper boundary of the 95 percent inversion set for the
indicated gross spillover budget. Coverage is candidatewise within each
fixed \((L,\rho)\) specification. The paths connect nine separately evaluated
specifications with fixed \(L\) from \(1\) to \(3\) in increments of \(0.25\); they
provide neither simultaneous coverage of the complete identified set nor a
simultaneous band over \(L\) or \(\rho\). The horizontal dashed line marks
zero. The three vertical gray lines mark the maximum factor placebo index for
specifications with one, two, or three factors.
\end{minipage}
\end{figure}

The main \(\rho=2\) inversion set is
\([-4.19,0.17]\) percentage points at \(L=1\),
\([-4.66,0.64]\) at \(L=2\), and
\([-5.13,1.05]\) at \(L=3\). Thus the inversion set contains zero even at the
smallest reported envelope, \(L=1\), which lies below all three factor
placebo benchmarks. Relaxing comparison validity further widens the
range of treatment effects not rejected as compatible.

Allowing a larger gross response primarily extends the negative boundary. At
\(L=2\), the widths of the inversion sets are \(5.03\), \(5.30\), and \(5.83\)
percentage points for \(\rho=1,2,4\), respectively. The upper boundaries are
not ordered in the same way. Population compatibility sets are nested in
\(\rho\), but the separately constructed inversion sets need not be: both the
near-optimal certificate set and its bootstrap critical value depend on the
specification.

Across all reported \((L,\rho)\) specifications, the inversion sets contain
both zero and the decline of 1.50 percentage points reported by
\citet{bohn2014lawa}. The maintained model and sampling uncertainty therefore
do not yield a sign conclusion: neither zero nor a decline of the original
magnitude is rejected. This conclusion is conditional on validity over the full
simplex and the reported gross spillover budgets.

\section{Monte Carlo Evidence}\label{sec:monte_carlo}

The simulation addresses two questions. First, how much do restrictions over
all convex weights narrow the population set when aggregation cancels
pre-treatment discrepancies? Second, how do the tests perform in these
settings? The design holds the empirical post-treatment inputs and admissible
rule fixed while varying only the pattern of pre-treatment discrepancies and
sampling precision.

\subsection{Design and Population Geometry}

The sampling experiment for both designs uses the full simplex comparison
domain, \(L=2\), the outcome support restrictions, and the main
gross spillover budget \(\rho=2\). The population exercise additionally solves
the benchmark using only donor vertices. Both designs preserve each comparison unit's
empirical total absolute pre-treatment discrepancy. The \emph{sign-aligned}
design makes every pre-treatment gap for every donor and period nonnegative.
Convex aggregation then
cannot reduce fit through sign cancellation. The \emph{near-exact} design
places five percent of each state's absolute discrepancy in a common positive
component and chooses donor signs for the remaining 95 percent to minimize the
fit of the convex comparison using population weights. The common component rules out an
exact convex fit. For \(q=(q_k)_{k\in\mathcal K_{\mathrm{cmp}}}\), define the
normalized population weights \(w_q=q/(\mathbbm 1^\top q)\). The ratio
\[
\frac{\sum_{t=2}^{T_0}|C_t(w_q)|}
     {\sum_k w_{q,k}\sum_{t=2}^{T_0}|C_t(e_k)|}
\]
is \(1\) in the sign-aligned design and \(0.0511\) in the near-exact design.
Appendix~\ref{app:simulations} gives the construction.

Table~\ref{tab:mc_population_geometry} reports the deterministic population
exercise before introducing sampling uncertainty. Under sign alignment, the
projections based on donor vertices and the full simplex coincide. Under near-exact fit, the
vertex projection is unchanged because every donor's absolute pre-treatment
discrepancy is held fixed, whereas the projection over the full simplex contracts from
\([-3.365,-1.070]\) to \([-2.130,-1.807]\) percentage points. Its width falls
from \(2.295\) to \(0.324\) percentage points, a reduction of \(85.9\) percent.
Comparing the two weighting domains within each design shows how much interior
weights add when pre-treatment discrepancies cancel.

\begin{table}[!htbp]
\centering
\caption{Population Identification: Donor Vertices and Full Simplex}
\label{tab:mc_population_geometry}
\begin{threeparttable}
\footnotesize
\setlength{\tabcolsep}{7pt}
\begin{tabular}{llccc}
\toprule
Geometry
& Comparison domain
& Lower endpoint
& Upper endpoint
& Width \\
\midrule
Sign-aligned & Donor vertices & \(-3.365\) & \(-1.070\) & \(2.295\) \\
Sign-aligned & Full simplex   & \(-3.365\) & \(-1.070\) & \(2.295\) \\
Near-exact   & Donor vertices & \(-3.365\) & \(-1.070\) & \(2.295\) \\
Near-exact   & Full simplex   & \(-2.130\) & \(-1.807\) & \(0.324\) \\
\bottomrule
\end{tabular}
\begin{tablenotes}[flushleft]
\footnotesize
\item \textit{Notes:} Endpoints and widths are in percentage points. The
donor vertex domain maintains comparison validity only for weights that place
all mass on one donor; the full simplex domain maintains it for every convex
weight. All rows use \(L=2\), \(\rho=2\), outcome support, and the same
post-treatment population inputs.
\end{tablenotes}
\end{threeparttable}
\end{table}

The sampling experiment uses the full simplex model in both geometries. It
follows a finite-dimensional Gaussian experiment calibrated to the
household-cluster variation in the application. The covariance is estimated
from \(399\) cluster multiplier draws. In a cell with precision multiplier
\(m\in\{0.5,1,2\}\), the covariance is divided by \(m\), so larger values of
\(m\) correspond to greater precision. Every cell uses \(500\) Monte Carlo
replications and \(299\) bootstrap draws. Appendix~\ref{app:simulations}
describes the reconstruction of the candidate systems, bootstrap calibration,
and use of common random numbers across cells.

For each cell, I report false exclusion at the lower and upper population
endpoints and the median total width of the inverted acceptance set. Power is
evaluated at candidates one and two percentage points below the lower endpoint
and above the upper endpoint. Fixed absolute distances are used because the
two population projections have very different widths; normalizing the
alternatives by projection width would test economically different deviations.

\subsection{Results}

Table~\ref{tab:mc_main_results} reports the performance in finite samples of
the candidate tests and their inversion for the two population geometries.

\begin{table}[!htbp]
\centering
\caption{Monte Carlo Performance of Candidate Tests}
\label{tab:mc_main_results}
\begin{threeparttable}
\scriptsize
\setlength{\tabcolsep}{2.3pt}
\begin{tabular*}{\textwidth}{@{\extracolsep{\fill}}llccccccc@{}}
\toprule
Geometry
& \(m\)
& \multicolumn{2}{c}{False exclusion}
& \multicolumn{2}{c}{Power: 1 p.p.}
& \multicolumn{2}{c}{Power: 2 p.p.}
& Median width \\
\cmidrule(lr){3-4}\cmidrule(lr){5-6}\cmidrule(lr){7-8}
& & Lower & Upper & Below & Above & Below & Above & (p.p.) \\
\midrule
Sign-aligned & 0.5 & 0.018 & 0.004 & 0.088 & 0.042 & 0.310 & 0.276 & 7.410 \\
Sign-aligned & 1   & 0.024 & 0.006 & 0.202 & 0.100 & 0.632 & 0.548 & 5.851 \\
Sign-aligned & 2   & 0.026 & 0.006 & 0.366 & 0.300 & 0.934 & 0.890 & 4.786 \\
Near-exact   & 0.5 & 0.014 & 0.008 & 0.056 & 0.060 & 0.252 & 0.340 & 5.469 \\
Near-exact   & 1   & 0.014 & 0.006 & 0.120 & 0.164 & 0.520 & 0.656 & 3.922 \\
Near-exact   & 2   & 0.016 & 0.008 & 0.254 & 0.390 & 0.896 & 0.948 & 2.815 \\
\bottomrule
\end{tabular*}
\begin{tablenotes}[flushleft]
\scriptsize
\item \textit{Notes:} The abbreviation p.p. denotes percentage points. Each row uses
\(500\) Monte Carlo replications and
\(299\) bootstrap draws. Lower and Upper are rejection frequencies at the
compatible population endpoints. Below and Above are rejection frequencies at
fixed distances from the corresponding endpoint. Median width is the median
total width of the inversion set, in percentage points. Endpoint tests use
the population widths under the full simplex: \(2.295\) percentage points under sign
alignment and \(0.324\) under near-exact fit. The largest Monte Carlo standard
error in the table is
\(0.022\).
\end{tablenotes}
\end{threeparttable}
\end{table}

Estimated rates of false exclusion range from \(0.004\) to \(0.026\) across the
twelve endpoint experiments. With \(500\) replications per cell, these
estimates provide no evidence of excessive rejection relative to the nominal
\(0.05\) level.

Power rises with both absolute separation and sampling precision. At a
distance of two percentage points, rejection rises from \(0.252\)--\(0.340\) to
\(0.896\)--\(0.948\) in the near-exact design as \(m\) increases from \(0.5\)
to \(2\). The corresponding rise under sign alignment is from
\(0.276\)--\(0.310\) to \(0.890\)--\(0.934\). Alternatives one percentage point
away remain materially harder to distinguish, particularly at the
lower precision levels.

The near-exact geometry also yields narrower inverted sets at every precision
level. Relative to sign alignment, the median width is lower by \(26\), \(33\),
and \(41\) percent at \(m=0.5,1,2\). The population projection falls by \(86\)
percent, so sampling uncertainty attenuates the additional identifying content
of convex comparisons whose discrepancies nearly cancel.

\section{Conclusion}

This paper asks what panel comparisons reveal about treatment effects when
interference to comparison units are unknown. Each comparison relates the
treatment effect to a weighted average of spillovers rather than identifying it
separately. Requiring the same validity rule scaled by fit for every convex donor
weight restricts these relative effects, while restrictions tailored to the
application determine which treatment effect values remain possible. Their
intersection yields the sharp identified set.

Under the finite representation, checking a candidate reduces exactly to
solving a finite linear system, even though validity is required for infinitely
many donor weights. Farkas' alternative provides a certificate when the system
has no solution. Building on the normalized solvability test and bootstrap
calibration of \citet{goff2025inference}, I establish uniform candidatewise
validity for this structured family under model-tailored feasibility and
variance conditions. Inverting the tests controls false exclusion of each
compatible value in large samples. The procedure remains defined when the
identified set is a half-line or the real line, with reported inversion
limited to the compact domain \(\mathcal T\).

The LAWA application combines outcome support with a gross spillover budget
scaled by population. Under the main calibration \(\rho=2\), as well as the
tighter and looser calibrations \(\rho\in\{1,4\}\), the data and maintained
restrictions are consistent with both no effect and the original study's
estimate of a 1.50 percentage point decline. Factor placebos benchmark the
reported values of \(L\) against observed variation without selecting \(L\) or
estimating its latent threshold. The maintained comparison validity envelopes,
gross spillover budgets, and sampling uncertainty therefore do not determine
the sign of the treatment effect.

The simulations isolate when restrictions over all weights matter. Interior
weights do not tighten population identification when donor discrepancies are
sign aligned, but they can sharply narrow the set when convex aggregation
nearly eliminates pre-treatment mismatch. That gain depends on requiring
validity for every donor weight and, in the baseline model, a zero floor at
exact population fit. Sampling uncertainty
attenuates the population gain, while power rises and inversion width falls
with precision.

Panel comparisons can therefore remain informative without an exposure mapping
or advance knowledge of which comparison units are affected. Restrictions
tailored to the application determine what can be learned about the treatment
effect and can be combined with richer models of untreated outcomes.

\clearpage
\begingroup
\raggedright
\bibliographystyle{plainnat}
\begin{thebibliography}{83}
\expandafter\ifx\csname urlstyle\endcsname\relax
  \else
  \fi

\bibitem[Abadie and Gardeazabal(2003)]{abadie2003economic}
Alberto Abadie and Javier Gardeazabal.
\newblock The economic costs of conflict: A case study of the Basque Country.
\newblock \emph{American Economic Review}, 93\penalty0 (1):\penalty0 113--132,
  2003.

\bibitem[Abadie et~al.(2010)Abadie, Diamond, and
  Hainmueller]{abadie2010synthetic}
Alberto Abadie, Alexis Diamond, and Jens Hainmueller.
\newblock Synthetic control methods for comparative case studies: Estimating
  the effect of California's tobacco control program.
\newblock \emph{Journal of the American Statistical Association}, 105\penalty0
  (490):\penalty0 493--505, 2010.

\bibitem[Abadie et~al.(2015)Abadie, Diamond, and
  Hainmueller]{abadie2015comparative}
Alberto Abadie, Alexis Diamond, and Jens Hainmueller.
\newblock Comparative politics and the synthetic control method.
\newblock \emph{American Journal of Political Science}, 59\penalty0
  (2):\penalty0 495--510, 2015.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.1111/ajps.12116}.

\bibitem[Abadie(2021)]{abadie2021using}
Alberto Abadie.
\newblock Using synthetic controls: Feasibility, data requirements, and
  methodological aspects.
\newblock \emph{Journal of Economic Literature}, 59\penalty0 (2):\penalty0
  391--425, 2021.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.1257/jel.20191450}.

\bibitem[Agrawal(2015)]{agrawal2015taxgradient}
David~R. Agrawal.
\newblock The tax gradient: Spatial aspects of fiscal competition.
\newblock \emph{American Economic Journal: Economic Policy}, 7\penalty0
  (2):\penalty0 1--29, 2015.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.1257/pol.20120360}.

\bibitem[Alexander and Richards(2023)]{alexander2023hospitalclosures}
Diane~E. Alexander and Michael~R. Richards.
\newblock Economic consequences of hospital closures.
\newblock \emph{Journal of Public Economics}, 221:\penalty0 104821, 2023.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.1016/j.jpubeco.2023.104821}.

\bibitem[Amuedo-Dorantes and Lozano(2019)]{amuedodorantes2019interstate}
Catalina Amuedo-Dorantes and Fernando~A. Lozano.
\newblock Interstate mobility patterns of likely unauthorized immigrants:
  Evidence from Arizona.
\newblock \emph{Journal of Economics, Race, and Policy}, 2\penalty0
  (1--2):\penalty0 109--120, 2019.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.1007/s41996-018-0023-7}.

\bibitem[Andrews and Soares(2010)]{andrews2010inference}
Donald W.~K. Andrews and Gustavo Soares.
\newblock Inference for parameters defined by moment inequalities using
  generalized moment selection.
\newblock \emph{Econometrica}, 78\penalty0 (1):\penalty0 119--157, 2010.

\bibitem[Andrews et~al.(2023)Andrews, Roth, and Pakes]{andrews2023linear}
Isaiah Andrews, Jonathan Roth, and Ariel Pakes.
\newblock Inference for linear conditional moment inequalities.
\newblock \emph{Review of Economic Studies}, 90\penalty0 (6):\penalty0
  2763--2791, 2023.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.1093/restud/rdad004}.

\bibitem[Arkhangelsky et~al.(2021)Arkhangelsky, Athey, Hirshberg, Imbens, and
  Wager]{arkhangelsky2021synthetic}
Dmitry Arkhangelsky, Susan Athey, David A. Hirshberg, Guido W. Imbens, and Stefan
  Wager.
\newblock Synthetic difference-in-differences.
\newblock \emph{American Economic Review}, 111\penalty0 (12):\penalty0
  4088--4118, 2021.

\bibitem[Aronow and Samii(2017)]{aronow2017general}
Peter M. Aronow and Cyrus Samii.
\newblock Estimating average causal effects under general interference, with
  application to a social network experiment.
\newblock \emph{The Annals of Applied Statistics}, 11\penalty0 (4):\penalty0
  1912--1947, 2017.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.1214/16-AOAS1005}.

\bibitem[Bai(2009)]{bai2009panel}
Jushan Bai.
\newblock Panel data models with interactive fixed effects.
\newblock \emph{Econometrica}, 77\penalty0 (4):\penalty0 1229--1279, 2009.

\bibitem[Bai et~al.(2026)Bai, Ponomarev, Santos, Shaikh, Tabord-Meehan, and
  Torgovitsky]{bai2026linear}
Yuehao Bai, Kirill Ponomarev, Andres Santos, Azeem~M. Shaikh, Max
  Tabord-Meehan, and Alexander Torgovitsky.
\newblock Inference for linear systems with unknown coefficients, 2026.
\newblock URL \texttt{https://arxiv.org/abs/2604.24904}.

\bibitem[Banerjee et~al.(2013)Banerjee, Chandrasekhar, Duflo, and
  Jackson]{banerjee2013diffusion}
Abhijit Banerjee, Arun~G. Chandrasekhar, Esther Duflo, and Matthew~O. Jackson.
\newblock The diffusion of microfinance.
\newblock \emph{Science}, 341\penalty0 (6144):\penalty0 1236498, 2013.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.1126/science.1236498}.

\bibitem[Ben-Michael et~al.(2021)Ben-Michael, Feller, and
  Rothstein]{benmichael2021augmented}
Eli Ben-Michael, Avi Feller, and Jesse Rothstein.
\newblock The augmented synthetic control method.
\newblock \emph{Journal of the American Statistical Association}, 116\penalty0
  (536):\penalty0 1789--1803, 2021.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.1080/01621459.2021.1929245}.

\bibitem[Bohn et~al.(2014)Bohn, Lofstrom, and Raphael]{bohn2014lawa}
Sarah Bohn, Magnus Lofstrom, and Steven Raphael.
\newblock Did the 2007 Legal Arizona Workers Act reduce the state's
  unauthorized immigrant population?
\newblock \emph{Review of Economics and Statistics}, 96\penalty0 (2):\penalty0
  258--269, 2014.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.1162/REST_a_00429}.

\bibitem[Bugni et~al.(2017)Bugni, Canay, and Shi]{bugni2017inference}
Federico A. Bugni, Ivan A. Canay, and Xiaoxia Shi.
\newblock Inference for subvectors and other functions of partially identified
  parameters in moment inequality models.
\newblock \emph{Quantitative Economics}, 8\penalty0 (1):\penalty0 1--38, 2017.

\bibitem[Butts(2021)]{butts2021difference}
Kyle Butts.
\newblock Difference-in-differences estimation with spatial spillovers.
\newblock \emph{arXiv preprint arXiv:2105.03737}, 2021.

\bibitem[Cai et~al.(2015)Cai, de~Janvry, and Sadoulet]{cai2015social}
Jing Cai, Alain de~Janvry, and Elisabeth Sadoulet.
\newblock Social networks and the decision to insure.
\newblock \emph{American Economic Journal: Applied Economics}, 7\penalty0
  (2):\penalty0 81--108, 2015.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.1257/app.20130442}.

\bibitem[Callaway et~al.(2025)Callaway, Dyal, Sant'Anna, and
  Tsyawo]{callaway2025beyond}
Brantly Callaway, Derek Dyal, Pedro H.~C. Sant'Anna, and Emmanuel S. Tsyawo.
\newblock Beyond parallel trends: An identification-strategy-robust approach to
  causal inference with panel data.
\newblock \emph{arXiv preprint arXiv:2511.21977}, 2025.

\bibitem[Cao and Dowd(2019)]{cao2019estimation}
Jianfei Cao and Connor Dowd.
\newblock Estimation and inference for synthetic control methods with spillover
  effects.
\newblock \emph{arXiv preprint arXiv:1902.07343}, 2019.

\bibitem[Celli et~al.(2026)Celli, Cerqua, and
  Pellegrini]{celli2026identifying}
Viviana Celli, Augusto Cerqua, and Guido Pellegrini.
\newblock Identifying treatment and spillover effects with control-based and
  forecast-based counterfactuals.
\newblock \emph{arXiv preprint arXiv:2607.20156}, 2026.
\newblock URL \texttt{https://arxiv.org/abs/2607.20156}.

\bibitem[Chernozhukov et~al.(2007)Chernozhukov, Hong, and
  Tamer]{chernozhukov2007estimation}
Victor Chernozhukov, Han Hong, and Elie Tamer.
\newblock Estimation and confidence regions for parameter sets in econometric
  models.
\newblock \emph{Econometrica}, 75\penalty0 (5):\penalty0 1243--1284, 2007.

\bibitem[Chernozhukov et~al.(2013)Chernozhukov, Lee, and
  Rosen]{chernozhukov2013intersection}
Victor Chernozhukov, Sokbae Lee, and Adam~M. Rosen.
\newblock Intersection bounds: Estimation and inference.
\newblock \emph{Econometrica}, 81\penalty0 (2):\penalty0 667--737, 2013.

\bibitem[Cho and Russell(2024)]{cho2023simple}
JoonHwan Cho and Thomas~M. Russell.
\newblock Simple inference on functionals of set-identified parameters defined
  by linear moments.
\newblock \emph{Journal of Business \& Economic Statistics}, 42\penalty0
  (2):\penalty0 563--578, 2024.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.1080/07350015.2023.2203768}.

\bibitem[Cox et~al.(2025)Cox, Shi, and Shimizu]{cox2025testing}
Gregory Fletcher Cox, Xiaoxia Shi, and Yuya Shimizu.
\newblock Testing inequalities linear in nuisance parameters, 2025.
\newblock URL \texttt{https://arxiv.org/abs/2510.27633}.
\newblock arXiv:2510.27633.

\bibitem[Di~Stefano and Mellace(2024)]{distefano2024inclusive}
Roberta Di~Stefano and Giovanni Mellace.
\newblock The inclusive synthetic control method.
\newblock \emph{arXiv preprint arXiv:2403.17624}, 2024.

\bibitem[Ellis et~al.(2014)Ellis, Wright, Townley, and
  Copeland]{ellis2014migrationresponse}
Mark Ellis, Richard Wright, Matthew Townley, and Kristy Copeland.
\newblock The migration response to the Legal Arizona Workers Act.
\newblock \emph{Political Geography}, 42:\penalty0 46--56, 2014.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.1016/j.polgeo.2014.06.001}.

\bibitem[Fang and Santos(2019)]{fang2019inference}
Zheng Fang and Andres Santos.
\newblock Inference on directionally differentiable functions.
\newblock \emph{The Review of Economic Studies}, 86\penalty0 (1):\penalty0
  377--412, 2019.

\bibitem[Fang et~al.(2023)Fang, Santos, Shaikh, and Torgovitsky]{fang2023large}
Zheng Fang, Andres Santos, Azeem~M. Shaikh, and Alexander Torgovitsky.
\newblock Inference for large-scale linear systems with known coefficients.
\newblock \emph{Econometrica}, 91\penalty0 (1):\penalty0 299--327, 2023.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.3982/ECTA18979}.

\bibitem[Ferguson and Ross(2021)]{ferguson2021assessing}
Billy Ferguson and Brad Ross.
\newblock Assessing the sensitivity of synthetic control treatment effect
  estimates to misspecification error.
\newblock \emph{arXiv preprint arXiv:2012.15367}, 2021.
\newblock URL \texttt{https://arxiv.org/abs/2012.15367}.

\bibitem[Ferman and Pinto(2021)]{ferman2021synthetic}
Bruno Ferman and Cristine Pinto.
\newblock Synthetic controls with imperfect pretreatment fit.
\newblock \emph{Quantitative Economics}, 12\penalty0 (4):\penalty0 1197--1221,
  2021.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.3982/QE1596}.

\bibitem[Fern{\'a}ndez-Morales et~al.(2026)Fern{\'a}ndez-Morales, Oganisian,
  and Lee]{fernandezmorales2026bayesian}
Esteban Fern{\'a}ndez-Morales, Arman Oganisian, and Youjin Lee.
\newblock Bayesian shrinkage priors for penalized synthetic control estimators
  in the presence of spillovers.
\newblock \emph{Biometrics}, 82\penalty0 (2):\penalty0 ujag054, 2026.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.1093/biomtc/ujag054}.

\bibitem[Fischer et~al.(2024)Fischer, Royer, and White]{fischer2024obstetric}
Stefanie~J. Fischer, Heather Royer, and Corey~D. White.
\newblock Health care centralization: The health impacts of obstetric unit
  closures in the United States.
\newblock \emph{American Economic Journal: Applied Economics}, 16\penalty0
  (3):\penalty0 113--141, 2024.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.1257/app.20220341}.

\bibitem[Flood et~al.(2025)Flood, King, Rodgers, Ruggles, Warren, Backman,
  Breton, Cooper, Rivera~Drew, Richards, Van~Riper, and
  Williams]{flood2025ipumscps}
Sarah Flood, Miriam King, Renae Rodgers, Steven Ruggles, J.~Robert Warren,
Daniel Backman, Etienne Breton, Grace Cooper, Julia~A. Rivera Drew, Stephanie
Richards, David Van~Riper, and Kari~C.W. Williams.
\newblock \emph{IPUMS CPS: Version 13.0} [dataset].
\newblock IPUMS, Minneapolis, MN, 2025.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.18128/D030.V13.0}.

\bibitem[Forastiere et~al.(2021)Forastiere, Airoldi, and
  Mealli]{forastiere2021identification}
Laura Forastiere, Edoardo M. Airoldi, and Fabrizia Mealli.
\newblock Identification and estimation of treatment and interference effects
  in observational studies on networks.
\newblock \emph{Journal of the American Statistical Association}, 116\penalty0
  (534):\penalty0 901--918, 2021.

\bibitem[Gafarov(2025)]{gafarov2025simple}
Bulat Gafarov.
\newblock Simple subvector inference on sharp identified set in affine models.
\newblock \emph{Journal of Econometrics}, 249\penalty0 (PB):\penalty0 105952,
  2025.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.1016/j.jeconom.2025.105952}.

\bibitem[Gobillon and Magnac(2016)]{gobillon2016regional}
Laurent Gobillon and Thierry Magnac.
\newblock Regional policy evaluation: Interactive fixed effects and synthetic
  controls.
\newblock \emph{The Review of Economics and Statistics}, 98\penalty0
  (3):\penalty0 535--551, 2016.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.1162/REST_a_00537}.

\bibitem[Goff and Mbakop(2026)]{goff2025inference}
Leonard Goff and Eric Mbakop.
\newblock Testing the solvability of systems of linear inequalities, 2026.
\newblock URL \texttt{https://arxiv.org/abs/2506.06776}.
\newblock arXiv v3, May 8, 2026; previously circulated as ``Inference on the
  Value of a Linear Program''.

\bibitem[Greenstone et~al.(2010)Greenstone, Hornbeck, and
  Moretti]{greenstone2010agglomeration}
Michael Greenstone, Richard Hornbeck, and Enrico Moretti.
\newblock Identifying agglomeration spillovers: Evidence from winners and
  losers of large plant openings.
\newblock \emph{Journal of Political Economy}, 118\penalty0 (3):\penalty0
  536--598, 2010.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.1086/653714}.

\bibitem[Grossi et~al.(2025)Grossi, Mariani, Mattei, Lattarulo, and
  {\"O}ner]{grossi2025direct}
Giulio Grossi, Marco Mariani, Alessandra Mattei, Patrizia Lattarulo, and
  {\"O}zge {\"O}ner.
\newblock Direct and spillover effects of a new tramway line on the commercial
  vitality of peripheral streets: a synthetic-control approach.
\newblock \emph{Journal of the Royal Statistical Society Series A: Statistics
  in Society}, 188\penalty0 (1):\penalty0 223--240, 2025.

\bibitem[Hoffman(1952)]{hoffman1952approximate}
Alan~J. Hoffman.
\newblock On approximate solutions of systems of linear inequalities.
\newblock \emph{Journal of Research of the National Bureau of Standards},
  49\penalty0 (4):\penalty0 263--265, 1952.

\bibitem[Hsu and Small(2013)]{hsu2013calibrating}
Jesse~Y. Hsu and Dylan~S. Small.
\newblock Calibrating sensitivity analyses to observed covariates in
  observational studies.
\newblock \emph{Biometrics}, 69\penalty0 (4):\penalty0 803--811, 2013.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.1111/biom.12101}.

\bibitem[Huangfu and Hall(2018)]{huangfu2018parallelizing}
Qi Huangfu and Julian A.~J. Hall.
\newblock Parallelizing the dual revised simplex method.
\newblock \emph{Mathematical Programming Computation}, 10\penalty0
  (1):\penalty0 119--142, 2018.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.1007/s12532-017-0130-5}.

\bibitem[Huber and Steinmayr(2021)]{huber2021framework}
Martin Huber and Andreas Steinmayr.
\newblock A framework for separating individual-level treatment effects from
  spillover effects.
\newblock \emph{Journal of Business \& Economic Statistics}, 39\penalty0
  (2):\penalty0 422--436, 2021.

\bibitem[Hudgens and Halloran(2008)]{hudgens2008toward}
Michael G. Hudgens and M. Elizabeth Halloran.
\newblock Toward causal inference with interference.
\newblock \emph{Journal of the American Statistical Association}, 103\penalty0
  (482):\penalty0 832--842, 2008.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.1198/016214508000000292}.

\bibitem[Hyndman and Fan(1996)]{hyndman1996sample}
Rob~J. Hyndman and Yanan Fan.
\newblock Sample quantiles in statistical packages.
\newblock \emph{The American Statistician}, 50\penalty0 (4):\penalty0
  361--365, 1996.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.1080/00031305.1996.10473566}.

\bibitem[Imbens and Viviano(2023)]{imbens2023identification}
Guido~W. Imbens and Davide Viviano.
\newblock Identification and inference for synthetic controls with confounding,
  2023.
\newblock URL \texttt{https://arxiv.org/abs/2312.00955}.

\bibitem[Kaido et~al.(2019)Kaido, Molinari, and Stoye]{kaido2019confidence}
Hiroaki Kaido, Francesca Molinari, and J{\"o}rg Stoye.
\newblock Confidence intervals for projections of partially identified
  parameters.
\newblock \emph{Econometrica}, 87\penalty0 (4):\penalty0 1397--1432, 2019.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.3982/ECTA14075}.

\bibitem[Kosorok(2008)]{kosorok2008introduction}
Michael~R. Kosorok.
\newblock \emph{Introduction to Empirical Processes and Semiparametric
  Inference}.
\newblock Springer, 2008.

\bibitem[L{\'e}pissier and Mildenberger(2021)]{lepissier2021climate}
Alice L{\'e}pissier and Matto Mildenberger.
\newblock Unilateral climate policies can substantially reduce national carbon
  pollution.
\newblock \emph{Climatic Change}, 166\penalty0 (3-4):\penalty0 31, 2021.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.1007/s10584-021-03111-2}.
\newblock URL
  \texttt{https://link.springer.com/article/10.1007/s10584-021-03111-2}.

\bibitem[Leung(2020)]{leung2020treatment}
Michael P. Leung.
\newblock Treatment and spillover effects under network interference.
\newblock \emph{Review of Economics and Statistics}, 102\penalty0 (2):\penalty0
  368--380, 2020.

\bibitem[Leung(2022)]{leung2022causal}
Michael P. Leung.
\newblock Causal inference under approximate neighborhood interference.
\newblock \emph{Econometrica}, 90\penalty0 (1):\penalty0 267--293, 2022.

\bibitem[Liu(2025)]{liu2025synthetic}
Yiqi Liu.
\newblock Synthetic parallel trends.
\newblock \emph{arXiv preprint arXiv:2511.05870}, 2025.

\bibitem[Manski(2013)]{manski2013identification}
Charles~F. Manski.
\newblock Identification of treatment response with social interactions.
\newblock \emph{The Econometrics Journal}, 16\penalty0 (1):\penalty0 S1--S23,
  2013.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.1111/j.1368-423X.2012.00368.x}.

\bibitem[Manski and Pepper(2000)]{manski2000monotone}
Charles~F. Manski and John~V. Pepper.
\newblock Monotone instrumental variables: With an application to the returns
  to schooling.
\newblock \emph{Econometrica}, 68\penalty0 (4):\penalty0 997--1010, 2000.
\newblock ISSN 00129682, 14680262.
\newblock URL \texttt{http://www.jstor.org/stable/2999533}.

\bibitem[Manski and Pepper(2018)]{manski2018right}
Charles F. Manski and John V. Pepper.
\newblock How do right-to-carry laws affect crime rates? Coping with ambiguity
  using bounded-variation assumptions.
\newblock \emph{Review of Economics and Statistics}, 100\penalty0 (2):\penalty0
  232--244, 2018.

\bibitem[Mealli and Viviens(2026)]{mealli2026difference}
Fabrizia Mealli and Javier Viviens.
\newblock Difference-in-differences in the presence of unknown interference.
\newblock \emph{arXiv preprint arXiv:2512.21176}, 2026.
\newblock URL \texttt{https://arxiv.org/abs/2512.21176}.

\bibitem[Melnychuk(2024)]{melnychuk2024synthetic}
Andrii Melnychuk.
\newblock Synthetic controls with spillover effects: A comparative study.
\newblock \emph{arXiv preprint arXiv:2405.01645}, 2024.

\bibitem[Miguel and Kremer(2004)]{miguel2004worms}
Edward Miguel and Michael Kremer.
\newblock Worms: Identifying impacts on education and health in the presence of
  treatment externalities.
\newblock \emph{Econometrica}, 72\penalty0 (1):\penalty0 159--217, 2004.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.1111/j.1468-0262.2004.00481.x}.

\bibitem[Orrenius and Zavodny(2016)]{orrenius2016everify}
Pia~M. Orrenius and Madeline Zavodny.
\newblock Do state work eligibility verification laws reduce unauthorized
  immigration?
\newblock \emph{IZA Journal of Migration}, 5\penalty0 (1):\penalty0 5, 2016.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.1186/s40176-016-0053-3}.

\bibitem[O'Riordan and Gilligan-Lee(2025)]{oriordan2025spillover}
Michael O'Riordan and Ciar{\'a}n~M. Gilligan-Lee.
\newblock Spillover detection for donor selection in synthetic control models.
\newblock \emph{Journal of Causal Inference}, 13\penalty0 (1):\penalty0
  20240036, 2025.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.1515/jci-2024-0036}.

\bibitem[Powell(2026)]{powell2026imperfect}
David Powell.
\newblock Imperfect synthetic controls.
\newblock \emph{Journal of Applied Econometrics}, 41\penalty0 (3):\penalty0
  253--264, 2026.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.1002/jae.70035}.

\bibitem[R Core Team(2024)]{rcoreteam2024r}
R Core Team.
\newblock \emph{R: A Language and Environment for Statistical Computing}.
\newblock R Foundation for Statistical Computing, Vienna, Austria, 2024.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.32614/R.manuals}.

\bibitem[Rambachan and Roth(2023)]{rambachan2023more}
Ashesh Rambachan and Jonathan Roth.
\newblock A more credible approach to parallel trends.
\newblock \emph{Review of Economic Studies}, 90\penalty0 (5):\penalty0
  2555--2591, 2023.

\bibitem[Romano and Shaikh(2010)]{romano2010inference}
Joseph~P. Romano and Azeem~M. Shaikh.
\newblock Inference for the identified set in partially identified econometric
  models.
\newblock \emph{Econometrica}, 78\penalty0 (1):\penalty0 169--211, 2010.

\bibitem[Romano et~al.(2014)Romano, Shaikh, and Wolf]{romano2014practical}
Joseph~P. Romano, Azeem~M. Shaikh, and Michael Wolf.
\newblock A practical two-step method for testing moment inequalities.
\newblock \emph{Econometrica}, 82\penalty0 (5):\penalty0 1979--2002, 2014.

\bibitem[Rosen(2008)]{rosen2008confidence}
Adam M. Rosen.
\newblock Confidence sets for partially identified parameters that satisfy a
  finite number of moment inequalities.
\newblock \emph{Journal of Econometrics}, 146\penalty0 (1):\penalty0 107--117,
  2008.

\bibitem[Sakaguchi and Tagawa(2026)]{sakaguchi2026bayesian}
Shosei Sakaguchi and Hayato Tagawa.
\newblock Identification and Bayesian inference for synthetic control methods
  with spillover effects.
\newblock \emph{The Econometrics Journal}, art. utag006, 2026.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.1093/ectj/utag006}.

\bibitem[S{\"a}vje et~al.(2021)S{\"a}vje, Aronow, and
  Hudgens]{savje2021average}
Fredrik S{\"a}vje, Peter~M. Aronow, and Michael~G. Hudgens.
\newblock Average treatment effects in the presence of unknown interference.
\newblock \emph{The Annals of Statistics}, 49\penalty0 (2):\penalty0 673--701,
  2021.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.1214/20-AOS1973}.

\bibitem[S{\"a}vje(2024)]{savje2024causal}
Fredrik S{\"a}vje.
\newblock Causal inference with misspecified exposure mappings: Separating
  definitions and assumptions.
\newblock \emph{Biometrika}, 111\penalty0 (1):\penalty0 1--15, 2024.

\bibitem[Schr{\"o}der et~al.(2026)Schr{\"o}der, Oprescu, Feuerriegel, and
  Kallus]{schroder2026causal}
Maresa Schr{\"o}der, Miruna Oprescu, Stefan Feuerriegel, and Nathan Kallus.
\newblock Causal inference on networks under misspecified exposure mappings: A
  partial identification framework, 2026.
\newblock URL \texttt{https://arxiv.org/abs/2602.03459}.

\bibitem[Sion(1958)]{sion1958general}
Maurice Sion.
\newblock On general minimax theorems.
\newblock \emph{Pacific Journal of Mathematics}, 8\penalty0 (1):\penalty0
  171--176, 1958.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.2140/pjm.1958.8.171}.

\bibitem[Sobel(2006)]{sobel2006randomized}
Michael~E. Sobel.
\newblock What do randomized studies of housing mobility demonstrate? Causal
  inference in the face of interference.
\newblock \emph{Journal of the American Statistical Association}, 101\penalty0
  (476):\penalty0 1398--1407, 2006.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.1198/016214506000000636}.

\bibitem[Strassen(1965)]{strassen1965existence}
Volker Strassen.
\newblock The existence of probability measures with given marginals.
\newblock \emph{The Annals of Mathematical Statistics}, 36\penalty0
  (2):\penalty0 423--439, 1965.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.1214/aoms/1177700153}.

\bibitem[U.S. Census Bureau(2012)]{uscensus2012intercensal}
U.S. Census Bureau, Population Division.
\newblock Intercensal estimates of the resident population by single year of
  age and sex for states and the United States: April 1, 2000 to July 1, 2010.
\newblock \emph{ST-EST00INT-AGESEX}, October 2012.
\newblock URL
  \texttt{https://www2.census.gov/programs-surveys/popest/datasets/2000-2010/intercensal/state/st-est00int-agesex.csv}.

\bibitem[U.S. Census Bureau and U.S. Bureau of Labor
  Statistics(2000)]{censusbls2000cpsdesign}
U.S. Census Bureau and U.S. Bureau of Labor Statistics.
\newblock \emph{Current Population Survey: Design and Methodology}.
\newblock Technical Paper 63, March 2000.
\newblock URL
  \texttt{https://www2.census.gov/programs-surveys/cps/methodology/tp63.pdf}.

\bibitem[van~der Vaart(1998)]{vandervaart1998asymptotic}
Aad~W. van~der Vaart.
\newblock \emph{Asymptotic Statistics}.
\newblock Cambridge University Press, 1998.

\bibitem[van~der Vaart and Wellner(1996)]{vandervaartwellner1996weak}
Aad~W. van~der Vaart and Jon~A. Wellner.
\newblock \emph{Weak Convergence and Empirical Processes: With Applications to
  Statistics}.
\newblock Springer, 1996.

\bibitem[V{\'a}zquez-Bare(2023)]{vazquezbare2023spillover}
Gonzalo V{\'a}zquez-Bare.
\newblock Identification and estimation of spillover effects in randomized
  experiments.
\newblock \emph{Journal of Econometrics}, 237\penalty0 (1):\penalty0 105237,
  2023.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.1016/j.jeconom.2021.10.014}.

\bibitem[Virtanen et~al.(2020)]{virtanen2020scipy}
Pauli Virtanen et~al.
\newblock {SciPy} 1.0: Fundamental algorithms for scientific computing in
  Python.
\newblock \emph{Nature Methods}, 17\penalty0 (3):\penalty0 261--272, 2020.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.1038/s41592-019-0686-2}.

\bibitem[Xu(2026)]{xu2023difference}
Ruonan Xu.
\newblock Difference-in-differences with interference.
\newblock \emph{Journal of Econometrics}, 257:\penalty0 106304, 2026.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.1016/j.jeconom.2026.106304}.

\bibitem[Xu(2017)]{xu2017generalized}
Yiqing Xu.
\newblock Generalized synthetic control method: Causal inference with
  interactive fixed effects models.
\newblock \emph{Political Analysis}, 25\penalty0 (1):\penalty0 57--76, 2017.
\newblock doi: \begingroup \urlstyle{rm}\Url{10.1017/pan.2016.2}.

\end{thebibliography}
\endgroup
\clearpage