The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
51,676 characters
\doublespacing
\setlength{\emergencystretch}{2em}
\begin{center}
{\Large \textbf{Partial Identification of Spatial Production Networks}\textsuperscript{*}\par}
\vspace{1.5em}
\begin{tabular}{>{\centering\arraybackslash}m{0.3\textwidth} >{\centering\arraybackslash}m{0.3\textwidth} >{\centering\arraybackslash}m{0.3\textwidth}}
\textbf{Shaowen Luo} & \textbf{Kwok Ping Tsang} & \textbf{Zichao Yang} \\
\end{tabular}
\vspace{0.75em}
{\today\par}
\end{center}
\footnotetext[1]{Luo: Department of Economics, Virginia Tech, [email removed]. Tsang: Department of Economics, Virginia Tech, [email removed]. Yang: Wenlan School of Business, Zhongnan University of Economics and Law, [email removed].}
\setcounter{footnote}{0}
\begingroup
\setstretch{1.35}
\begin{abstract}
Which regional exposure conclusions are identified when public data do not observe buyer-seller links across states? We study this question by treating the missing intermediate-input spatial kernel as an unknown coupling constrained by regional activity margins, support restrictions, and auxiliary shipment moments. For linear exposure statistics, the sharp identified set is computed by transportation linear programs. Applying the method to U.S. state-sector data, we find that shipment data are inconsistent with the spatial diffuseness implied by proportional regionalization in key goods sectors. However, they do not identify a unique regional production network or a precise ranking of state exposure to local shocks. Bilateral shipment restrictions tighten the bounds, but much of the remaining uncertainty comes from large service and mixed sectors that are weakly covered by goods-movement data. The results show which exposure conclusions are supported by public data and which are imposed by maintained regionalization assumptions.
\end{abstract}
\endgroup
\noindent \textbf{Keywords}: production networks, spatial general equilibrium, partial identification, transportation polytopes, local shocks
\noindent \textbf{JEL Codes}: C67, D57, E23, R12
\clearpage
\section{Introduction}
Regional production-network calculations require assumptions about who buys from whom across space. A national input-output table tells us how much each industry buys from each other industry. Regional activity data tell us where industries are located. They do not tell us which supplier states sell intermediate inputs to which buyer states. The missing matrix is a bilateral state-sector buyer-seller matrix.
This paper studies what can be learned about that matrix before imposing a regionalization rule. We focus on the intermediate-input spatial kernel, the joint distribution of supplier and buyer locations for purchases from each supplier industry. Public data restrict its origin and destination margins only after the researcher chooses proxy mappings from regional activity to intermediate supply and demand. They do not identify the cells of the joint distribution. A proportional allocation, gravity completion, or support rule is therefore a maintained restriction on an unidentified coupling, not an observed regional IO table.
We develop a partial-identification approach to this problem. The admissible set consists of all nonnegative spatial kernels that satisfy maintained proxy margins, support restrictions, and auxiliary shipment moments. For linear one-step exposure statistics, the sharp identified interval is computed by transportation linear programs. This lets us separate exposure claims that survive the admissible set from claims that are imposed by a particular completed regional network.
The distinction matters because the same missing links are harmless for some propagation questions and first order for others. A pure national industry shock has no within-industry geographic variation, so the spatial kernel cancels in first-round exposure and in the Leontief accounting multiplier. This does not imply that industry shocks have small aggregate effects. It means that they cannot identify the spatial kernel. Local shocks are different. For regional and region-sector shocks, the kernel determines where downstream exposure lands, even when aggregate exposure looks stable.
We make the point in a regional trade-production model. The national input-output table gives the industry input share. What it does not give is the spatial sourcing share that says which supplier locations sell to which buyer locations. In an unrestricted regionalization, the regional intermediate-spending coefficient is
\begin{equation}
\label{eq:main_object_intro}
W_{(r,i),(s,j)}=\omega_{ji}\pi^{ij}_{r|s}.
\end{equation}
The national IO table identifies \(\omega_{ji}\), but not \(\pi^{ij}_{r|s}\). The empirical application imposes a lower-dimensional supplier-sector kernel, so buyer industries share the same spatial sourcing pattern within a supplier sector. This reduces the dimensionality of the missing bilateral matrix, but it does not make the remaining supplier-sector coupling identified.
The empirical application uses U.S. states, sixteen sectors, national IO coefficients, QCEW wage-bill and employment proxies for state-sector activity, and the 2017 Commodity Flow Survey. The CFS moments show that the proportional conditional-independence completion is too spatially diffuse in shipment-covered sectors. The average within-state shipment share is 0.506{} in the CFS, compared with 0.028{} under conditional independence and 0.384{} under a structured gravity completion. Conditional independence also has a mean distance-bin total variation gap of 0.578, compared with 0.123{} for structured gravity.
These shipment moments restrict the admissible set, but they do not identify local-shock incidence. For a Gulf regional shock, conditional independence lies outside the final sharp exposure interval for the directly shocked state of Louisiana: its point exposure is 7.5e-05{}, below the final sharp lower endpoint of 0.0009{}. Thus the proportional completion is not simply one admissible point inside the final set. The same calculation also shows the limit of the available public data. Adding bilateral CFS cell bands lowers the median state exposure interval width by 11.9{} percent relative to home-share and distance-bin moments. The final exact-margin specification still determines only 0.6{} percent of sharp state-pair exposure rankings, and only 1.0{} percent of pairs involving the 20 states with the largest upper endpoints. No state is guaranteed to be in the top exposure decile, and all 51{} states remain possible top-decile states.
The residual interval width is not concentrated only in the shipment-covered goods sectors. Services and other pooled sectors account for 53.5{} percent of total state interval width, with FIRE and professional and business services alone accounting for 26.5{} and 19.4{} percent. CFS-covered goods and logistics account for 38.7{} percent. Additional goods-shipment information therefore narrows the set, but it cannot by itself resolve state-level incidence when large service-sector sourcing relationships remain weakly observed. When wage-bill and employment margins are treated as bands rather than exact margins, aggregate exposure is no longer pinned down by exact-margin aggregation. Its width is 0.0019{}, while the median state width remains 0.0036{}.
We make three contributions. First, we characterize sharp bounds for unknown regional production-network couplings. The endpoints are transportation linear programs, and the dual variables summarize which margins or moments bind a given bound. Second, we compare common regionalization restrictions using CFS shipment moments and public state-sector activity proxies. The comparison covers conditional independence, structured gravity completion, local support restrictions, bilateral CFS cell bands, and banded margins. Third, we apply the admissible set to local-shock exposure and show which state-incidence conclusions survive. The analysis does not estimate a full regional IO table, employment effect, output effect, or welfare effect. A full counterfactual still requires final-demand sourcing, behavioral parameters, and equilibrium closure. A causal event study also needs a design for the outcome response.
\subsection{Relation to the literature}
The closest substantive literature is macroeconomic work on production networks and spatial propagation. Input-output linkages shape the transmission of shocks across sectors and, in regional models, across locations \citep{Hulten1978,LongPlosser1983,Horvath1998,Horvath2000,Gabaix2011,FoersterEtAl2011,AcemogluEtAl2012,Atalay2017,CaliendoEtAl2018,BaqaeeFarhi2019}. A natural structural reference point is \citet{CaliendoEtAl2018}, who specify a quantitative spatial model to study regional and sectoral productivity shocks, equilibrium reallocation, aggregate effects, and welfare. \citet{AdaoArkolakisEsposito2020} provide a complementary reduced-form bridge between shift-share designs and spatial general equilibrium effects by estimating bilateral reduced-form elasticities across local labor markets. We ask a different question. We study what public data identify about one input into these calculations before imposing the trade, substitution, labor-market, final-demand, and market-clearing structure needed for a full counterfactual.
The analysis is also related to evidence on firm-level production networks. Firm-to-firm data show that actual buyer-seller links matter for shock transmission \citep{BarrotSauvagnat2016,BoehmEtAl2019,CarvalhoEtAl2021}. Those papers use more detailed link-level information than is available in public regional IO construction. We ask what remains identifiable when the researcher has national industry linkages, state-sector proxy margins, and shipment moments, but not buyer-seller intermediate-input links.
The paper also uses partial identification and data combination. The identified-set logic follows the partial-identification perspective of \citet{Manski2003}, \citet{Tamer2010}, and \citet{Molinari2020}. The empirical problem also resembles data-combination settings in which separate sources identify different margins or moments of an unobserved joint distribution \citep{CrossManski2002,RidderMoffitt2007}. We do not report confidence intervals for the identified set. If sampling inference over estimated bounds were added, the relevant econometric issues would be those studied by \citet{ImbensManski2004} and \citet{Stoye2009}.
Finally, the computation uses the same mathematical structure as discrete optimal transport with fixed marginals. We use that structure as an identification and accounting tool, not as a behavioral transport-cost model. The linear-program representation is standard in optimal transport \citep{Galichon2016}. The regional IO literature has long developed non-survey, partial-survey, location-quotient, and balancing methods for constructing regional and interregional tables \citep{MillerBlair2009,FleggEtAl1995,FleggWebber2000,BoeroEtAl2018}. We do not propose a new regionalization formula. We provide an identified-set analysis of the intermediate-input spatial kernel that such calculations often take as an input.
The remainder of the paper proceeds as follows. Section \ref{sec:general_framework} defines the spatial-kernel coupling problem, the maintained proxy margins, and the sharp linear-programming bounds. Section \ref{sec:shock_irf} explains which exposure and multiplier calculations depend on the kernel before structural closure. Section \ref{sec:ci} describes conditional independence, structured gravity completion, support restrictions, and CFS moment bands. Section \ref{sec:implementation} presents the U.S. state-sector evidence. Section \ref{sec:experiments} reports the identified-set exposure results. Section \ref{sec:empirical_applications} compares the Gulf regional shock with a national manufacturing shock. Section \ref{sec:conclusion} concludes. The appendix gives proofs, data details, nonlinear multiplier calculations, and robustness tables.
\section{Spatial Kernels and Sharp Bounds}
\label{sec:general_framework}
\label{sec:model}
\label{sec:object}
The empirical problem is a missing joint distribution. A national input-output
table gives supplier-industry spending shares \(\omega_{ji}\) and buyer
industry intermediate-input intensities \(\mu_j\). Their product is
\begin{equation}
\label{eq:Bji}
B_{ji}=\mu_j\omega_{ji}.
\end{equation}
The missing component is the spatial coupling for intermediate purchases of a
fixed supplier sector \(i\). Let \(K^i_{rs}\) be the share of supplier-\(i\)
intermediate purchases that pairs supplier state \(r\) with buyer state \(s\).
The destination-conditional sourcing share is
\begin{equation}
\label{eq:model_joint}
\pi^i_{r|s}=\frac{K^i_{rs}}{b^i_s},
\end{equation}
where \(b^i_s\) is the destination margin. The regional input coefficient is
\begin{equation}
\label{eq:model_A}
A^K_{(r,i),(s,j)}=B_{ji}\pi^i_{r|s}.
\end{equation}
Thus the national IO block identifies industry linkages, while \(K^i\)
allocates those linkages across space.
The baseline uses maintained proxy margins. The origin margin is the
state-sector activity share,
\begin{equation}
\label{eq:intermediate_origin_marginal}
a^i_r=\frac{x_{ri}}{\sum_\ell x_{\ell i}},
\end{equation}
with \(x_{ri}\) measured by the relevant state-sector activity proxy. The
destination margin measures where supplier-\(i\) inputs are used:
\begin{equation}
\label{eq:intermediate_destination_marginal}
D^I_{si}=\sum_j \mu_j\omega_{ji}x_{sj},
\qquad
b^i_s=\frac{D^I_{si}}{\sum_\ell D^I_{\ell i}}.
\end{equation}
These margins are not observed bilateral intermediate-input flows. They are
maintained mappings from public state-sector activity data and the national IO
block.
\label{subsec:two_layers}
This creates two layers of non-identification. First, state-sector activity
does not separately identify output sold to intermediate users, household
final demand, and residual final use. Second, even if the relevant origin and
destination totals were known, the bilateral matrix matching supplier states
to buyer states would remain unidentified. The baseline therefore studies the
narrower intermediate-input kernel. Final-demand sourcing and equilibrium
closure are left to Appendix~\ref{app:final_demand_accounting}.
\begin{proposition}[Non-identification from national IO and regional marginals]
\label{prop:nonidentification}
Fix the national IO matrix. In the unrestricted pair-specific case, suppose
that for an industry pair the researcher observes origin and destination
margins that have the same total mass. If there are at least two regions, then
the joint spatial coupling is not identified by the national IO matrix and
those margins. Under the supplier-sector restriction used below, even if the
researcher observes the corresponding supplier-sector margins, the joint
spatial coupling is still not identified. Unless the relevant margins are
degenerate, there are multiple couplings that match the same observed margins
but imply different regional networks.
\end{proposition}
A two-region example makes the non-identification concrete. Let the origin
and destination margins both be \((1/2,1/2)\). The two couplings
\[
\begin{pmatrix}
1/2 & 0\\
0 & 1/2
\end{pmatrix}
\qquad \text{and} \qquad
\begin{pmatrix}
0 & 1/2\\
1/2 & 0
\end{pmatrix}
\]
match the same margins. The first is entirely local sourcing, while the
second is entirely cross-region sourcing. National IO coefficients and
regional margins alone cannot distinguish them.
For a generic finite coupling problem, let \(K\) be a nonnegative matrix with
row margin \(a\) and column margin \(b\). Support restrictions set selected
entries to zero. Auxiliary moments have the form
\[
m_h(K)=\sum_{x,y}g_{h,xy}K_{xy}.
\]
With a support set \(\mathcal S\) and a moment set \(\mathcal M\), the
admissible set is
\begin{equation}
\label{eq:generic_admissible_set}
\mathcal A(a,b,\mathcal S,\mathcal M)
=
\{K\geq0:\sum_y K_{xy}=a_x,\ \sum_x K_{xy}=b_y,\
K_{xy}=0 \text{ outside }\mathcal S,\ m(K)\in\mathcal M\}.
\end{equation}
Conditional independence is one feasible point when \(K_{xy}=a_xb_y\).
Structured gravity is another point completion. Support restrictions and CFS
moment bands define subsets of the transport polytope.
\begin{theorem}[Sharp linear identified sets and LP dual]
\label{thm:generic_lp_dual}
Suppose \(\mathcal M=\{m:\underline m\leq m\leq \overline m\}\), the
admissible set in equation \eqref{eq:generic_admissible_set} is nonempty, and
\(T(K)=c'k\) is linear in the vectorized coupling \(k\). Then the sharp
identified set for \(T(K)\) is \([\underline T,\overline T]\), where
\begin{equation}
\label{eq:generic_primal_min}
\underline T
=
\min_{k\geq0} c'k
\quad
\text{s.t.}
\quad
P_X k=a,\ P_Y k=b,\ Gk\leq\overline m,\ -Gk\leq-\underline m.
\end{equation}
Variables are restricted to \(\mathcal S\), and \(\overline T\) is obtained by
reversing the objective. The dual of the lower-bound program is
\begin{equation}
\label{eq:generic_dual}
\max_{u,v,\lambda^+,\lambda^-}
a'u+b'v+\overline m'\lambda^+-\underline m'\lambda^-,
\end{equation}
subject to
\[
u_x+v_y+\sum_{h=1}^J(\lambda^+_h-\lambda^-_h)g_{h,xy}
\leq c_{xy}
\quad \text{for all }(x,y)\in\mathcal S,
\]
with \(\lambda^+\leq0\) and \(\lambda^-\leq0\). Strong duality holds under the
maintained feasibility and boundedness conditions.
\end{theorem}
\begin{proposition}[Nested information]
\label{prop:nested_information}
Let \(\mathcal M_1\subseteq\mathcal M_0\). If the corresponding admissible
sets are nonempty, then the identified intervals for any target \(T\) satisfy
\[
\underline T(\mathcal M_0)\leq \underline T(\mathcal M_1)
\leq \overline T(\mathcal M_1)\leq \overline T(\mathcal M_0).
\]
\end{proposition}
When auxiliary moments are estimated or reconciled from imperfect data, we use
moment bands:
\begin{equation}
\label{eq:moment_band_set}
\mathcal A_\alpha
=
\{K\in\mathcal A(a,b,\mathcal S,\mathbb R^J):
\hat m_h-c_{h,\alpha}\leq m_h(K)\leq \hat m_h+c_{h,\alpha}
\text{ for all }h\}.
\end{equation}
\begin{proposition}[Moment-band coverage]
\label{prop:moment_band_coverage}
Let \(K_0\) be the true coupling and suppose it satisfies the maintained
margins and support restriction. If
\[
\Pr\left(
\hat m_h-c_{h,\alpha}\leq m_h(K_0)\leq \hat m_h+c_{h,\alpha}
\text{ for all }h
\right)\geq 1-\alpha,
\]
then the interval obtained by minimizing and maximizing a linear functional
\(T(K)\) over \(\mathcal A_\alpha\) covers \(T(K_0)\) with probability at
least \(1-\alpha\).
\end{proposition}
For exposure, the target is linear. Let \(Q(E)=\sum_{s,j}q_{sj}E_{sj}\). For a
shock \(z\),
\begin{equation}
\label{eq:linear_exposure_functional}
Q(E^K(z))
=
\sum_i\sum_{r,s}\ell^i_{rs}(q,z)K^i_{rs},
\qquad
\ell^i_{rs}(q,z)
=
\frac{z_{ri}}{b^i_s}\sum_j q_{sj}B_{ji}.
\end{equation}
\begin{proposition}[Sharp bounds for linear exposure]
\label{prop:admissible_exposure}
If each sectoral admissible set is nonempty, compact, convex, and defined by
linear margins, support restrictions, or moment bands, then the admissible
values of \(Q(E^K(z))\) form the interval obtained by minimizing and
maximizing equation~\eqref{eq:linear_exposure_functional} over the product of
sectoral admissible sets. The endpoints are sharp and are computed by
sector-by-sector transportation linear programs.
\end{proposition}
The theorem and propositions are the main tools used below. They also define
the limits of the analysis. The bounds are sharp for linear exposure given the
maintained margins and moment bands. They do not identify behavioral
elasticities, final-demand substitution, factor adjustment, or welfare.
\section{Propagation Before Structural Closure}
\label{sec:shock_irf}
What does the missing kernel affect? Given an admissible intermediate-input
kernel \(K\), define the regional input matrix
\begin{equation}
\label{eq:macro_A}
A^K_{(r,i),(s,j)}
=
B_{ji}\pi^i_{r|s}.
\end{equation}
The first-round exposure of destination node \((s,j)\) to a shock vector \(z\)
is
\begin{equation}
\label{eq:general_exposure}
E^K_{sj}(z)
=
\sum_i B_{ji}\sum_r \pi^i_{r|s}z_{ri}.
\end{equation}
This one-step exposure is linear in \(K\). With a unit shock, a value of
0.001 means one-tenth of one percentage point in this accounting exposure
index before equilibrium responses. The corresponding Leontief accounting
multiplier is
\begin{equation}
\label{eq:macro_multiplier}
M^K
=
\left(I-\left(A^K\right)'\right)^{-1},
\end{equation}
whenever \(\rho(A^K)<1\). A Domar-style accounting exposure index is
\begin{equation}
\label{eq:domar_loss}
\mathcal L^K(z)
=
\sum_{s,j}\lambda_{sj}\left[M^Kz\right]_{sj}.
\end{equation}
These are accounting quantities. Employment, output, and welfare responses
require final demand, prices, factor adjustment, financing, and market
clearing.
For a pure industry shock, \(z_{ri}=z_i\) for every origin \(r\), the
first-round exposure of buyer node \((s,j)\) is
\begin{equation}
\label{eq:industry_shock_exposure}
E_{sj}^{I}(z)
=
\sum_i B_{ji}z_i,
\end{equation}
because the sourcing shares sum to one. The same cancellation holds for
Leontief accounting exposure because each term in the Leontief series maps a
vector that is constant across origins within an industry into another vector
with the same property. Pure industry shocks therefore cannot validate a
spatial regionalization rule.
For a pure regional shock, \(z_{ri}=z_r\), the kernel generally determines
where exposure lands. One aggregate incidence measure still cancels. If
exposure is first aggregated over destination regions using the
supplier-sector destination margins \(b^i_s\), then
\begin{equation}
\label{eq:regional_aggregate_industry}
\bar E_j^R(z)
=
\sum_i B_{ji}\sum_s b^i_s\sum_r \pi^i_{r|s}z_r
=
\sum_i B_{ji}\sum_r a^i_r z_r .
\end{equation}
\begin{proposition}[Invariance and kernel dependence]
\label{prop:shock_cancellation}
For first-round exposure in equation~\eqref{eq:general_exposure}, every
\(K\in\mathcal A(\mathcal M)\) gives the same buyer-industry exposure to a
pure industry shock. In the Leontief accounting system in
equation~\eqref{eq:macro_multiplier}, the same invariance holds for the full
multiplier response to pure industry shocks. For a pure regional shock, the
kernel is not needed for the destination-margin-weighted industrial incidence
in equation~\eqref{eq:regional_aggregate_industry}. The kernel is needed to
allocate exposure across destination regions whenever the shock has geographic
content.
\end{proposition}
The proposition gives the paper's boundary result. Industry shocks can answer
industry propagation questions, but they say little about regional incidence.
Regional and region-sector shocks require the spatial kernel for local
incidence.
Leontief exposure is harder to bound sharply because \(M^K\) is nonlinear in
\(K\). Exact global bounds are nonlinear optimization problems. We therefore
use one local outer-bound calculation only as an appendix result. Appendix
\ref{app:additional_evidence} gives the sensitivity formula and the
conservative remainder bound used in the nonlinear multiplier table. The main
empirical results below are sharp for linear one-step exposure, not for the
full nonlinear multiplier.
\section{Admissible Spatial-Kernel Restrictions}
\label{sec:ci}
Common regionalization rules are restrictions on the feasible coupling set.
Some select a point completion. Others define an admissible subset over which
linear exposure can be bounded. The question is whether the restriction is
admissible relative to maintained margins, support restrictions, and auxiliary
shipment moments.
The proportional completion fills in the missing joint coupling as
\begin{equation}
\label{eq:ci}
K^{i,CI}_{rs}=a^i_r b^i_s .
\end{equation}
In destination-conditional form, every buyer region sources supplier-\(i\)
inputs from the same origin distribution. Conditional independence is also
the maximum-entropy completion, maximizing
\begin{equation}
\label{eq:entropy}
-\sum_{r,s}K^i_{rs}\log K^i_{rs}.
\end{equation}
\begin{proposition}[Conditional independence as maximum entropy]
\label{prop:max_entropy}
Among all nonnegative couplings with origin marginal \(a^i\) and destination
marginal \(b^i\), the proportional matrix \(K^{i,CI}_{rs}=a^i_r b^i_s\) is the
unique maximum-entropy coupling when the marginals are positive. Equivalently,
it imposes zero mutual information between supplier and buyer locations
conditional on supplier industry.
\end{proposition}
This makes conditional independence a useful benchmark, not an identified
coupling. It rules out home bias, distance decay, corridor structure, and other
dependence between supplier and buyer locations after conditioning on supplier
industry.
\label{subsec:dgp_ci_bias}
We also use a structured gravity completion as a point comparison. For each
supplier industry \(i\), the unbalanced kernel is
\begin{equation}
\label{eq:dgp_kernel}
\widetilde K^i_{rs}(\eta,\tau)
=
a^i_r b^i_s
\exp\left\{
\eta \mathbf{1}\{r=s\}
-\tau \log(1+d_{rs})
\right\}.
\end{equation}
It is rebalanced by iterative proportional fitting to match the maintained
margins.
\label{sec:workflow}
The empirical analysis uses different restrictions by sector. For
shipment-covered sectors, we estimate a gravity-style point completion,
\begin{equation}
\label{eq:gravity}
K^i_{rs}
\propto
\exp\left\{
\alpha^i_r+\delta^i_s+\eta_i\mathbf{1}\{r=s\}
-\tau_i \log(1+d_{rs})
\right\},
\end{equation}
where origin and destination fixed effects absorb the marginals. For mixed
sectors, we pool spatial parameters. For local sectors, we impose support
restrictions rather than treating shipment-like observations as sectoral
intermediate-input flows. The baseline local support includes same-state
pairs, adjacent-state pairs, and the minimum additional nearest-state radius
needed for feasibility.
Let \(\mathcal S_i\) be a support set and let
\(\mathcal K_i(a^i,b^i,\mathcal S_i)\) be the nonnegative matrices that match
the margins and place zero mass outside \(\mathcal S_i\). For any linear
exposure functional \(L(K^i)=\sum_{r,s}\ell_{rs}K^i_{rs}\), sharp bounds are
\begin{equation}
\label{eq:transport_bounds}
\underline L_i
=
\min_{K^i\in\mathcal K_i(a^i,b^i,\mathcal S_i)}
L(K^i),
\qquad
\overline L_i
=
\max_{K^i\in\mathcal K_i(a^i,b^i,\mathcal S_i)}
L(K^i).
\end{equation}
These are sharp under the maintained support restriction because the feasible
set is exactly the transport polytope matching \(a^i\), \(b^i\), and
\(\mathcal S_i\).
All restrictions are fixed before propagation outcomes are inspected. This
differs from a standard regional IO construction because a completed
matrix can fit shipment geography and still reveal little about
local-shock incidence.
\section{U.S. State-Sector Evidence}
\label{sec:implementation}
We next compare common spatial-kernel restrictions with observed shipment
geography. The empirical application uses U.S. states and a sixteen-sector
aggregation. The national input-output block provides the industry shares
\(\omega_{ji}\) and the intermediate-input intensities \(\mu_j\). The baseline
origin margins, destination margins, and exposure weights use 2019 QCEW
wage-bill shares. The destination marginal is constructed from destination
activity, sector intermediate-input intensities, and the national IO matrix,
as in equation~\eqref{eq:intermediate_destination_marginal}. It is a proxy
for intermediate-demand geography, not regional final expenditure. Wage bills
are the baseline because they are a public, consistently available measure of
state-sector economic activity and they weight high-productivity state-sector
cells more than headcount alone. QCEW employment shares are the first
robustness margin, and model-output margins are reported only in
robustness. State-to-state shipment flows come from the
2017 Commodity Flow Survey. State distances are computed from Census state
centroids.
We therefore interpret the empirical inputs as maintained proxy measures.
The spatial kernel is the unknown intermediate-input coupling.
Conditional independence is the proportional benchmark. Structured gravity
completion is a low-dimensional point completion estimated from shipment
geography. One-step exposure is an accounting exposure measure before
employment responses, price responses, and welfare effects.
\begin{table}[htbp]
\centering
\caption{Sector Treatment in the Main Specification}
\label{tab:sector_classification}
\normalsize
\setlength{\tabcolsep}{4pt}
\begin{tabular}{@{}p{0.22\textwidth}p{0.38\textwidth}p{0.32\textwidth}@{}}
\toprule
Group & Sectors & Treatment \\
\midrule
Shipment-covered & Manufacturing, mining, transportation, wholesale & CFS moment and selected bilateral cell bands \\
Tradable pooled & Agriculture & Pooled tradable parameters \\
Mixed pooled & Utilities, information, finance, insurance, real estate, professional and business services & Pooled mixed-sector parameters \\
Local support-restricted & Construction, retail, education, health care, leisure, other services, government & Support bounds for local sourcing \\
\bottomrule
\end{tabular}
\begin{flushleft}
\footnotesize Notes: The classification is fixed before propagation outcomes are interpreted. Retail is treated as local in the main specification because retail output is conceptually closer to local service provision than to an interregional supplier sector. Appendix Table~\ref{tab:retail_robustness} reports a shipment-informed retail variant.
\end{flushleft}
\end{table}
\subsection{Why shipment data do not observe the coupling}
CFS moments are auxiliary shipment moments, not the full intermediate-input
coupling. The distinction matters for interpretation. First, CFS
shipments include final goods as well as goods that may become intermediate
inputs. Second, CFS commodity classifications do not map one-for-one into
the production industries in the IO table. Third, wholesale shipments and
re-shipments can break the link between the producing origin and the
intermediate-input seller relevant for a buyer. Fourth, services are missing
or weakly covered, which matters because many local and business-service
sectors are large in the IO block. Finally, CFS observes goods movement. It
does not observe the buyer-seller use of intermediate inputs by destination
industry. For this reason, the CFS restrictions below restrict shipment
geography within selected sectors, but they do not convert the spatial kernel
into an observed matrix.
We use point completions only as comparisons. Conditional
independence is the proportional benchmark. Structured gravity, with pooled
variants where data are thin, is the shipment-informed comparison. For local
sectors, the relevant calculation is not a point completion but the support-restricted
admissible set. Lower and upper support-bound kernels summarize feasible
ranges. All comparisons hold fixed the same national IO block and the same
intermediate-demand marginals. They differ only in the intermediate-input
spatial coupling.
Figure~\ref{fig:home_bias_wedge} shows the main empirical fact. In the
four strict shipment-covered sectors, observed state-to-state CFS flows display
large within-state shares. The average observed CFS home share is
0.506. Conditional independence implies 0.028. The
structured gravity completion implies 0.384. Thus conditional
independence does not merely smooth bilateral flows. It removes most of the
same-state mass observed in shipment data. Appendix
Table~\ref{tab:home_bias_wedge} reports the corresponding home-share and
held-out RMSE values.
\begin{figure}[htbp]
\centering
\includegraphics[width=0.82\textwidth]{figures/fig_home_bias_wedge.png}
\caption{Home Bias in Shipment-Covered Kernel Restrictions}
\label{fig:home_bias_wedge}
\begin{flushleft}
\footnotesize Notes: The figure reports within-state shipment or kernel shares for the four strict shipment-covered sectors. CFS is the observed within-state shipment share. CI is the proportional conditional-independence completion. Gravity is the structured spatial completion. The interpretation of CFS moments is described in Section~\ref{sec:implementation}.
\end{flushleft}
\end{figure}
Table~\ref{tab:admissibility_frontier} reports admissibility frontiers using
the same evidence. For each restriction, we ask how much the CFS moment restrictions must be relaxed before that restriction becomes admissible. A larger tolerance means the restriction is farther from the shipment evidence. Formally, for a restriction \(R\), let
\[
\tau_R=\inf\{\tau:m(K_R)\in\mathcal M(\tau)\}
\]
be the smallest tolerance under which the restriction is admissible for the maintained moment. For the shipment-covered sectors, the moments are the CFS home share, the CFS distance-bin distribution, and held-out flow fit. Conditional independence requires an average home-share tolerance of
0.478, while sector gravity requires 0.122.
The corresponding distance-bin total variation gaps are
0.578{} and 0.123. This should not be read as a formal statistical rejection of conditional independence, because it
does not use CFS sampling variances. It shows that conditional independence is admissible only under a much looser moment set than sector gravity.
\begin{table}[htbp]
\centering
\caption{Admissibility Frontier for Shipment-Covered Kernel Restrictions}
\label{tab:admissibility_frontier}
\normalsize
\begin{tabular}{lcccc}
\toprule
Restriction & Home gap & Distance TV & Avg. dist. gap & RMSE/width \\
\midrule
Conditional independence & 0.451 & 0.578 & 784.0 & 0.0120 \\
Sector gravity & 0.066 & 0.123 & 126.0 & 0.0061 \\
Support restriction & -- & -- & -- & 0.714 \\
\bottomrule
\end{tabular}
\begin{flushleft}
\footnotesize Notes: The table reports admissibility-frontier comparisons for the four strict shipment-covered sectors. Home gap is the mean absolute gap between
the completion-implied home share and the observed CFS home share. Distance TV is the mean total variation distance between the completion-implied and observed CFS distance-bin distributions. Avg. dist. gap is the mean absolute
gap in average shipment distance, in miles. RMSE is held-out normalized flow RMSE.
\end{flushleft}
\end{table}
Local support restrictions are different from the shipment-covered
point-completion comparisons in Table~\ref{tab:admissibility_frontier}. They define feasible sets rather than CFS flow-fit statistics. Appendix Figure~\ref{fig:local_bounds} reports the corresponding local-sector home-share intervals. The mean interval width is 0.714.
The gap is economically large. The proportional completion does more than smooth flows at the margin. It almost eliminates home bias in the sectors where shipment data show home bias most clearly. In manufacturing, observed CFS flows imply a within-state share of 0.371{}, while the CI completion implies 0.026{}. In wholesale, the corresponding numbers are 0.606{} and
0.028{}. Gravity is closer on both home shares and held-out flow fit, but it remains a maintained low-dimensional restriction rather than an identified intermediate-input spatial kernel.
This distinction motivates the sector treatment in
Table~\ref{tab:sector_classification}: sector-specific CFS restrictions for
shipment-covered sectors, pooled parameters where shipment evidence is thin,
and support-restricted bounds for local sectors.
\section{Identified-Set Exposure Results}
\label{sec:experiments}
We now report sharp bounds for one-step exposure over the admissible set of
intermediate-input kernels. The bounds are not ranges across selected point
completions. They are the minimum and maximum exposure values attainable by
any kernel satisfying the maintained proxy margins, support restrictions, and
CFS moment bands. The calculations hold fixed the national IO table,
intermediate-demand proxy margins, sectoral intermediate-input intensities,
and state-sector wage-bill exposure weights. Only the intermediate-input
spatial coupling changes.
The results show that the public data determine some incidence
comparisons but leave many state rankings unresolved. The calibrated
economy has 51 states and 16 sectors. For each coupling \(K\), we construct
the regional input matrix in equation~\eqref{eq:macro_A} using
\(B_{ji}=\mu_j\omega_{ji}\), with
\(\mu_j=\text{intermediate}_j/\text{output}_j\). The maximum spectral radius
across all calibrated and comparator matrices is 0.4483{}, so the accounting
inverse exists in every reported case.
\subsection{Sharp one-step bounds for the Gulf regional shock}
We first study a Gulf regional shock that hits
Louisiana and Mississippi in all supplier sectors. This shock has geographic
content, so the spatial kernel matters for where exposure lands. At the same
time, the exact-margin aggregate target has a cancellation property. Because
the aggregate weights and destination margins use the same wage-bill proxy,
the unknown bilateral kernel collapses to maintained origin margins after
destination aggregation. The zero aggregate width in the exact-margin
baseline should therefore not be read as evidence that bilateral sourcing is
precisely identified. We therefore focus on state-level incidence.
Table~\ref{tab:information_content} reports how the Gulf exposure set changes
as information is added. The first four rows add exact proxy margins,
local-sector support restrictions, CFS home-share bands, and CFS distance-bin
bands. The fifth row adds selected bilateral CFS cell bands for shipment-
covered sectors. The final two rows replace exact wage-bill margins with
bands whose lower and upper endpoints are the QCEW wage-bill and employment
shares. The mean home-share feasibility tolerance is 0.035, and
the maximum is 0.141. Appendix Table~\ref{tab:moment_tolerances}
reports the sector-specific tolerances.
\begin{table}[htbp]
\centering
\caption{Information Content of Spatial Restrictions}
\label{tab:information_content}
\normalsize
\resizebox{\textwidth}{!}{\begin{tabular}{lccccc}
\toprule
Admissible set & Agg. interval & Agg. width & Median width & P90 width & Possible top \\
\midrule
Margins only & [0.0066, 0.0066] & 0 & 0.00474 & 0.00658 & 51 \\
Local support & [0.0066, 0.0066] & 0 & 0.00312 & 0.00421 & 51 \\
CFS home band & [0.0066, 0.0066] & 0 & 0.00309 & 0.00419 & 51 \\
CFS distance-bin band & [0.0066, 0.0066] & 0 & 0.00309 & 0.00419 & 51 \\
CFS bilateral cell bands & [0.0066, 0.0066] & 0 & 0.00272 & 0.00340 & 51 \\
\midrule
Banded margins & [0.0065, 0.0085] & 0.00197 & 0.00388 & 0.00553 & 51 \\
Banded margins + CFS bilateral & [0.0066, 0.0085] & 0.00189 & 0.00364 & 0.00470 & 51 \\
\bottomrule
\end{tabular}}
\begin{flushleft}
\footnotesize Notes: The table reports sharp one-step exposure bounds for the
Gulf regional shock. Agg. interval and Agg. width refer to the
wage-bill-weighted aggregate one-step exposure interval. Median width and P90
width summarize state-level exposure intervals. Possible top is the number of
states that can be in the top exposure decile for some admissible kernel,
using conservative interval classification. The banded-margin rows are a
separate admissible-set layer, not a nested refinement of the exact wage-bill
rows. No row has a robust top-decile state.
\end{flushleft}
\end{table}
The top-decile classification is deliberately conservative, so we also solve
sharp pairwise bounds. For each unordered state pair, we bound \(E_s-E_{s'}\).
One state is classified as dominating the other only when the sharp lower
bound is positive or the sharp upper bound is negative. Under the final
exact-margin bilateral-CFS set, only eight of 1,275 state pairs are
determined, or 0.6{} percent. Only one state has any
robust dominance relation. The determined share among pairs involving the 20
states with the largest upper endpoints is
1.0{} percent. This pattern is visible in Figure
\ref{fig:gulf_state_intervals}: except for Louisiana, most lower endpoints
among high-upper-bound states remain close to zero.
Figure~\ref{fig:gulf_state_intervals} shows why many rankings remain
unresolved. Most lower endpoints among the high-upper-bound states are close
to zero, and many intervals overlap even after CFS restrictions. The figure
also shows that conditional independence is outside the final sharp interval
for one directly shocked state. For Louisiana, conditional independence gives
exposure 7.5e-05{}, below the sharp lower endpoint
0.0009{}. Structured gravity remains inside all displayed
intervals.
\begin{figure}[htbp]
\centering
\includegraphics[width=0.86\textwidth]{figures/fig_gulf_state_exposure_intervals.png}
\caption{Top Gulf State Exposure Intervals}
\label{fig:gulf_state_intervals}
\begin{flushleft}
\footnotesize Notes: Horizontal lines report the sharp one-step Gulf regional-
shock exposure interval for the 20 states with the largest upper endpoint.
States are sorted by the upper endpoint. Crosses mark conditional
independence, and diamonds mark structured gravity. Asterisks mark Louisiana
and Mississippi. Louisiana has a strictly positive lower endpoint, and its
conditional-independence point lies below that lower bound. Most other lower
endpoints are close to zero.
\end{flushleft}
\end{figure}
The sector decomposition in Table~\ref{tab:sector_width_decomposition}
explains where residual uncertainty comes from. The decomposition is exact:
the state exposure interval width is the sum of sector-level endpoint
differences because the transport problem separates by supplier sector.
FIRE, professional and business services, manufacturing, wholesale, and
transportation are the largest contributors to median state width. Grouping
the sectors shows that services and other pooled sectors account for
53.5{} percent of total state interval width. FIRE and
professional and business services alone account for 26.5{} and
19.4{} percent. CFS-covered goods and logistics account for
38.7{} percent. The residual uncertainty is
therefore not only a problem of missing goods-shipment cells. It also reflects
large service and mixed sectors whose intermediate-input geography is not
directly observed in CFS.
\begin{table}[htbp]
\centering
\caption{Sector Decomposition of Gulf Interval Width}
\label{tab:sector_width_decomposition}
\normalsize
\resizebox{\textwidth}{!}{\begin{tabular}{llccc}
\toprule
Supplier sector & Group & Median contribution & P90 contribution & Mean share (\%) \\
\midrule
FIRE & services and other pooled sectors & 0.00076 & 0.00076 & 26.5 \\
Prof. Business Services & services and other pooled sectors & 0.00059 & 0.00059 & 19.4 \\
Manufacturing & CFS-covered goods and logistics & 0.00057 & 0.00057 & 21.1 \\
Wholesale & CFS-covered goods and logistics & 0.00028 & 0.00039 & 10.5 \\
Transportation & CFS-covered goods and logistics & 0.00014 & 0.00034 & 6.9 \\
Information & services and other pooled sectors & 0.00010 & 0.00010 & 4.0 \\
Utilities & services and other pooled sectors & 0.00008 & 0.00015 & 3.5 \\
Agriculture & services and other pooled sectors & 0.00004 & 0.00006 & 1.6 \\
\bottomrule
\end{tabular}}
\begin{flushleft}
\footnotesize Notes: Contributions are sector-level differences between the
upper and lower exposure endpoint for each state, aggregated over states. The
table lists the largest sector contributors by median state contribution.
\end{flushleft}
\end{table}
The bilateral CFS cells lower the median state width by
11.9{} percent, so they add information, but
the gain is modest relative to the remaining interval width. Appendix
Table~\ref{tab:binding_bilateral_cfs_cells} reports which selected bilateral
cells bind in the Gulf exposure endpoint problems.
The banded-margin rows in Table~\ref{tab:information_content} evaluate the
role of exact wage-bill margins. In those rows, each
sector's total mass remains one, but origin and destination shares are allowed
to range between the QCEW wage-bill and employment shares. Under banded
margins with bilateral CFS restrictions, aggregate exposure width is
0.0019{} and median state width is
0.0036{}. These rows should be interpreted as a separate
admissible-set layer, not as a nested refinement of the exact wage-bill
baseline.
Appendix Table~\ref{tab:moment_band_bounds} reports the CFS moment-band
calculation. Appendix Table~\ref{tab:alternative_margin_identified_sets}
repeats the identified-set calculation with wage-bill, employment, mixed, and
model-output margins. The model-output row is a robustness specification rather than
the baseline. Across those margin constructions, all states remain possible
top-decile exposure states. Appendix Table~\ref{tab:mining_robustness} also
shows that excluding mining CFS moments leaves the median and p90 state
interval widths at 0.0031{} and
0.0042{}. The exact-margin aggregate cancellation is
therefore a property of the maintained aggregate target. Weak identification
of regional incidence is not.
Appendix Table~\ref{tab:dual_shadow_prices} reports the largest nonzero LP
moment shadows for state-level Gulf exposure bounds. The rows are not
whole-state exposure endpoints. They are sector-specific dual results inside
the state-level bounds. They show which CFS moment restrictions bind
particular sector-level endpoints, while the maximum primal-dual gap in the
full dual output is only 0{}.
Nonlinear Leontief multiplier exposure is harder to bound sharply. Appendix
Table~\ref{tab:nonlinear_multiplier_outer_bounds} reports the perturbation
outer-bound calculation. In the Gulf application, the first-order LP interval
is much narrower than the final outer interval because the conservative
perturbation radius, 2.947, is above one. The resulting
outer interval is valid and contains the reported point-completion multiplier
values, but it is too wide to support a sharp nonlinear identified-set claim
in this application. Without additional structure, nonlinear Leontief exposure
is much harder to sharply bound than linear one-step exposure. For this
reason, the empirical results in the main text are interpreted as sharp bounds
for one-step exposure, not for full nonlinear propagation.
\section{Local Shock Exposure Applications}
\label{sec:empirical_applications}
The preceding results imply a simple distinction. A shock can have economic
effects without identifying the spatial kernel. If the shock is
national within a supplier industry, every origin in that industry is hit and
the destination-conditional sourcing shares sum out. A shock with geographic
content is different because the spatial kernel determines where downstream
exposure lands.
Table~\ref{tab:empirical_application_summary} applies this distinction to two
shock designs. The first is a Katrina-style Gulf regional shock:
\[
z_{ri}^{Gulf}
=
\mathbf{1}\{r\in\{\mathrm{LA},\mathrm{MS}\}\}.
\]
The shock hits Louisiana and Mississippi in every supplier sector. Its
geographic content is sharp while its industry content is broad. The second is a national manufacturing-input shock:
\[
z_{ri}^{Mfg}
=
\mathbf{1}\{i=\mathrm{manufacturing}\}.
\]
It is best interpreted here as an industry-shock limiting case rather than a
detailed tariff counterfactual.
\begin{table}[htbp]
\centering
\caption{Gulf and Manufacturing Shock Comparisons}
\label{tab:empirical_application_summary}
\normalsize
\begin{tabular}{lrrrrr}
\toprule
Application & Domar error & Regional TV & Rank corr. & Top overlap & Loss ratio \\
\midrule
Gulf regional & 0.008 & 0.137 & 0.887 & 0.598 & 0.860 \\
Manufacturing industry & 0 & 0 & 1.000 & 1.000 & 1.000 \\
\bottomrule
\end{tabular}
\begin{flushleft}
\footnotesize Notes: The table compares conditional independence with the
structured-gravity completion. Domar error is the absolute difference in
output-weighted accounting loss. Regional TV compares state-level loss
allocations. Top overlap compares the top decile of state-sector losses.
\end{flushleft}
\end{table}
For the Gulf regional shock, the point-completion aggregate difference is not
negligible. The Domar-weighted loss under conditional independence is
0.860{} of the structured-gravity loss. The regional allocation also
changes substantially: state-level allocation TV is 0.137{}, and
top-decile overlap is 0.598{}. Thus the proportional completion
changes both aggregate accounting loss and which downstream places are
classified as highly exposed. For the national manufacturing shock, all
metrics are invariant up to numerical precision, as predicted by
Proposition~\ref{prop:shock_cancellation}.
The ranking changes are most visible outside the directly shocked states.
Louisiana and Mississippi remain the two most exposed destination states under
both completions. But Arkansas is ranked 3{} under
structured gravity and 36{} under conditional independence.
California moves from rank 10{} to
3{}, and New York moves from rank
20{} to 5{}, under conditional
independence. These are point-completion comparisons, not sharp identified
rankings, but they show why regional incidence cannot be inferred from
aggregate exposure alone.
\section{Conclusion}
\label{sec:conclusion}
Regional production-network calculations require assumptions about who buys
from whom across space. Standard data provide national industry IO tables and
regional sectoral activity, but they do not observe the bilateral
state-sector buyer-seller matrix. This paper shows that those data identify
an admissible set of intermediate-input spatial kernels, not a unique regional
network or a full regional IO table. A completed regional IO matrix should
therefore be interpreted as a maintained restriction on the missing coupling,
not as an observed input.
This distinction matters because the missing spatial kernel is not equally
relevant for all propagation questions. Pure national industry shocks are
invariant to the kernel and therefore provide a limiting case for
regionalization assumptions. Local regional and region-sector shocks are
different because the kernel determines where downstream exposure lands. For
linear one-step exposure, the admissible set delivers sharp
transportation-program bounds. For nonlinear Leontief accounting exposure,
the same problem becomes a nonlinear identified-set problem. The perturbation
calculation provides conservative outer bounds rather than sharp global
bounds.
The U.S. state-sector application shows that the issue is quantitatively
relevant. CFS shipment moments imply substantially more home bias and distance
concentration than conditional independence, so the proportional completion is
too spatially diffuse in shipment-covered sectors. At the same time, these
moments do not identify the full intermediate-input coupling or the
state-level incidence of local shocks. For a Gulf regional shock, selected
bilateral CFS cells narrow the state exposure intervals, but they determine
few state-pair rankings. Banded wage-bill and employment margins remove the
exact-margin aggregate cancellation, yet state-level rankings remain weakly
identified.
These bounds can restrict quantitative spatial models, but they do not
replace them. A full counterfactual still requires final-demand sourcing,
household-demand geography, behavioral elasticities, factor adjustment,
financing, and market clearing. The point is to separate the
intermediate-input spatial network features supported by proxy margins and
auxiliary shipment moments from those imposed by regionalization assumptions.
Rather than evaluating a structural model at a single proportional
regionalization, researchers can ask whether its exposure and counterfactual
conclusions survive over the admissible set of spatial kernels.
\clearpage
\onehalfspacing
\bibliographystyle{aer}
\bibliography{state_spatial_kernel_refs}
\clearpage