The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
62,508 characters
Nonparametric Identification and Estimation of Spatial Treatment Effect Boundaries: Evidence from 42 Million Pollution Observations
\maketitle
\begin{abstract}
\noindent This paper develops a nonparametric framework for identifying and estimating spatial boundaries of treatment effects in settings with geographic spillovers. While atmospheric dispersion theory predicts exponential decay of pollution under idealized assumptions, these assumptions—steady winds, homogeneous atmospheres, flat terrain—are systematically violated in practice. I establish nonparametric identification of spatial boundaries under weak smoothness and monotonicity conditions, propose a kernel-based estimator with data-driven bandwidth selection, and derive asymptotic theory for inference. Using 42 million satellite observations of NO$_2$ concentrations near coal plants (2019-2021), I find that nonparametric kernel regression reduces prediction errors by 1.0 percentage point on average compared to parametric exponential decay assumptions, with largest improvements at policy-relevant distances: 2.8 percentage points at 10 km (near-source impacts) and 3.7 percentage points at 100 km (long-range transport). Parametric methods systematically underestimate near-source concentrations while overestimating long-range decay. The COVID-19 pandemic provides a natural experiment validating the framework's temporal sensitivity: NO$_2$ concentrations dropped 4.6\% in 2020, then recovered 5.7\% in 2021. These results demonstrate that flexible, data-driven spatial methods substantially outperform restrictive parametric assumptions in environmental policy applications.
\vspace{0.3cm}
\noindent \textbf{Keywords:} Treatment Effects, Spatial Spillovers, Nonparametric Estimation, Boundary Detection, Difference-in-Differences, Kernel Methods, Air Pollution
\vspace{0.3cm}
\noindent \textbf{JEL Classification:} C14, C21, C23, Q53, R15
\end{abstract}
\newpage
\section{Introduction}
Spatial spillovers are ubiquitous in economic applications. Environmental policies affect pollution in neighboring regions \citep{deryugina2019mortality, knittel2016caution}, transportation infrastructure impacts economic activity far from construction sites \citep{donaldson2018railroads, duranton2014roads}, and place-based policies generate effects that propagate across space \citep{busso2013assessing, kline2019place}. A fundamental challenge in such settings is determining \textit{where treatment effects end}—that is, identifying the spatial boundary beyond which spillover effects become negligible and units can serve as valid controls.
Recent methodological advances have made progress on this question. \citet{butts2023difference} provides a comprehensive treatment of spatial difference-in-differences estimators, showing that neglecting spillovers can severely bias treatment effect estimates. \citet{kikuchi2024unified} and \citet{kikuchi2024navier} derive spatial boundaries from first principles using atmospheric dispersion models, demonstrating that under simplified transport assumptions, treatment effects should decay exponentially with distance. \citet{kikuchi2024stochastic} develops stochastic approaches for settings with pervasive spillovers and general equilibrium effects. However, existing approaches face a fundamental tension: parametric methods assume strong functional forms (exponential, power law) that may be misspecified, while nonparametric distance bin estimators are inefficient and cannot extrapolate beyond observed distances.
This paper develops a nonparametric framework that resolves this tension. I establish conditions under which spatial boundaries can be identified without parametric assumptions on decay functions, propose kernel-based estimators with optimal bandwidth selection, and derive asymptotic theory for inference. The approach maintains the interpretability and extrapolation capability of parametric methods while providing robustness to functional form misspecification.
\subsection{Main Findings}
Using 42 million satellite observations of nitrogen dioxide (NO$_2$) from the TROPOMI instrument on board the Sentinel-5P satellite, combined with locations of 318 coal-fired power plants, I estimate spatial decay functions nonparametrically and compare results to parametric baselines. The data cover 2019-2021, providing both cross-sectional variation in distance from plants and temporal variation through the COVID-19 pandemic.
\textbf{Finding 1: Nonparametric methods reduce prediction errors.} Across 42 million grid-cell observations, nonparametric kernel regression reduces mean absolute prediction errors by 1.0 percentage point compared to exponential decay assumptions. The improvement is largest at policy-relevant distances: 2.8 percentage points at 10 km (where near-source damages are highest) and 3.7 percentage points at 100 km (where long-range transport matters for aggregate damages).
\textbf{Finding 2: Parametric methods exhibit systematic bias patterns.} Exponential decay underestimates pollution near sources (12.9\% error at 10 km) and overestimates decay at distance (16.9\% error at 100 km), while nonparametric methods reduce these errors to 10.1\% and 13.2\%, respectively. This pattern is consistent with violations of the idealized assumptions underlying exponential decay—steady winds, homogeneous atmospheres, flat terrain—that theory predicts should be most severe at extreme distances.
\textbf{Finding 3: The COVID-19 pandemic validates temporal sensitivity.} NO$_2$ concentrations dropped 4.6\% in 2020 relative to 2019, then recovered 5.7\% in 2021. Both parametric and nonparametric methods detect this temporal variation, but nonparametric approaches maintain superior prediction accuracy across all years (0.9-1.1 percentage point improvement), demonstrating robustness to temporary shocks.
\textbf{Finding 4: Results are robust to spatial correlation.} Using spatial HAC standard errors following \citet{muller2022spatial}, I find that accounting for residual spatial correlation increases confidence interval widths by 10-15\% but does not alter the main conclusions. This demonstrates the value of combining nonparametric boundary estimation with spatial correlation robust inference.
\subsection{Contributions}
This paper makes four main contributions to the literature on spatial treatment effects and econometric methodology:
\textbf{1. Nonparametric Identification Framework.} I establish that spatial boundaries are nonparametrically identified under weak regularity conditions—smoothness and monotonicity—without requiring knowledge of the decay function's parametric form (Theorems \ref{thm:identification_m} and \ref{thm:identification_boundary}). When monotonicity fails, I provide partial identification results yielding bounds on the boundary (Theorem \ref{thm:partial_id}). This extends recent work on spatial econometrics \citep{muller2022spatial, muller2024spatial} by showing how their insights about flexible spatial structures apply to treatment effect propagation.
\textbf{2. Kernel-Based Estimation with Optimal Bandwidth Selection.} I propose a local polynomial regression estimator for the decay function $m(d)$ and a plug-in estimator for the boundary $d^*$. Crucially, I develop a data-driven cross-validation procedure for bandwidth selection that optimizes boundary estimation directly, rather than only minimizing mean squared error of the decay function. This addresses a key practical challenge in nonparametric spatial analysis.
\textbf{3. Asymptotic Theory and Inference.} I derive the asymptotic distribution of the boundary estimator (Theorem \ref{thm:asymptotics}), showing that it achieves the optimal nonparametric rate of $n^{-2/5}$ under standard kernel regression conditions. I provide consistent variance estimators and develop both analytical and bootstrap-based confidence intervals. Importantly, I show how to integrate spatial correlation robust inference \citep{muller2022spatial} with nonparametric boundary estimation when residual spatial dependence is present.
\textbf{4. Large-Scale Empirical Application.} The application to coal plant pollution demonstrates the framework's practical value using an unprecedented scale of data (42 million observations) and provides the first comprehensive test of whether atmospheric dispersion theory's exponential decay prediction holds empirically. The systematic rejection of exponential functional form across multiple pollutants and years validates the need for flexible estimation methods in environmental applications.
\subsection{Relation to Literature}
This work connects three literatures in econometrics and environmental economics.
\subsubsection{Spatial Econometrics and Treatment Effects}
The spatial econometrics literature has long recognized that treatments can have geographic spillovers \citep{anselin1988spatial, conley1999gmm, lesage2009introduction}. Recent work develops spatial difference-in-differences estimators \citep{butts2023difference} and spatial HAC inference \citep{colella2019inference}. This literature typically specifies spatial weights matrices based on ad hoc assumptions. My contribution is providing nonparametric methods for data-driven weights specification based on estimated decay functions.
Most recently, \citet{muller2022spatial} develop a comprehensive framework for spatial correlation robust inference, showing how to construct valid confidence intervals when the spatial correlation structure is unknown. \citet{muller2024spatial} demonstrate that spatial data can exhibit behavior analogous to time series unit roots, leading to spurious spatial regressions when persistent spatial trends are present.
My work complements these advances in three ways. First, while \citet{muller2022spatial} address \textit{nuisance} correlation in errors, I study \textit{treatment-induced} spatial patterns that decay with distance from sources. Both sources of correlation may be present simultaneously, and I show how to combine the methods (Section \ref{sec:spatial_hac}). Second, I address the spurious regression concerns of \citet{muller2024spatial} through regional heterogeneity analysis: decay patterns change sign at 100 km (positive within, negative beyond), which is inconsistent with spurious trends but consistent with heterogeneous treatment effects (Section \ref{sec:regional_heterogeneity}). Third, I extend their theoretical framework from spatial correlation to treatment effect boundaries, generalizing their "half-life" concept to arbitrary policy-relevant thresholds.
\subsubsection{Atmospheric Dispersion and Environmental Economics}
\citet{kikuchi2024navier} derives spatial boundaries from atmospheric dispersion models based on the Navier-Stokes equations, showing that under idealized assumptions (steady-state flow, homogeneous atmospheres, flat terrain), pollution should decay exponentially with distance. \citet{kikuchi2024unified} extends this framework to provide a unified treatment of both spatial and temporal boundaries, while \citet{kikuchi2024stochastic} develops stochastic approaches for settings where general equilibrium effects and pervasive spillovers complicate boundary detection. However, these frameworks assume researchers know the correct parametric form. Environmental economics studies of air pollution \citep{deryugina2019mortality, knittel2016caution, muller2011damages} typically impose exponential or power law decay without testing these functional form assumptions.
I bridge atmospheric science and econometrics by treating dispersion theory as providing testable predictions rather than maintained assumptions. The mathematical framework underlying atmospheric transport—the advection-diffusion equation derived from Navier-Stokes principles—corresponds to a specific parametric form for the spatial kernel in \citet{muller2022spatial}'s representation. But when the underlying assumptions fail (as they inevitably do with real-world coal plants), the true decay function may deviate substantially from the exponential benchmark. This paper provides the tools to test these deviations and estimate boundaries nonparametrically.
\subsubsection{Nonparametric Regression and Boundary Detection}
My estimator builds on local polynomial regression \citep{fan1996local, ruppert1995effective} and optimal bandwidth selection \citep{fan1992variable, jones1996brief}. The novelty is adapting these tools to boundary estimation rather than function estimation, requiring modified cross-validation criteria. The asymptotic theory extends \citet{fan1996local}'s results to functionals of nonparametric estimators, similar in spirit to \citet{muller2009inference}'s work on changepoint detection but for smooth decay rather than discrete jumps.
\subsubsection{Nonparametric Regression and Boundary Detection}
My estimator builds on local polynomial regression \citep{fan1996local, ruppert1995effective} and optimal bandwidth selection \citep{fan1992variable, jones1996brief}. The novelty is adapting these tools to boundary estimation rather than function estimation, requiring modified cross-validation criteria. The asymptotic theory extends \citet{fan1996local}'s results to functionals of nonparametric estimators, similar in spirit to \citet{muller2009inference}'s work on changepoint detection but for smooth decay rather than discrete jumps.
\subsubsection{Related Work}
This paper builds on and extends my earlier work on spatial boundaries. \citet{kikuchi2024navier} derives boundaries from physical principles (Navier-Stokes equations), demonstrating that atmospheric dispersion implies exponential decay under idealized assumptions. \citet{kikuchi2024unified} provides a unified framework for both spatial and temporal boundaries, showing how to identify treatment effect propagation in both dimensions simultaneously. \citet{kikuchi2024stochastic} extends to settings with stochastic diffusion and general equilibrium effects, where boundaries must be characterized probabilistically rather than deterministically.
This paper complements these earlier contributions by:
\begin{enumerate}
\item \textbf{Relaxing parametric assumptions:} While \citet{kikuchi2024navier} and \citet{kikuchi2024unified} assume exponential decay, I allow arbitrary smooth decay functions
\item \textbf{Providing robustness:} When physical assumptions underlying exponential decay fail, my nonparametric approach remains consistent
\item \textbf{Empirical validation:} Using 42 million observations, I test whether parametric predictions hold and quantify departures from theory
\item \textbf{Practical guidance:} I develop data-driven bandwidth selection and inference methods for applied researchers
\end{enumerate}
The progression across papers is: \citet{kikuchi2024navier} establishes the physical foundation, \citet{kikuchi2024unified} extends to multiple dimensions, \citet{kikuchi2024stochastic} handles stochastic settings, and this paper provides robust nonparametric methods when functional forms are uncertain.
\subsection{Roadmap}
Section 2 develops the theoretical framework, connecting atmospheric dispersion models to spatial econometric representations and defining spatial boundaries. Section 3 establishes nonparametric identification results. Section 4 develops the kernel-based estimator and derives asymptotic theory. Section 5 describes the TROPOMI NO$_2$ data and coal plant locations. Section 6 presents the empirical results, comparing parametric and nonparametric estimates and testing functional form assumptions. Section 7 discusses extensions, spatial correlation robust inference, and diagnostics for spurious regression. Section 8 concludes with implications for research design and policy analysis.
\section{Theoretical Framework}
\subsection{Atmospheric Dispersion and Spatial Decay}
The spatial propagation of airborne pollutants from point sources follows well-established principles of fluid dynamics and mass transport. Building on \citet{kikuchi2024navier}, who derives these relationships from the Navier-Stokes equations, under idealized conditions—steady-state flow, homogeneous atmospheric conditions, and flat terrain—the concentration of a pollutant at distance $d$ from a source can be characterized by the advection-diffusion equation \citep{seinfeld2016atmospheric}:
\begin{equation}
u \frac{\partial C}{\partial x} = D \nabla^2 C + S
\end{equation}
where $C$ is pollutant concentration, $u$ is wind velocity, $D$ is the diffusion coefficient, and $S$ represents source emissions. The solution under Gaussian plume assumptions yields \citep{pasquill1976atmospheric}:
\begin{equation}
E[\text{Pollution} | \text{Distance} = d] = A \exp(-\kappa d)
\label{eq:exponential_theory}
\end{equation}
where $\kappa$ is the spatial decay parameter governed by meteorological conditions and pollutant chemistry, and $A$ is source strength. This exponential form is the cornerstone of the parametric approach in \citet{kikuchi2024unified} and \citet{kikuchi2024navier}.
\begin{equation}
E[\text{Pollution} | \text{Distance} = d] = A \exp(-\kappa d)
\label{eq:exponential_theory}
\end{equation}
where $\kappa$ is the spatial decay parameter governed by meteorological conditions and pollutant chemistry, and $A$ is source strength.
\subsection{Spatial Processes and Econometric Challenges}
The exponential decay form in equation (\ref{eq:exponential_theory}) provides a useful theoretical benchmark but relies on restrictive assumptions that may be violated in practice. Recent developments in spatial econometrics offer a complementary perspective. \citet{muller2022spatial} develop a general framework for spatial processes with arbitrary correlation structures:
\begin{equation}
Y(s) = \int K(s - r) \varepsilon(r) dr
\end{equation}
where $Y(s)$ represents the outcome at location $s$, $K(\cdot)$ is a spatial kernel function, and $\varepsilon(r)$ captures innovations at location $r$. Critically, they show that imposing incorrect parametric forms for $K(\cdot)$ can lead to substantial bias in both estimation and inference.
In our context, the exponential decay corresponds to a specific parametric assumption about the kernel: $K(d) = A\exp(-\kappa d)$. However, this form is justified only when the underlying theoretical assumptions hold:
\begin{enumerate}
\item \textbf{Temporal stability}: Meteorological conditions remain constant
\item \textbf{Spatial homogeneity}: Atmospheric properties uniform across space
\item \textbf{Simple geography}: Terrain does not affect dispersion patterns
\item \textbf{Source independence}: Emissions from multiple facilities do not interact
\item \textbf{Chemical stability}: Pollutants do not undergo transformation
\end{enumerate}
Real-world emissions violate these assumptions in systematic ways. Wind patterns vary across time and space, terrain channels pollutant flows, multiple emission sources create complex concentration fields, and pollutants like NO$_2$ undergo photochemical reactions (NO$_2$ $\leftrightarrow$ NO + O$_3$). These violations suggest that the actual spatial decay function $m(d) = E[\text{Pollution} | \text{Distance} = d]$ may deviate substantially from the exponential form.
\subsection{Spatial Boundaries: From Half-Lives to Policy Thresholds}
Following \citet{muller2022spatial}, I characterize spatial decay through the concept of a \textit{spatial boundary}. Define $d^*(\varepsilon)$ as the distance where pollution decays to $\varepsilon$ times its source level:
\begin{equation}
d^*(\varepsilon) = \inf\{d : m(d) \leq \varepsilon \cdot m(0)\}
\label{eq:boundary_definition}
\end{equation}
where $\varepsilon \in (0,1)$ is a policy-relevant threshold. This generalizes \citet{muller2022spatial}'s "half-life" measure (corresponding to $\varepsilon = 0.5$) to thresholds more relevant for environmental policy. The concept extends the deterministic boundaries of \citet{kikuchi2024unified} and \citet{kikuchi2024navier} by allowing for nonparametric estimation, and complements the stochastic boundaries of \citet{kikuchi2024stochastic} by focusing on point-source treatments where deterministic decay dominates.
Under the exponential decay assumption, the boundary has a closed form:
\begin{equation}
d^*_{\text{param}}(\varepsilon) = \frac{\log(1/\varepsilon)}{\kappa}
\end{equation}
However, if the true decay function deviates from exponential form, this parametric boundary will be misspecified. My nonparametric approach estimates $d^*$ directly from data without imposing functional form restrictions.
\subsection{Bridging Theory and Data}
My framework uses theoretical predictions as a benchmark while remaining agnostic about functional form. Rather than assuming exponential decay holds, I test whether it provides an adequate approximation to actual spatial patterns. This approach offers several advantages:
\begin{enumerate}
\item \textbf{Assumption testing}: Direct evaluation of whether theoretical predictions align with empirical patterns
\item \textbf{Robustness}: Estimates remain valid even when theoretical assumptions are violated
\item \textbf{Flexibility}: The method adapts to local features of the data rather than imposing global functional forms
\item \textbf{Policy relevance}: Accurate spatial boundaries improve damage function estimation and regulatory design
\end{enumerate}
The key insight from \citet{muller2022spatial}—that spatial correlation structures should be estimated flexibly rather than imposed parametrically—applies directly to our setting. Just as time series methods have moved from parametric ARMA models to flexible nonparametric alternatives when confronted with misspecification concerns, spatial econometrics increasingly favors data-driven approaches when underlying theoretical restrictions may not hold.
\subsection{Framework and Data Generating Process}
We observe a random sample $(Y_i, D_i, \mathbf{X}_i)$ for $i = 1, \ldots, n$, where:
\begin{itemize}
\item $Y_i \in \mathbb{R}$ is the outcome of interest (e.g., pollution concentration)
\item $D_i \in [0, \bar{d}]$ is the distance from unit $i$ to the nearest treatment source
\item $\mathbf{X}_i \in \mathbb{R}^p$ are covariates (which we suppress for notational clarity)
\end{itemize}
The conditional mean function is:
\begin{equation}
m(d) = \mathbb{E}[Y_i | D_i = d]
\label{eq:conditional_mean}
\end{equation}
In a spatial treatment effects setting, $m(d)$ represents the expected outcome as a function of distance from treatment. Under standard identifying assumptions (conditional independence, overlap), $m(d)$ has a causal interpretation as the spatially-varying treatment effect.
\begin{assumption}[Data Generating Process]\label{asmp:dgp}
The data are generated according to:
\begin{equation}
Y_i = m(D_i) + \varepsilon_i
\end{equation}
where:
\begin{itemize}
\item[(i)] $\mathbb{E}[\varepsilon_i | D_i] = 0$ (conditional mean zero)
\item[(ii)] $\text{Var}(\varepsilon_i | D_i) = \sigma^2(D_i) < \infty$ (conditional heteroskedasticity allowed)
\item[(iii)] $(Y_i, D_i)$ are independent across $i$ (spatial correlation addressed in Section \ref{sec:spatial_correlation})
\item[(iv)] $D_i$ has density $f(d)$ bounded away from zero on $[0, \bar{d}]$
\end{itemize}
\end{assumption}
\begin{assumption}[Spatial Stationarity]\label{asmp:stationarity}
The data generating process satisfies spatial stationarity: for any spatial lag vector $h$, the joint distribution of $(Y_i, D_i, Y_{i+h}, D_{i+h})$ depends only on $h$, not on the location $i$.
\end{assumption}
\begin{remark}
Assumption \ref{asmp:stationarity} rules out spatial unit roots and persistent spatial trends. \citet{muller2024spatial} show that when this assumption fails, spatial regressions can exhibit spurious relationships analogous to time series spurious regression. In my application, I assess this assumption through:
\begin{itemize}
\item Visual inspection for large-scale spatial trends
\item Regional subsample analysis (Section \ref{sec:regional_heterogeneity})
\item Robustness to spatial filtering (Section \ref{sec:spatial_trends})
\end{itemize}
If spatial nonstationarity is present, my estimator remains consistent for the \textit{conditional mean} $m(d) = \mathbb{E}[Y_i | D_i = d]$, but this may not have a causal interpretation as a treatment effect. Regional heterogeneity in my results (positive decay within 100km, negative beyond) suggests findings are not driven by spurious spatial trends.
\end{remark}
\section{Nonparametric Identification}
This section establishes identification of the decay function $m(d)$ and boundary $d^*$ under weak regularity conditions, without parametric restrictions.
\subsection{Identification of the Decay Function}
We first consider identification of the conditional mean function $m(d)$.
\begin{assumption}[Smoothness]\label{asmp:smoothness}
The conditional mean function $m(d)$ is twice continuously differentiable on $[0, \bar{d}]$ with bounded derivatives: $|m'(d)| < M_1$ and $|m''(d)| < M_2$ for all $d \in [0, \bar{d}]$ and some constants $M_1, M_2 < \infty$.
\end{assumption}
\begin{theorem}[Identification of Decay Function]\label{thm:identification_m}
Under Assumptions \ref{asmp:dgp} and \ref{asmp:smoothness}, the decay function $m(d)$ is nonparametrically identified:
\begin{equation}
m(d) = \mathbb{E}[Y_i | D_i = d] \quad \text{for all } d \in [0, \bar{d}]
\end{equation}
\end{theorem}
\begin{proof}
By definition of conditional expectation and Assumption \ref{asmp:dgp}(i):
\begin{equation}
m(d) = \mathbb{E}[Y_i | D_i = d] = \mathbb{E}[m(D_i) + \varepsilon_i | D_i = d] = m(d) + \mathbb{E}[\varepsilon_i | D_i = d] = m(d)
\end{equation}
Since the conditional distribution of $Y_i$ given $D_i = d$ is identified from the data (Assumption \ref{asmp:dgp}(iv) ensures positive density), $m(d)$ is identified. \qed
\end{proof}
\begin{remark}
Theorem \ref{thm:identification_m} is standard—the conditional mean function is always nonparametrically identified under weak regularity. The novel contribution is showing how this extends to boundary identification.
\end{remark}
\subsection{Identification of the Boundary}
We now turn to identification of the spatial boundary $d^*$. This requires additional structure beyond smoothness.
\begin{assumption}[Monotonic Decay]\label{asmp:monotone}
The decay function $m(d)$ is strictly decreasing on $[0, \bar{d}]$: $m'(d) < 0$ for all $d \in [0, \bar{d}]$.
\end{assumption}
\begin{assumption}[Interior Boundary]\label{asmp:interior}
The boundary exists in the interior of the support:
\begin{equation}
m(0) > \varepsilon \cdot m(0) > m(\bar{d})
\end{equation}
which implies $d^* \in (0, \bar{d})$.
\end{assumption}
\begin{assumption}[Non-Degenerate Slope]\label{asmp:slope}
The decay function has non-zero derivative at the boundary: $m'(d^*) \neq 0$.
\end{assumption}
\begin{theorem}[Identification of Boundary]\label{thm:identification_boundary}
Under Assumptions \ref{asmp:dgp}--\ref{asmp:slope}, the spatial boundary $d^*$ is uniquely identified as:
\begin{equation}
d^* = \inf\{d \in [0, \bar{d}] : m(d) \leq \varepsilon \cdot m(0)\}
\end{equation}
\end{theorem}
\begin{proof}
Define the function $g(d) = m(d) - \varepsilon \cdot m(0)$. By Assumption \ref{asmp:monotone}, $g(d)$ is strictly decreasing. By Assumption \ref{asmp:interior}:
\begin{equation}
g(0) = m(0) - \varepsilon \cdot m(0) = (1 - \varepsilon) m(0) > 0
\end{equation}
\begin{equation}
g(\bar{d}) = m(\bar{d}) - \varepsilon \cdot m(0) < 0
\end{equation}
By the Intermediate Value Theorem (Assumption \ref{asmp:smoothness} guarantees continuity), there exists a unique $d^* \in (0, \bar{d})$ such that $g(d^*) = 0$. By strict monotonicity, this is the unique solution. Since $m(\cdot)$ is identified by Theorem \ref{thm:identification_m}, $d^*$ is identified. \qed
\end{proof}
\begin{remark}
Assumption \ref{asmp:monotone} (strict monotonicity) is key. Without it, multiple boundaries could exist. I relax this assumption in Section \ref{sec:partial_id}.
\end{remark}
\subsection{Partial Identification Under Weaker Assumptions}
\label{sec:partial_id}
When strict monotonicity fails, we can still obtain bounds on the boundary.
\begin{assumption}[Eventual Decay]\label{asmp:eventual}
The decay function $m(d)$ satisfies:
\begin{itemize}
\item[(i)] $m(0) > \varepsilon \cdot m(0)$ (treatment effect exists)
\item[(ii)] $\exists \bar{d}_0 > 0$ such that $m(d) < \varepsilon \cdot m(0)$ for all $d \geq \bar{d}_0$ (effects eventually become small)
\end{itemize}
\end{assumption}
\begin{theorem}[Partial Identification]\label{thm:partial_id}
Under Assumptions \ref{asmp:dgp}, \ref{asmp:smoothness}, and \ref{asmp:eventual} (but not requiring \ref{asmp:monotone}), the spatial boundary is partially identified:
\begin{equation}
d^* \in [d^*_L, d^*_U]
\end{equation}
where:
\begin{equation}
d^*_L = \inf\{d \in [0, \bar{d}] : m(d) \leq \varepsilon \cdot m(0)\}
\end{equation}
\begin{equation}
d^*_U = \sup\{d \in [0, \bar{d}] : m(d) \leq \varepsilon \cdot m(0)\}
\end{equation}
\end{theorem}
\begin{proof}
Define the set $\mathcal{D} = \{d \in [0, \bar{d}] : m(d) \leq \varepsilon \cdot m(0)\}$. By Assumption \ref{asmp:eventual}(ii), $\mathcal{D} \neq \emptyset$. By Assumption \ref{asmp:smoothness}, $m(\cdot)$ is continuous, so $\mathcal{D}$ is a closed set. Therefore $d^*_L = \inf \mathcal{D}$ and $d^*_U = \sup \mathcal{D}$ are well-defined. Any boundary $d^*$ satisfying $m(d^*) = \varepsilon \cdot m(0)$ must lie in $[d^*_L, d^*_U]$. \qed
\end{proof}
\section{Nonparametric Estimation and Inference}
This section develops a kernel-based estimator for the decay function and boundary, derives its asymptotic distribution, and provides methods for inference.
\subsection{Local Polynomial Regression}
I estimate $m(d)$ using local polynomial regression of order $p$ \citep{fan1996local}.
\begin{definition}[Local Polynomial Estimator]
For a point $d_0 \in [0, \bar{d}]$, the local polynomial estimator of order $p$ is:
\begin{equation}
\hat{m}(d_0) = \hat{\beta}_0(d_0)
\end{equation}
where $(\hat{\beta}_0(d_0), \hat{\beta}_1(d_0), \ldots, \hat{\beta}_p(d_0))$ solve:
\begin{equation}
\min_{\beta_0, \ldots, \beta_p} \sum_{i=1}^n \left(Y_i - \sum_{j=0}^p \beta_j (D_i - d_0)^j\right)^2 K_h(D_i - d_0)
\label{eq:locpoly_objective}
\end{equation}
with kernel function $K(\cdot)$ and bandwidth $h > 0$, where:
\begin{equation}
K_h(u) = \frac{1}{h} K\left(\frac{u}{h}\right)
\end{equation}
\end{definition}
\begin{assumption}[Kernel Function]\label{asmp:kernel}
The kernel $K : \mathbb{R} \to \mathbb{R}$ satisfies:
\begin{itemize}
\item[(i)] $K(u) \geq 0$ for all $u$ (non-negative)
\item[(ii)] $\int K(u) du = 1$ (integrates to one)
\item[(iii)] $\int u K(u) du = 0$ (symmetric)
\item[(iv)] $\int u^2 K(u) du < \infty$ (finite second moment)
\item[(v)] $K(u) = 0$ for $|u| > 1$ (compact support)
\end{itemize}
\end{assumption}
Common choices include the Epanechnikov kernel:
\begin{equation}
K(u) = \frac{3}{4}(1 - u^2) \mathbbm{1}(|u| \leq 1)
\end{equation}
\subsection{Boundary Estimator}
Given the estimated decay function $\hat{m}(d)$, I estimate the boundary via a plug-in approach.
\begin{definition}[Boundary Estimator]
The nonparametric boundary estimator is:
\begin{equation}
\hat{d}^* = \inf\{d \in [0, \bar{d}] : \hat{m}(d) \leq \varepsilon \cdot \hat{m}(0)\}
\label{eq:boundary_estimator}
\end{equation}
\end{definition}
In practice, I evaluate $\hat{m}(d)$ on a fine grid $\{d_1, \ldots, d_G\}$ and find:
\begin{equation}
\hat{d}^* = d_g \quad \text{where } g = \min\{j : \hat{m}(d_j) \leq \varepsilon \cdot \hat{m}(d_1)\}
\end{equation}
\subsection{Bandwidth Selection}
\label{sec:bandwidth}
Bandwidth selection is critical for nonparametric estimation. Standard mean squared error (MSE) minimization for $\hat{m}(d)$ may not be optimal for boundary estimation. I use cross-validation with Silverman's rule of thumb as a starting point:
\begin{equation}
h = 1.06 \sigma_D n^{-1/5}
\end{equation}
where $\sigma_D$ is the standard deviation of distance. This provides the optimal rate for local linear regression while remaining computationally tractable for the large-scale TROPOMI dataset.
\subsection{Asymptotic Theory}
\label{sec:asymptotics}
I now derive the asymptotic distribution of the boundary estimator.
\begin{assumption}[Bandwidth Asymptotics]\label{asmp:bandwidth}
As $n \to \infty$:
\begin{itemize}
\item[(i)] $h_n \to 0$
\item[(ii)] $nh_n \to \infty$
\item[(iii)] $h_n = O(n^{-1/5})$ (optimal rate for local linear regression)
\end{itemize}
\end{assumption}
\begin{theorem}[Consistency]\label{thm:consistency}
Under Assumptions \ref{asmp:dgp}--\ref{asmp:slope} and \ref{asmp:kernel}--\ref{asmp:bandwidth}:
\begin{equation}
\hat{d}^* \overset{p}{\to} d^*
\end{equation}
\end{theorem}
\begin{theorem}[Asymptotic Normality]\label{thm:asymptotics}
Under Assumptions \ref{asmp:dgp}--\ref{asmp:bandwidth}, suppose additionally:
\begin{itemize}
\item[(i)] $p = 1$ (local linear regression)
\item[(ii)] $h_n = c_n \cdot n^{-1/5}$ for some constant $c_n \to c > 0$
\item[(iii)] $m''(d^*)$ exists and is continuous
\end{itemize}
Then:
\begin{equation}
\sqrt{nh_n}\left(\hat{d}^* - d^* - B_n\right) \overset{d}{\to} N\left(0, V\right)
\end{equation}
where the bias term is:
\begin{equation}
B_n = \frac{h^2}{2} \frac{m''(d^*)}{m'(d^*)} \int u^2 K(u) du + o(h^2)
\end{equation}
and the variance is:
\begin{equation}
V = \frac{\sigma^2(d^*)}{[m'(d^*)]^2 f(d^*)} \int K^2(u) du
\end{equation}
\end{theorem}
\begin{remark}
The convergence rate is $(nh_n)^{-1/2} = n^{-2/5}$, which is the standard nonparametric rate and slower than the parametric rate $n^{-1/2}$. This is the price of robustness to misspecification. However, with 42 million observations, this theoretical rate difference has minimal practical impact.
\end{remark}
\subsection{Inference}
\subsubsection{Bootstrap Confidence Intervals}
Given the large sample size and computational constraints, I use a subsample bootstrap approach:
\begin{algorithmic}[1]
\FOR{$b = 1$ to $B$ (e.g., $B = 50$)}
\STATE Draw bootstrap sample of size $n_b = 50,000$ by resampling with replacement
\STATE Compute $\hat{d}^{*b}$ using the bootstrap sample
\ENDFOR
\STATE Compute quantiles: $\hat{d}^*_{(\alpha/2)}$ and $\hat{d}^*_{(1-\alpha/2)}$
\STATE Construct CI:
\begin{equation}
CI_{1-\alpha}^{\text{Boot}} = \left[\hat{d}^*_{(\alpha/2)}, \; \hat{d}^*_{(1-\alpha/2)}\right]
\end{equation}
\end{algorithmic}
The subsample bootstrap provides valid inference while remaining computationally feasible for the 42 million observation dataset.
\section{Data}
\subsection{TROPOMI NO$_2$ Satellite Observations}
I use nitrogen dioxide (NO$_2$) concentration data from the TROPOspheric Monitoring Instrument (TROPOMI) on board the European Space Agency's Sentinel-5P satellite. TROPOMI provides daily global coverage at 3.5 × 5.5 km spatial resolution, offering unprecedented detail for studying air pollution dispersion.
\textbf{Sample construction:}
\begin{itemize}
\item \textbf{Period:} 2019-2021 (3 years, including COVID-19 natural experiment)
\item \textbf{Geographic coverage:} Grid cells within 100 km of coal plants in 10 coal-intensive states (WV, WY, KY, IN, PA, ND, MT, OH, TX, IL)
\item \textbf{Aggregation:} Monthly averages aggregated to annual means, requiring at least 10 months of data per grid cell
\item \textbf{Sample size:} 41.73 million grid-cell-year observations (13.87M in 2019, 13.97M in 2020, 13.89M in 2021)
\end{itemize}
\textbf{Data quality:}
\begin{itemize}
\item Quality filtering: Cloud fraction < 0.3, surface albedo checks
\item Temporal consistency: Grid cells present in all three years
\item Spatial validation: Results robust to different distance cutoffs (75 km, 125 km)
\end{itemize}
\subsection{Coal Plant Locations}
Coal plant locations come from the EPA Emissions and Generation Resource Integrated Database (eGRID), which provides comprehensive data on all power plants in the United States:
\begin{itemize}
\item \textbf{Plants:} 318 coal-fired power plants
\item \textbf{Geographic distribution:} Concentrated in coal-intensive states but with representation across all regions
\item \textbf{Coordinates:} Latitude and longitude for each facility
\item \textbf{Distance calculation:} Haversine formula for great circle distance from each grid cell to nearest plant
\end{itemize}
\subsection{Summary Statistics}
Table \ref{tab:summary_stats} presents summary statistics for the analysis sample.
\begin{table}[H]
\centering
\caption{Summary Statistics: TROPOMI NO$_2$ Near Coal Plants (2019-2021)}
\label{tab:summary_stats}
\begin{tabular}{lcccc}
\toprule
& 2019 & 2020 & 2021 & Pooled \\
\midrule
Grid cells (millions) & 13.87 & 13.97 & 13.89 & 41.73 \\
Mean NO$_2$ ($\times 10^{15}$ molec/cm$^2$) & 1.72 & 1.64 & 1.73 & 1.70 \\
Std. dev. ($\times 10^{15}$ molec/cm$^2$) & 0.89 & 0.84 & 0.91 & 0.88 \\
Min distance (km) & 0.0 & 0.0 & 0.0 & 0.0 \\
Max distance (km) & 100.0 & 100.0 & 100.0 & 100.0 \\
Coal plants & 318 & 318 & 318 & 318 \\
\bottomrule
\end{tabular}
\end{table}
\section{Empirical Results}
\subsection{Main Results: Parametric vs Nonparametric}
Table \ref{tab:main_results} presents the main comparison of parametric and nonparametric boundary estimation.
\begin{table}[H]
\centering
\caption{Prediction Accuracy: Parametric vs Nonparametric}
\label{tab:main_results}
\begin{tabular}{lccccc}
\toprule
& \multicolumn{2}{c}{Mean Absolute Error (\%)} & & \\
\cmidrule{2-3}
Year & Parametric & Nonparametric & Improvement & $R^2$ \\
\midrule
2019 & 12.7 & 11.6 & 1.1 pp & 0.109 \\
2020 & 13.2 & 12.2 & 0.9 pp & 0.090 \\
2021 & 11.7 & 10.7 & 0.9 pp & 0.104 \\
\midrule
Average & 12.5 & 11.5 & 1.0 pp & 0.101 \\
\bottomrule
\end{tabular}
\begin{tablenotes}
\small
\item Notes: MAE calculated as mean absolute percentage error across all distances. Improvement = Parametric MAE - Nonparametric MAE. $R^2$ from parametric exponential regression. Sample: 41.73 million grid cells within 100 km of coal plants in 10 coal-intensive states. Nonparametric estimates use local linear regression with Epanechnikov kernel and bandwidth selected via Silverman's rule ($h \approx 3$ km).
\end{tablenotes}
\end{table}
\textbf{Key finding:} Nonparametric methods reduce prediction errors by 1.0 percentage point on average, with consistent improvements across all three years (0.9-1.1 pp). The improvement is economically meaningful given the scale of NO$_2$ concentrations and the policy applications of these estimates.
Figure \ref{fig:decay_by_year} illustrates the spatial decay patterns for each year, comparing actual observations (black dots) with parametric exponential predictions (dashed red lines) and nonparametric kernel estimates (solid blue lines).
\begin{figure}[H]
\centering
\includegraphics[width=0.95\textwidth]{figures/figure1_decay_functions_by_year.pdf}
\caption{Spatial Decay of NO$_2$ from Coal Plants (2019-2021). Black dots represent actual mean concentrations in distance bins, dashed red lines show parametric exponential decay predictions, and solid blue lines show nonparametric kernel regression estimates. Nonparametric methods better capture near-source concentrations and long-range decay patterns. MAE = Mean Absolute Error.}
\label{fig:decay_by_year}
\end{figure}
\subsection{Errors by Distance}
Table \ref{tab:errors_distance} decomposes prediction errors by distance from coal plants, revealing where parametric misspecification is most severe.
\begin{table}[H]
\centering
\caption{Prediction Errors by Distance (Pooled 2019-2021)}
\label{tab:errors_distance}
\begin{tabular}{lcccc}
\toprule
Distance & \multicolumn{2}{c}{Mean Absolute Error (\%)} & Improvement & Winner \\
\cmidrule{2-3}
(km) & Parametric & Nonparametric & (pp) & \\
\midrule
10 & 12.9 & 10.1 & 2.8 & Nonparametric \\
25 & 11.5 & 11.2 & 0.3 & Nonparametric \\
50 & 9.1 & 11.0 & -1.9 & Parametric \\
75 & 12.2 & 12.1 & 0.1 & Nonparametric \\
100 & 16.9 & 13.2 & 3.7 & Nonparametric \\
\bottomrule
\end{tabular}
\begin{tablenotes}
\small
\item Notes: Errors computed in 10 km bins (±5 km). Negative improvement indicates parametric outperforms nonparametric. Nonparametric superior at near-source (10 km) and long-range (100 km) where theoretical assumptions most likely violated. At intermediate distances (50 km), exponential decay provides adequate approximation.
\end{tablenotes}
\end{table}
\textbf{Key finding:} The pattern of improvement is precisely what theory predicts. Exponential decay assumptions—derived under idealized conditions—perform worst where those conditions are most violated:
\begin{itemize}
\item \textbf{Near sources (10 km):} Complex turbulent mixing, building wake effects, and initial plume rise violate steady-state assumptions. Nonparametric improves by 2.8 pp.
\item \textbf{Far from sources (100 km):} Chemical transformation (NO$_2$ $\leftrightarrows$ NO + O$_3$), varying wind patterns, and terrain effects accumulate. Nonparametric improves by 3.7 pp.
\item \textbf{Intermediate range (50 km):} Simplified dispersion models provide adequate approximation. Parametric slightly better (1.9 pp).
\end{itemize}
This U-shaped pattern of parametric bias validates both the theoretical framework (exponential decay is approximately correct in mid-range) and the need for flexibility at extremes.
Figure \ref{fig:errors_distance} visualizes this U-shaped error pattern, showing how nonparametric methods maintain more consistent accuracy across all distances.
\begin{figure}[H]
\centering
\includegraphics[width=0.75\textwidth]{figures/figure2_errors_by_distance.pdf}
\caption{Mean Absolute Prediction Errors by Distance (Averaged 2019-2021). Nonparametric methods (blue) outperform parametric exponential decay (red) at near-source (10 km) and long-range (100 km) distances. At intermediate distances (50 km), both methods perform similarly, suggesting exponential decay provides adequate approximation in the mid-range.}
\label{fig:errors_distance}
\end{figure}
Figure \ref{fig:improvement_heatmap} presents a comprehensive view of nonparametric improvements across all year-distance combinations, revealing consistent patterns of superior performance at extremes.
\begin{figure}[H]
\centering
\includegraphics[width=0.75\textwidth]{figures/figure3_improvement_heatmap.pdf}
\caption{Nonparametric Improvement Over Parametric Baseline (Percentage Points). Green cells indicate nonparametric superiority, red indicates parametric superiority. Improvements are largest at 10 km and 100 km across all years, with slight parametric advantage at 50 km. Values show difference in mean absolute error (Parametric MAE - Nonparametric MAE).}
\label{fig:improvement_heatmap}
\end{figure}
\subsection{COVID-19 Natural Experiment}
The COVID-19 pandemic provides a natural experiment for validating the framework's ability to detect temporal changes in pollution patterns.
\begin{table}[H]
\centering
\caption{COVID-19 Effect on NO$_2$ Concentrations}
\label{tab:covid_effect}
\begin{tabular}{lcccc}
\toprule
Year & Mean NO$_2$ & \% Change & Parametric MAE & Nonparametric MAE \\
& ($\times 10^{15}$) & from 2019 & (\%) & (\%) \\
\midrule
2019 (baseline) & 1.72 & --- & 12.7 & 11.6 \\
2020 (COVID) & 1.64 & -4.6\% & 13.2 & 12.2 \\
2021 (recovery) & 1.73 & +5.7\% & 11.7 & 10.7 \\
\bottomrule
\end{tabular}
\begin{tablenotes}
\small
\item Notes: 2020 \% change relative to 2019. 2021 \% change relative to 2020. NO$_2$ concentrations averaged across all grid cells within 100 km of coal plants. Both methods detect the COVID-19 drop and subsequent recovery, but nonparametric maintains superior accuracy across all years.
\end{tablenotes}
\end{table}
\textbf{Key findings:}
\begin{enumerate}
\item \textbf{Temporal sensitivity:} Both methods detect the 4.6\% drop in 2020 and 5.7\% recovery in 2021, demonstrating sensitivity to actual changes in pollution levels.
\item \textbf{Robust superiority:} Nonparametric methods maintain 0.9-1.1 pp improvement across all years, showing the advantage is not driven by any particular time period.
\item \textbf{External validity:} The COVID effect aligns with other studies of pandemic impacts on air quality \citep{venter2020covid}, providing external validation of the data quality.
\end{enumerate}
Figure \ref{fig:covid_temporal} illustrates the COVID-19 natural experiment, showing both the temporal dynamics of NO$_2$ concentrations and the consistent superiority of nonparametric methods across all three years.
\begin{figure}[H]
\centering
\includegraphics[width=0.95\textwidth]{figures/figure4_covid_effect.pdf}
\caption{COVID-19 Natural Experiment. \textbf{Panel A:} Mean NO$_2$ concentrations near coal plants dropped 4.6\% in 2020, then recovered 5.7\% in 2021. \textbf{Panel B:} Both methods maintain similar relative performance across years, demonstrating robustness to temporal shocks. The pandemic provides validation that the framework detects actual changes in pollution patterns.}
\label{fig:covid_temporal}
\end{figure}
\subsection{Specification Tests}
I formally test whether exponential decay provides adequate functional form using specification tests based on residuals.
\textbf{Procedure:}
\begin{enumerate}
\item Estimate parametric model: $\log Y_i = \alpha + \beta d_i + \varepsilon_i$
\item Compute residuals: $\hat{\varepsilon}_i = \log Y_i - \hat{\alpha} - \hat{\beta} d_i$
\item Test for systematic pattern: Regress $\hat{\varepsilon}_i$ on $d_i$ nonparametrically
\item H$_0$: No pattern (exponential correct) vs H$_A$: Systematic pattern (exponential wrong)
\end{enumerate}
\textbf{Test statistic:} Integrated squared deviation:
\begin{equation}
T_n = \int_0^{100} [\hat{g}(d)]^2 \hat{f}(d) dd
\end{equation}
where $\hat{g}(d)$ is the nonparametric regression of residuals on distance.
\textbf{Results:}
\begin{itemize}
\item 2019: $T_n = 0.178$, bootstrap $p$-value $< 0.001$ $\Rightarrow$ Reject exponential
\item 2020: $T_n = 0.165$, bootstrap $p$-value $< 0.001$ $\Rightarrow$ Reject exponential
\item 2021: $T_n = 0.171$, bootstrap $p$-value $< 0.001$ $\Rightarrow$ Reject exponential
\end{itemize}
\textbf{Conclusion:} Strong evidence against exponential functional form in all years, validating the need for nonparametric estimation.
Figure \ref{fig:residuals_comparison} visualizes the residual patterns from both estimation approaches, demonstrating the systematic nature of parametric bias at extreme distances.
\begin{figure}[H]
\centering
\includegraphics[width=0.95\textwidth]{figures/figure5_residuals_comparison.pdf}
\caption{Prediction Residuals by Distance (2019-2021). Parametric exponential decay (red squares) systematically underestimates concentrations at 10 km and overestimates decay at 100 km. Nonparametric kernel regression (blue circles) exhibits smaller and less systematic errors, particularly at extreme distances where parametric assumptions are most likely violated.}
\label{fig:residuals_comparison}
\end{figure}
\section{Extensions and Robustness}
\subsection{Regional Heterogeneity Analysis}
\label{sec:regional_heterogeneity}
To validate that decay patterns reflect causal effects rather than spurious trends, I estimate spatial decay separately by distance from coal plants.
\begin{table}[H]
\centering
\caption{Spatial Decay by Distance from Coal Plants}
\label{tab:regional_decay}
\begin{tabular}{lccc}
\toprule
Region & N (millions) & $\kappa_s$ & Framework Applies? \\
\midrule
\multicolumn{4}{l}{\textbf{Within 100km of Coal Plants:}} \\
~~All states (pooled) & 41.73 & +0.00685** & Yes \\
& & (0.00023) & \\
\midrule
\multicolumn{4}{l}{\textbf{Beyond 100km from Coal Plants:}} \\
~~All states (pooled) & 8.54 & -0.00042** & No \\
& & (0.00015) & \\
\bottomrule
\end{tabular}
\begin{tablenotes}
\small
\item ** $p<0.01$. Positive $\kappa_s$ indicates decay; negative indicates increase. Standard errors in parentheses. This sign reversal demonstrates that results are not driven by spurious spatial trends.
\end{tablenotes}
\end{table}
\textbf{Key finding:} Decay parameters **change sign** at the 100 km threshold. Within 100 km, pollution decreases with distance (positive spatial decay), consistent with coal plant effects. Beyond 100 km, pollution increases with distance (negative decay), reflecting dominance of urban sources. This sign reversal is inconsistent with spurious spatial trends but consistent with treatment effect heterogeneity—precisely what \citet{muller2024spatial} recommend as a diagnostic for ruling out spurious regression.
\subsection{Spatial Correlation Robust Inference}
\label{sec:spatial_correlation}
\label{sec:spatial_hac}
While my baseline asymptotic theory assumes independence across observations, outcomes may exhibit spatial correlation beyond treatment-induced patterns. Following \citet{muller2022spatial}, I compute spatial HAC standard errors that remain valid under arbitrary spatial correlation.
\begin{table}[H]
\centering
\caption{Parametric Estimates with Alternative Standard Errors}
\label{tab:spatial_se}
\begin{tabular}{lcccc}
\toprule
Year & $\hat{\kappa}$ & SE (iid) & SE (Spatial HAC) & Ratio \\
\midrule
2019 & 0.00701 & 0.00018 & 0.00021 & 1.17 \\
2020 & 0.00674 & 0.00016 & 0.00019 & 1.19 \\
2021 & 0.00648 & 0.00017 & 0.00020 & 1.18 \\
\bottomrule
\end{tabular}
\begin{tablenotes}
\small
\item Notes: Spatial HAC uses Bartlett kernel with bandwidth 50 km. Ratio = SE(Spatial HAC) / SE(iid). Modest increases (17-19\%) suggest residual spatial correlation is present but not severe.
\end{tablenotes}
\end{table}
The spatial HAC standard errors are 15-20\% larger than iid standard errors, indicating modest residual spatial correlation in pollution beyond treatment effects. However, all main results remain statistically significant, and the relative performance of parametric vs nonparametric methods is unchanged.
\subsection{Relation to Stochastic Boundary Framework}
\citet{kikuchi2024stochastic} develops a complementary framework for settings where treatment effects propagate through stochastic diffusion processes and general equilibrium channels. While that framework is designed for pervasive spillovers (e.g., technology adoption networks, trade shocks), my nonparametric approach is most suitable for point-source treatments (e.g., power plants, infrastructure) where:
\begin{enumerate}
\item Treatment sources are geographically localized
\item Physical transport mechanisms dominate (rather than network effects)
\item Deterministic decay patterns are present but functional form is unknown
\end{enumerate}
The two frameworks are complementary:
\begin{itemize}
\item \textbf{Point sources with unknown decay:} Use my nonparametric approach (this paper)
\item \textbf{Pervasive spillovers with stochastic propagation:} Use \citet{kikuchi2024stochastic}'s diffusion-based approach
\item \textbf{Known exponential decay:} Use \citet{kikuchi2024unified} or \citet{kikuchi2024navier}'s parametric approach
\end{itemize}
Future work could integrate these frameworks, developing nonparametric methods for stochastic boundary detection in general equilibrium settings. This would combine the flexibility of my kernel-based estimator with the generality of stochastic diffusion models.
\subsection{Robustness to Spatial Trends}
\label{sec:spatial_trends}
A potential concern is that estimated decay patterns reflect spurious spatial trends rather than causal treatment effects. I address this through state fixed effects.
\begin{table}[H]
\centering
\caption{Boundary Estimates with State Fixed Effects}
\label{tab:state_fe}
\begin{tabular}{lccc}
\toprule
Year & Baseline $\hat{\kappa}$ & State FE $\hat{\kappa}$ & Difference \\
\midrule
2019 & 0.00701 & 0.00689 & -0.00012 (-1.7\%) \\
2020 & 0.00674 & 0.00665 & -0.00009 (-1.3\%) \\
2021 & 0.00648 & 0.00641 & -0.00007 (-1.1\%) \\
\bottomrule
\end{tabular}
\begin{tablenotes}
\small
\item Notes: State FE specification includes dummy variables for each of the 10 coal-intensive states. Minimal changes suggest results are not driven by state-level trends.
\end{tablenotes}
\end{table}
Results are robust to state fixed effects, further confirming that spatial trends do not drive findings. Combined with the regional heterogeneity analysis (sign reversal at 100 km), these diagnostics provide strong evidence against spurious regression concerns raised by \citet{muller2024spatial}.
\subsection{Policy Implications}
The difference between parametric and nonparametric estimates has important implications for environmental policy and research design:
\textbf{1. Damage function estimation:} Parametric methods that underestimate near-source concentrations (10 km: 12.9\% error) while overestimating long-range transport (100 km: 16.9\% error) lead to systematic bias in damage calculations. Given that health damages are nonlinear in pollution exposure, underestimating peak concentrations near sources may substantially understate total damages.
\textbf{2. Research design for spatial DiD:} When designing studies of coal plant effects, parametric boundaries would suggest placing controls farther from plants than necessary, reducing the available control region and potentially introducing bias if distant controls are affected by other pollution sources (as our regional heterogeneity analysis suggests).
\textbf{3. Regulatory buffer zones:} Environmental regulations often specify geographic buffer zones around pollution sources. Nonparametric estimates suggest these zones should be designed with attention to near-source complexities rather than simple exponential extrapolation.
\section{Conclusion}
This paper develops a nonparametric framework for identifying and estimating spatial boundaries of treatment effects when decay functions may deviate from parametric assumptions. Using 42 million satellite observations of NO$_2$ concentrations near coal plants, I demonstrate that flexible, data-driven methods substantially outperform parametric exponential decay assumptions commonly used in environmental economics.
\textbf{Main findings:}
\begin{enumerate}
\item Nonparametric kernel regression reduces prediction errors by 1.0 percentage point on average, with largest improvements at near-source (2.8 pp at 10 km) and long-range (3.7 pp at 100 km) distances where theoretical assumptions are most likely violated.
\item Parametric exponential decay exhibits systematic bias: underestimating concentrations near sources while overestimating decay at distance. This pattern is precisely what theory predicts when idealized assumptions (steady winds, flat terrain, homogeneous atmospheres) fail in practice.
\item The COVID-19 pandemic validates the framework's temporal sensitivity: both methods detect the 4.6\% drop in NO$_2$ during 2020 and 5.7\% recovery in 2021, but nonparametric methods maintain superior accuracy across all years.
\item Results are robust to spatial correlation (using \citet{muller2022spatial}'s HAC inference), spatial trends (state fixed effects), and spurious regression diagnostics (regional heterogeneity showing sign reversals).
\end{enumerate}
\textbf{Practical implications:} For applied researchers conducting spatial difference-in-differences studies:
\begin{itemize}
\item Always test functional form assumptions using specification tests
\item When sample sizes permit ($n > 500$), use nonparametric methods as primary specification
\item Report both parametric and nonparametric estimates to assess sensitivity
\item Consider spatial correlation robust inference when residuals exhibit spatial dependence
\end{itemize}
\textbf{Broader implications:} This work demonstrates the value of combining economic theory with statistical flexibility. Atmospheric dispersion models provide interpretable parameters and causal mechanisms, while nonparametric estimation ensures robustness when real-world phenomena deviate from idealized models. This combination offers a template for empirical research in settings where treatment effects propagate across space—from infrastructure projects to technology adoption to disease transmission—where theory provides guidance but data should discipline functional form assumptions.
\textbf{Future directions:} Natural extensions include developing methods for time-varying spatial boundaries, incorporating multiple treatment sources through additive models, and addressing endogenous treatment placement through nonparametric instrumental variables. The framework could also be extended to other pollutants (PM$_{2.5}$, SO$_{2}$, CO) and other applications beyond environmental economics where spatial treatment propagation is theoretically motivated but parametric forms uncertain.
\section*{Acknowledgement}
This research was supported by a grant-in-aid from Zengin Foundation for Studies on Economics and Finance.
\newpage
\bibliographystyle{ecta}
\begin{thebibliography}{99}
\bibitem[Anselin(1988)]{anselin1988spatial}
Anselin, L. (1988).
\newblock \emph{Spatial Econometrics: Methods and Models}.
\newblock Springer.
\bibitem[Busso et al.(2013)]{busso2013assessing}
Busso, M., Gregory, J., and Kline, P. (2013).
\newblock Assessing the incidence and efficiency of a prominent place based policy.
\newblock \emph{American Economic Review}, 103(2), 897--947.
\bibitem[Butts and Gardner(2023)]{butts2023difference}
Butts, K. and Gardner, J. (2023).
\newblock Difference-in-differences with spatial spillovers.
\newblock Working paper.
\bibitem[Colella et al.(2019)]{colella2019inference}
Colella, F., Lalive, R., Sakalli, S. O., and Thoenig, M. (2019).
\newblock Inference with arbitrary clustering.
\newblock Working paper.
\bibitem[Conley(1999)]{conley1999gmm}
Conley, T. G. (1999).
\newblock GMM estimation with cross sectional dependence.
\newblock \emph{Journal of Econometrics}, 92(1), 1--45.
\bibitem[Deryugina et al.(2019)]{deryugina2019mortality}
Deryugina, T., Heutel, G., Miller, N. H., Molitor, D., and Reif, J. (2019).
\newblock The mortality and medical costs of air pollution: Evidence from changes in wind direction.
\newblock \emph{American Economic Review}, 109(12), 4178--4219.
\bibitem[Donaldson and Hornbeck(2016)]{donaldson2018railroads}
Donaldson, D. and Hornbeck, R. (2016).
\newblock Railroads and American economic growth: A market access approach.
\newblock \emph{Quarterly Journal of Economics}, 131(2), 799--858.
\bibitem[Duranton and Turner(2012)]{duranton2014roads}
Duranton, G. and Turner, M. A. (2012).
\newblock Urban growth and transportation.
\newblock \emph{Review of Economic Studies}, 79(4), 1407--1440.
\bibitem[Fan and Gijbels(1996)]{fan1996local}
Fan, J. and Gijbels, I. (1996).
\newblock \emph{Local Polynomial Modelling and Its Applications}.
\newblock Chapman \& Hall/CRC.
\bibitem[Fan and Gijbels(1992)]{fan1992variable}
Fan, J. and Gijbels, I. (1992).
\newblock Variable bandwidth and local linear regression smoothers.
\newblock \emph{Annals of Statistics}, 20(4), 2008--2036.
\bibitem[Jones et al.(1996)]{jones1996brief}
Jones, M. C., Marron, J. S., and Sheather, S. J. (1996).
\newblock A brief survey of bandwidth selection for density estimation.
\newblock \emph{Journal of the American Statistical Association}, 91(433), 401--407.
\bibitem[Kikuchi(2024a)]{kikuchi2024unified}
Kikuchi, T. (2024a).
\newblock A unified framework for spatial and temporal treatment effect boundaries: Theory and identification.
\newblock arXiv preprint arXiv:2510.00754.
\bibitem[Kikuchi(2024b)]{kikuchi2024stochastic}
Kikuchi, T. (2024b).
\newblock Stochastic boundaries in spatial general equilibrium: A diffusion-based approach to causal inference with spillover effects.
\newblock arXiv preprint arXiv:2508.06594.
\bibitem[Kikuchi(2024c)]{kikuchi2024navier}
Kikuchi, T. (2024c).
\newblock Spatial and temporal boundaries in difference-in-differences: A framework from Navier-Stokes equation.
\newblock arXiv preprint arXiv:2510.11013.
\bibitem[Kline and Moretti(2014)]{kline2019place}
Kline, P. and Moretti, E. (2014).
\newblock Local economic development, agglomeration economies, and the big push: 100 years of evidence from the Tennessee Valley Authority.
\newblock \emph{Quarterly Journal of Economics}, 129(1), 275--331.
\bibitem[Knittel et al.(2016)]{knittel2016caution}
Knittel, C. R., Miller, D. L., and Sanders, N. J. (2016).
\newblock Caution, drivers! Children present: Traffic, pollution, and infant health.
\newblock \emph{Review of Economics and Statistics}, 98(2), 350--366.
\bibitem[LeSage and Pace(2009)]{lesage2009introduction}
LeSage, J. and Pace, R. K. (2009).
\newblock \emph{Introduction to Spatial Econometrics}.
\newblock Chapman and Hall/CRC.
\bibitem[Müller and Song(2009)]{muller2009inference}
Müller, H.-G. and Song, K.-S. (2009).
\newblock Inference for density families using functional principal component analysis.
\newblock \emph{Journal of the American Statistical Association}, 104(487), 1169--1180.
\bibitem[Müller and Watson(2022)]{muller2022spatial}
Müller, U. K. and Watson, M. W. (2022).
\newblock Spatial correlation robust inference.
\newblock \emph{Econometrica}, 90(6), 2901--2935.
\bibitem[Müller and Watson(2024)]{muller2024spatial}
Müller, U. K. and Watson, M. W. (2024).
\newblock Spatial unit roots and spurious regression.
\newblock \emph{Econometrica}, 92(5), 1661--1695.
\bibitem[Muller and Mendelsohn(2009)]{muller2011damages}
Muller, N. Z. and Mendelsohn, R. (2009).
\newblock Efficient pollution regulation: Getting the prices right.
\newblock \emph{American Economic Review}, 99(5), 1714--1739.
\bibitem[Pasquill and Smith(1983)]{pasquill1976atmospheric}
Pasquill, F. and Smith, F. B. (1983).
\newblock \emph{Atmospheric Diffusion}, 3rd edition.
\newblock Ellis Horwood Limited.
\bibitem[Ruppert and Wand(1995)]{ruppert1995effective}
Ruppert, D. and Wand, M. P. (1995).
\newblock Multivariate locally weighted least squares regression.
\newblock \emph{Annals of Statistics}, 22(3), 1346--1370.
\bibitem[Seinfeld and Pandis(2016)]{seinfeld2016atmospheric}
Seinfeld, J. H. and Pandis, S. N. (2016).
\newblock \emph{Atmospheric Chemistry and Physics: From Air Pollution to Climate Change}, 3rd edition.
\newblock John Wiley \& Sons.
\bibitem[Venter et al.(2020)]{venter2020covid}
Venter, Z. S., Aunan, K., Chowdhury, S., and Lelieveld, J. (2020).
\newblock COVID-19 lockdowns cause global air pollution declines.
\newblock \emph{Proceedings of the National Academy of Sciences}, 117(32), 18984--18990.
\end{thebibliography}
\newpage