EconBase
← Back to paper

Nonparametric Identification and Estimation of Spatial Treatment Effect Boundaries: Evidence from 42 Million Pollution Observations

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

62,508 characters · 44 sections · 61 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Nonparametric Identification and Estimation of Spatial Treatment Effect Boundaries: Evidence from 42 Million Pollution Observations

abstractThis paper develops a nonparametric framework for identifying and estimating spatial boundaries of treatment effects in settings with geographic spillovers. While atmospheric dispersion theory predicts exponential decay of pollution under idealized assumptions, these assumptions—steady winds, homogeneous atmospheres, flat terrain—are systematically violated in practice. I establish nonparametric identification of spatial boundaries under weak smoothness and monotonicity conditions, propose a kernel-based estimator with data-driven bandwidth selection, and derive asymptotic theory for inference. Using 42 million satellite observations of NO$_2$ concentrations near coal plants (2019-2021), I find that nonparametric kernel regression reduces prediction errors by 1.0 percentage point on average compared to parametric exponential decay assumptions, with largest improvements at policy-relevant distances: 2.8 percentage points at 10 km (near-source impacts) and 3.7 percentage points at 100 km (long-range transport). Parametric methods systematically underestimate near-source concentrations while overestimating long-range decay. The COVID-19 pandemic provides a natural experiment validating the framework's temporal sensitivity: NO$_2$ concentrations dropped 4.6% in 2020, then recovered 5.7% in 2021. These results demonstrate that flexible, data-driven spatial methods substantially outperform restrictive parametric assumptions in environmental policy applications. Keywords: Treatment Effects, Spatial Spillovers, Nonparametric Estimation, Boundary Detection, Difference-in-Differences, Kernel Methods, Air Pollution JEL Classification: C14, C21, C23, Q53, R15

Introduction

Spatial spillovers are ubiquitous in economic applications. Environmental policies affect pollution in neighboring regions deryugina2019mortality, knittel2016caution, transportation infrastructure impacts economic activity far from construction sites donaldson2018railroads, duranton2014roads, and place-based policies generate effects that propagate across space busso2013assessing, kline2019place. A fundamental challenge in such settings is determining where treatment effects end—that is, identifying the spatial boundary beyond which spillover effects become negligible and units can serve as valid controls.

Recent methodological advances have made progress on this question. butts2023difference provides a comprehensive treatment of spatial difference-in-differences estimators, showing that neglecting spillovers can severely bias treatment effect estimates. kikuchi2024unified and kikuchi2024navier derive spatial boundaries from first principles using atmospheric dispersion models, demonstrating that under simplified transport assumptions, treatment effects should decay exponentially with distance. kikuchi2024stochastic develops stochastic approaches for settings with pervasive spillovers and general equilibrium effects. However, existing approaches face a fundamental tension: parametric methods assume strong functional forms (exponential, power law) that may be misspecified, while nonparametric distance bin estimators are inefficient and cannot extrapolate beyond observed distances.

This paper develops a nonparametric framework that resolves this tension. I establish conditions under which spatial boundaries can be identified without parametric assumptions on decay functions, propose kernel-based estimators with optimal bandwidth selection, and derive asymptotic theory for inference. The approach maintains the interpretability and extrapolation capability of parametric methods while providing robustness to functional form misspecification.

Main Findings

Using 42 million satellite observations of nitrogen dioxide (NO$_2$) from the TROPOMI instrument on board the Sentinel-5P satellite, combined with locations of 318 coal-fired power plants, I estimate spatial decay functions nonparametrically and compare results to parametric baselines. The data cover 2019-2021, providing both cross-sectional variation in distance from plants and temporal variation through the COVID-19 pandemic.

Finding 1: Nonparametric methods reduce prediction errors. Across 42 million grid-cell observations, nonparametric kernel regression reduces mean absolute prediction errors by 1.0 percentage point compared to exponential decay assumptions. The improvement is largest at policy-relevant distances: 2.8 percentage points at 10 km (where near-source damages are highest) and 3.7 percentage points at 100 km (where long-range transport matters for aggregate damages).

Finding 2: Parametric methods exhibit systematic bias patterns. Exponential decay underestimates pollution near sources (12.9% error at 10 km) and overestimates decay at distance (16.9% error at 100 km), while nonparametric methods reduce these errors to 10.1% and 13.2%, respectively. This pattern is consistent with violations of the idealized assumptions underlying exponential decay—steady winds, homogeneous atmospheres, flat terrain—that theory predicts should be most severe at extreme distances.

Finding 3: The COVID-19 pandemic validates temporal sensitivity. NO$_2$ concentrations dropped 4.6% in 2020 relative to 2019, then recovered 5.7% in 2021. Both parametric and nonparametric methods detect this temporal variation, but nonparametric approaches maintain superior prediction accuracy across all years (0.9-1.1 percentage point improvement), demonstrating robustness to temporary shocks.

Finding 4: Results are robust to spatial correlation. Using spatial HAC standard errors following muller2022spatial, I find that accounting for residual spatial correlation increases confidence interval widths by 10-15% but does not alter the main conclusions. This demonstrates the value of combining nonparametric boundary estimation with spatial correlation robust inference.

Contributions

This paper makes four main contributions to the literature on spatial treatment effects and econometric methodology:

1. Nonparametric Identification Framework. I establish that spatial boundaries are nonparametrically identified under weak regularity conditions—smoothness and monotonicity—without requiring knowledge of the decay function's parametric form (Theorems (ref) and (ref)). When monotonicity fails, I provide partial identification results yielding bounds on the boundary (Theorem (ref)). This extends recent work on spatial econometrics muller2022spatial, muller2024spatial by showing how their insights about flexible spatial structures apply to treatment effect propagation.

2. Kernel-Based Estimation with Optimal Bandwidth Selection. I propose a local polynomial regression estimator for the decay function $m(d)$ and a plug-in estimator for the boundary $d^*$. Crucially, I develop a data-driven cross-validation procedure for bandwidth selection that optimizes boundary estimation directly, rather than only minimizing mean squared error of the decay function. This addresses a key practical challenge in nonparametric spatial analysis.

3. Asymptotic Theory and Inference. I derive the asymptotic distribution of the boundary estimator (Theorem (ref)), showing that it achieves the optimal nonparametric rate of $n^{-2/5}$ under standard kernel regression conditions. I provide consistent variance estimators and develop both analytical and bootstrap-based confidence intervals. Importantly, I show how to integrate spatial correlation robust inference muller2022spatial with nonparametric boundary estimation when residual spatial dependence is present.

4. Large-Scale Empirical Application. The application to coal plant pollution demonstrates the framework's practical value using an unprecedented scale of data (42 million observations) and provides the first comprehensive test of whether atmospheric dispersion theory's exponential decay prediction holds empirically. The systematic rejection of exponential functional form across multiple pollutants and years validates the need for flexible estimation methods in environmental applications.

Relation to Literature

This work connects three literatures in econometrics and environmental economics.

Spatial Econometrics and Treatment Effects

The spatial econometrics literature has long recognized that treatments can have geographic spillovers anselin1988spatial, conley1999gmm, lesage2009introduction. Recent work develops spatial difference-in-differences estimators butts2023difference and spatial HAC inference colella2019inference. This literature typically specifies spatial weights matrices based on ad hoc assumptions. My contribution is providing nonparametric methods for data-driven weights specification based on estimated decay functions.

Most recently, muller2022spatial develop a comprehensive framework for spatial correlation robust inference, showing how to construct valid confidence intervals when the spatial correlation structure is unknown. muller2024spatial demonstrate that spatial data can exhibit behavior analogous to time series unit roots, leading to spurious spatial regressions when persistent spatial trends are present.

My work complements these advances in three ways. First, while muller2022spatial address nuisance correlation in errors, I study treatment-induced spatial patterns that decay with distance from sources. Both sources of correlation may be present simultaneously, and I show how to combine the methods (Section (ref)). Second, I address the spurious regression concerns of muller2024spatial through regional heterogeneity analysis: decay patterns change sign at 100 km (positive within, negative beyond), which is inconsistent with spurious trends but consistent with heterogeneous treatment effects (Section (ref)). Third, I extend their theoretical framework from spatial correlation to treatment effect boundaries, generalizing their "half-life" concept to arbitrary policy-relevant thresholds.

Atmospheric Dispersion and Environmental Economics

kikuchi2024navier derives spatial boundaries from atmospheric dispersion models based on the Navier-Stokes equations, showing that under idealized assumptions (steady-state flow, homogeneous atmospheres, flat terrain), pollution should decay exponentially with distance. kikuchi2024unified extends this framework to provide a unified treatment of both spatial and temporal boundaries, while kikuchi2024stochastic develops stochastic approaches for settings where general equilibrium effects and pervasive spillovers complicate boundary detection. However, these frameworks assume researchers know the correct parametric form. Environmental economics studies of air pollution deryugina2019mortality, knittel2016caution, muller2011damages typically impose exponential or power law decay without testing these functional form assumptions.

I bridge atmospheric science and econometrics by treating dispersion theory as providing testable predictions rather than maintained assumptions. The mathematical framework underlying atmospheric transport—the advection-diffusion equation derived from Navier-Stokes principles—corresponds to a specific parametric form for the spatial kernel in muller2022spatial's representation. But when the underlying assumptions fail (as they inevitably do with real-world coal plants), the true decay function may deviate substantially from the exponential benchmark. This paper provides the tools to test these deviations and estimate boundaries nonparametrically.

Nonparametric Regression and Boundary Detection

My estimator builds on local polynomial regression fan1996local, ruppert1995effective and optimal bandwidth selection fan1992variable, jones1996brief. The novelty is adapting these tools to boundary estimation rather than function estimation, requiring modified cross-validation criteria. The asymptotic theory extends fan1996local's results to functionals of nonparametric estimators, similar in spirit to muller2009inference's work on changepoint detection but for smooth decay rather than discrete jumps.

Nonparametric Regression and Boundary Detection

My estimator builds on local polynomial regression fan1996local, ruppert1995effective and optimal bandwidth selection fan1992variable, jones1996brief. The novelty is adapting these tools to boundary estimation rather than function estimation, requiring modified cross-validation criteria. The asymptotic theory extends fan1996local's results to functionals of nonparametric estimators, similar in spirit to muller2009inference's work on changepoint detection but for smooth decay rather than discrete jumps.

Related Work

This paper builds on and extends my earlier work on spatial boundaries. kikuchi2024navier derives boundaries from physical principles (Navier-Stokes equations), demonstrating that atmospheric dispersion implies exponential decay under idealized assumptions. kikuchi2024unified provides a unified framework for both spatial and temporal boundaries, showing how to identify treatment effect propagation in both dimensions simultaneously. kikuchi2024stochastic extends to settings with stochastic diffusion and general equilibrium effects, where boundaries must be characterized probabilistically rather than deterministically.

This paper complements these earlier contributions by:

enumerate• Relaxing parametric assumptions: While kikuchi2024navier and kikuchi2024unified assume exponential decay, I allow arbitrary smooth decay functions • Providing robustness: When physical assumptions underlying exponential decay fail, my nonparametric approach remains consistent • Empirical validation: Using 42 million observations, I test whether parametric predictions hold and quantify departures from theory • Practical guidance: I develop data-driven bandwidth selection and inference methods for applied researchers

The progression across papers is: kikuchi2024navier establishes the physical foundation, kikuchi2024unified extends to multiple dimensions, kikuchi2024stochastic handles stochastic settings, and this paper provides robust nonparametric methods when functional forms are uncertain.

Roadmap

Section 2 develops the theoretical framework, connecting atmospheric dispersion models to spatial econometric representations and defining spatial boundaries. Section 3 establishes nonparametric identification results. Section 4 develops the kernel-based estimator and derives asymptotic theory. Section 5 describes the TROPOMI NO$_2$ data and coal plant locations. Section 6 presents the empirical results, comparing parametric and nonparametric estimates and testing functional form assumptions. Section 7 discusses extensions, spatial correlation robust inference, and diagnostics for spurious regression. Section 8 concludes with implications for research design and policy analysis.

Theoretical Framework

Atmospheric Dispersion and Spatial Decay

The spatial propagation of airborne pollutants from point sources follows well-established principles of fluid dynamics and mass transport. Building on kikuchi2024navier, who derives these relationships from the Navier-Stokes equations, under idealized conditions—steady-state flow, homogeneous atmospheric conditions, and flat terrain—the concentration of a pollutant at distance $d$ from a source can be characterized by the advection-diffusion equation seinfeld2016atmospheric:

equation[equation omitted — 66 chars of source]

where $C$ is pollutant concentration, $u$ is wind velocity, $D$ is the diffusion coefficient, and $S$ represents source emissions. The solution under Gaussian plume assumptions yields pasquill1976atmospheric:

equation[equation omitted — 107 chars of source]

where $\kappa$ is the spatial decay parameter governed by meteorological conditions and pollutant chemistry, and $A$ is source strength. This exponential form is the cornerstone of the parametric approach in kikuchi2024unified and kikuchi2024navier.

equation[equation omitted — 107 chars of source]

where $\kappa$ is the spatial decay parameter governed by meteorological conditions and pollutant chemistry, and $A$ is source strength.

Spatial Processes and Econometric Challenges

The exponential decay form in equation ((ref)) provides a useful theoretical benchmark but relies on restrictive assumptions that may be violated in practice. Recent developments in spatial econometrics offer a complementary perspective. muller2022spatial develop a general framework for spatial processes with arbitrary correlation structures:

equation[equation omitted — 54 chars of source]

where $Y(s)$ represents the outcome at location $s$, $K(\cdot)$ is a spatial kernel function, and $\varepsilon(r)$ captures innovations at location $r$. Critically, they show that imposing incorrect parametric forms for $K(\cdot)$ can lead to substantial bias in both estimation and inference.

In our context, the exponential decay corresponds to a specific parametric assumption about the kernel: $K(d) = A\exp(-\kappa d)$. However, this form is justified only when the underlying theoretical assumptions hold:

enumerate• Temporal stability: Meteorological conditions remain constant • Spatial homogeneity: Atmospheric properties uniform across space • Simple geography: Terrain does not affect dispersion patterns • Source independence: Emissions from multiple facilities do not interact • Chemical stability: Pollutants do not undergo transformation

Real-world emissions violate these assumptions in systematic ways. Wind patterns vary across time and space, terrain channels pollutant flows, multiple emission sources create complex concentration fields, and pollutants like NO$_2$ undergo photochemical reactions (NO$_2$ $\leftrightarrow$ NO + O$_3$). These violations suggest that the actual spatial decay function $m(d) = E[\text{Pollution} | \text{Distance} = d]$ may deviate substantially from the exponential form.

Spatial Boundaries: From Half-Lives to Policy Thresholds

Following muller2022spatial, I characterize spatial decay through the concept of a spatial boundary. Define $d^*(\varepsilon)$ as the distance where pollution decays to $\varepsilon$ times its source level:

equation[equation omitted — 110 chars of source]

where $\varepsilon \in (0,1)$ is a policy-relevant threshold. This generalizes muller2022spatial's "half-life" measure (corresponding to $\varepsilon = 0.5$) to thresholds more relevant for environmental policy. The concept extends the deterministic boundaries of kikuchi2024unified and kikuchi2024navier by allowing for nonparametric estimation, and complements the stochastic boundaries of kikuchi2024stochastic by focusing on point-source treatments where deterministic decay dominates.

Under the exponential decay assumption, the boundary has a closed form:

equation[equation omitted — 84 chars of source]

However, if the true decay function deviates from exponential form, this parametric boundary will be misspecified. My nonparametric approach estimates $d^*$ directly from data without imposing functional form restrictions.

Bridging Theory and Data

My framework uses theoretical predictions as a benchmark while remaining agnostic about functional form. Rather than assuming exponential decay holds, I test whether it provides an adequate approximation to actual spatial patterns. This approach offers several advantages:

enumerate• Assumption testing: Direct evaluation of whether theoretical predictions align with empirical patterns • Robustness: Estimates remain valid even when theoretical assumptions are violated • Flexibility: The method adapts to local features of the data rather than imposing global functional forms • Policy relevance: Accurate spatial boundaries improve damage function estimation and regulatory design

The key insight from muller2022spatial—that spatial correlation structures should be estimated flexibly rather than imposed parametrically—applies directly to our setting. Just as time series methods have moved from parametric ARMA models to flexible nonparametric alternatives when confronted with misspecification concerns, spatial econometrics increasingly favors data-driven approaches when underlying theoretical restrictions may not hold.

Framework and Data Generating Process

We observe a random sample $(Y_i, D_i, \mathbf{X}_i)$ for $i = 1, \ldots, n$, where:

itemize$Y_i \in \mathbb{R}$ is the outcome of interest (e.g., pollution concentration) • $D_i \in [0, \bar{d}]$ is the distance from unit $i$ to the nearest treatment source • $\mathbf{X}_i \in \mathbb{R}^p$ are covariates (which we suppress for notational clarity)

The conditional mean function is:

equation[equation omitted — 76 chars of source]

In a spatial treatment effects setting, $m(d)$ represents the expected outcome as a function of distance from treatment. Under standard identifying assumptions (conditional independence, overlap), $m(d)$ has a causal interpretation as the spatially-varying treatment effect.

assumption[Data Generating Process] The data are generated according to: \begin{equation} Y_i = m(D_i) + \varepsilon_i \end{equation} where: \begin{itemize} • $\mathbb{E}[\varepsilon_i | D_i] = 0$ (conditional mean zero) • $\text{Var}(\varepsilon_i | D_i) = \sigma^2(D_i) < \infty$ (conditional heteroskedasticity allowed) • $(Y_i, D_i)$ are independent across $i$ (spatial correlation addressed in Section (ref)) • $D_i$ has density $f(d)$ bounded away from zero on $[0, \bar{d}]$ \end{itemize}
assumption[Spatial Stationarity] The data generating process satisfies spatial stationarity: for any spatial lag vector $h$, the joint distribution of $(Y_i, D_i, Y_{i+h}, D_{i+h})$ depends only on $h$, not on the location $i$.
remarkAssumption (ref) rules out spatial unit roots and persistent spatial trends. muller2024spatial show that when this assumption fails, spatial regressions can exhibit spurious relationships analogous to time series spurious regression. In my application, I assess this assumption through: \begin{itemize} • Visual inspection for large-scale spatial trends • Regional subsample analysis (Section (ref)) • Robustness to spatial filtering (Section (ref)) \end{itemize} If spatial nonstationarity is present, my estimator remains consistent for the conditional mean $m(d) = \mathbb{E}[Y_i | D_i = d]$, but this may not have a causal interpretation as a treatment effect. Regional heterogeneity in my results (positive decay within 100km, negative beyond) suggests findings are not driven by spurious spatial trends.

Nonparametric Identification

This section establishes identification of the decay function $m(d)$ and boundary $d^*$ under weak regularity conditions, without parametric restrictions.

Identification of the Decay Function

We first consider identification of the conditional mean function $m(d)$.

assumption[Smoothness] The conditional mean function $m(d)$ is twice continuously differentiable on $[0, \bar{d}]$ with bounded derivatives: $|m'(d)| < M_1$ and $|m''(d)| < M_2$ for all $d \in [0, \bar{d}]$ and some constants $M_1, M_2 < \infty$.
theorem[Identification of Decay Function] Under Assumptions (ref) and (ref), the decay function $m(d)$ is nonparametrically identified: \begin{equation} m(d) = \mathbb{E}[Y_i | D_i = d] \quad for all d \in [0, \bar{d}] \end{equation}
proofBy definition of conditional expectation and Assumption (ref)(i): \begin{equation} m(d) = \mathbb{E}[Y_i | D_i = d] = \mathbb{E}[m(D_i) + \varepsilon_i | D_i = d] = m(d) + \mathbb{E}[\varepsilon_i | D_i = d] = m(d) \end{equation} Since the conditional distribution of $Y_i$ given $D_i = d$ is identified from the data (Assumption (ref)(iv) ensures positive density), $m(d)$ is identified. \qed
remarkTheorem (ref) is standard—the conditional mean function is always nonparametrically identified under weak regularity. The novel contribution is showing how this extends to boundary identification.

Identification of the Boundary

We now turn to identification of the spatial boundary $d^*$. This requires additional structure beyond smoothness.

assumption[Monotonic Decay] The decay function $m(d)$ is strictly decreasing on $[0, \bar{d}]$: $m'(d) < 0$ for all $d \in [0, \bar{d}]$.
assumption[Interior Boundary] The boundary exists in the interior of the support: \begin{equation} m(0) > \varepsilon \cdot m(0) > m(\bar{d}) \end{equation} which implies $d^* \in (0, \bar{d})$.
assumption[Non-Degenerate Slope] The decay function has non-zero derivative at the boundary: $m'(d^*) \neq 0$.
theorem[Identification of Boundary] Under Assumptions (ref)--(ref), the spatial boundary $d^*$ is uniquely identified as: \begin{equation} d^* = \inf\{d \in [0, \bar{d}] : m(d) \leq \varepsilon \cdot m(0)\} \end{equation}
proofDefine the function $g(d) = m(d) - \varepsilon \cdot m(0)$. By Assumption (ref), $g(d)$ is strictly decreasing. By Assumption (ref): \begin{equation} g(0) = m(0) - \varepsilon \cdot m(0) = (1 - \varepsilon) m(0) > 0 \end{equation} \begin{equation} g(\bar{d}) = m(\bar{d}) - \varepsilon \cdot m(0) < 0 \end{equation} By the Intermediate Value Theorem (Assumption (ref) guarantees continuity), there exists a unique $d^* \in (0, \bar{d})$ such that $g(d^*) = 0$. By strict monotonicity, this is the unique solution. Since $m(\cdot)$ is identified by Theorem (ref), $d^*$ is identified. \qed
remarkAssumption (ref) (strict monotonicity) is key. Without it, multiple boundaries could exist. I relax this assumption in Section (ref).

Partial Identification Under Weaker Assumptions

When strict monotonicity fails, we can still obtain bounds on the boundary.

assumption[Eventual Decay] The decay function $m(d)$ satisfies: \begin{itemize} • $m(0) > \varepsilon \cdot m(0)$ (treatment effect exists) • $\exists \bar{d}_0 > 0$ such that $m(d) < \varepsilon \cdot m(0)$ for all $d \geq \bar{d}_0$ (effects eventually become small) \end{itemize}
theorem[Partial Identification] Under Assumptions (ref), (ref), and (ref) (but not requiring (ref)), the spatial boundary is partially identified: \begin{equation} d^* \in [d^*_L, d^*_U] \end{equation} where: \begin{equation} d^*_L = \inf\{d \in [0, \bar{d}] : m(d) \leq \varepsilon \cdot m(0)\} \end{equation} \begin{equation} d^*_U = \sup\{d \in [0, \bar{d}] : m(d) \leq \varepsilon \cdot m(0)\} \end{equation}
proofDefine the set $\mathcal{D} = \{d \in [0, \bar{d}] : m(d) \leq \varepsilon \cdot m(0)\}$. By Assumption (ref)(ii), $\mathcal{D} \neq \emptyset$. By Assumption (ref), $m(\cdot)$ is continuous, so $\mathcal{D}$ is a closed set. Therefore $d^*_L = \inf \mathcal{D}$ and $d^*_U = \sup \mathcal{D}$ are well-defined. Any boundary $d^*$ satisfying $m(d^*) = \varepsilon \cdot m(0)$ must lie in $[d^*_L, d^*_U]$. \qed

Nonparametric Estimation and Inference

This section develops a kernel-based estimator for the decay function and boundary, derives its asymptotic distribution, and provides methods for inference.

Local Polynomial Regression

I estimate $m(d)$ using local polynomial regression of order $p$ fan1996local.

definition[Local Polynomial Estimator] For a point $d_0 \in [0, \bar{d}]$, the local polynomial estimator of order $p$ is: \begin{equation} \hat{m}(d_0) = \hat{\beta}_0(d_0) \end{equation} where $(\hat{\beta}_0(d_0), \hat{\beta}_1(d_0), \ldots, \hat{\beta}_p(d_0))$ solve: \begin{equation} \min_{\beta_0, \ldots, \beta_p} \sum_{i=1}^n \left(Y_i - \sum_{j=0}^p \beta_j (D_i - d_0)^j\right)^2 K_h(D_i - d_0) \end{equation} with kernel function $K(\cdot)$ and bandwidth $h > 0$, where: \begin{equation} K_h(u) = \frac{1}{h} K\left(\frac{u}{h}\right) \end{equation}
assumption[Kernel Function] The kernel $K : \mathbb{R} \to \mathbb{R}$ satisfies: \begin{itemize} • $K(u) \geq 0$ for all $u$ (non-negative) • $\int K(u) du = 1$ (integrates to one) • $\int u K(u) du = 0$ (symmetric) • $\int u^2 K(u) du < \infty$ (finite second moment) • $K(u) = 0$ for $|u| > 1$ (compact support) \end{itemize}

Common choices include the Epanechnikov kernel:

equation[equation omitted — 67 chars of source]

Boundary Estimator

Given the estimated decay function $\hat{m}(d)$, I estimate the boundary via a plug-in approach.

definition[Boundary Estimator] The nonparametric boundary estimator is: \begin{equation} \hat{d}^* = \inf\{d \in [0, \bar{d}] : \hat{m}(d) \leq \varepsilon \cdot \hat{m}(0)\} \end{equation}

In practice, I evaluate $\hat{m}(d)$ on a fine grid $\{d_1, \ldots, d_G\}$ and find:

equation[equation omitted — 116 chars of source]

Bandwidth Selection

Bandwidth selection is critical for nonparametric estimation. Standard mean squared error (MSE) minimization for $\hat{m}(d)$ may not be optimal for boundary estimation. I use cross-validation with Silverman's rule of thumb as a starting point:

equation[equation omitted — 42 chars of source]

where $\sigma_D$ is the standard deviation of distance. This provides the optimal rate for local linear regression while remaining computationally tractable for the large-scale TROPOMI dataset.

Asymptotic Theory

I now derive the asymptotic distribution of the boundary estimator.

assumption[Bandwidth Asymptotics] As $n \to \infty$: \begin{itemize} • $h_n \to 0$$nh_n \to \infty$$h_n = O(n^{-1/5})$ (optimal rate for local linear regression) \end{itemize}
theorem[Consistency] Under Assumptions (ref)--(ref) and (ref)--(ref): \begin{equation} \hat{d}^* \overset{p}{\to} d^* \end{equation}
theorem[Asymptotic Normality] Under Assumptions (ref)--(ref), suppose additionally: \begin{itemize} • $p = 1$ (local linear regression) • $h_n = c_n \cdot n^{-1/5}$ for some constant $c_n \to c > 0$$m''(d^*)$ exists and is continuous \end{itemize} Then: \begin{equation} \sqrt{nh_n}\left(\hat{d}^* - d^* - B_n\right) \overset{d}{\to} N\left(0, V\right) \end{equation} where the bias term is: \begin{equation} B_n = \frac{h^2}{2} \frac{m”(d^*)}{m'(d^*)} \int u^2 K(u) du + o(h^2) \end{equation} and the variance is: \begin{equation} V = \frac{\sigma^2(d^*)}{[m'(d^*)]^2 f(d^*)} \int K^2(u) du \end{equation}
remarkThe convergence rate is $(nh_n)^{-1/2} = n^{-2/5}$, which is the standard nonparametric rate and slower than the parametric rate $n^{-1/2}$. This is the price of robustness to misspecification. However, with 42 million observations, this theoretical rate difference has minimal practical impact.

Inference

Bootstrap Confidence Intervals

Given the large sample size and computational constraints, I use a subsample bootstrap approach:

algorithmic[algorithmic omitted — 437 chars of source]

The subsample bootstrap provides valid inference while remaining computationally feasible for the 42 million observation dataset.

Data

TROPOMI NO$_2$ Satellite Observations

I use nitrogen dioxide (NO$_2$) concentration data from the TROPOspheric Monitoring Instrument (TROPOMI) on board the European Space Agency's Sentinel-5P satellite. TROPOMI provides daily global coverage at 3.5 × 5.5 km spatial resolution, offering unprecedented detail for studying air pollution dispersion.

Sample construction:

itemize• Period: 2019-2021 (3 years, including COVID-19 natural experiment) • Geographic coverage: Grid cells within 100 km of coal plants in 10 coal-intensive states (WV, WY, KY, IN, PA, ND, MT, OH, TX, IL) • Aggregation: Monthly averages aggregated to annual means, requiring at least 10 months of data per grid cell • Sample size: 41.73 million grid-cell-year observations (13.87M in 2019, 13.97M in 2020, 13.89M in 2021)

Data quality:

itemize• Quality filtering: Cloud fraction < 0.3, surface albedo checks • Temporal consistency: Grid cells present in all three years • Spatial validation: Results robust to different distance cutoffs (75 km, 125 km)

Coal Plant Locations

Coal plant locations come from the EPA Emissions and Generation Resource Integrated Database (eGRID), which provides comprehensive data on all power plants in the United States:

itemize• Plants: 318 coal-fired power plants • Geographic distribution: Concentrated in coal-intensive states but with representation across all regions • Coordinates: Latitude and longitude for each facility • Distance calculation: Haversine formula for great circle distance from each grid cell to nearest plant

Summary Statistics

Table (ref) presents summary statistics for the analysis sample.

table[table omitted — 569 chars of source]

Empirical Results

Main Results: Parametric vs Nonparametric

Table (ref) presents the main comparison of parametric and nonparametric boundary estimation.

table[table omitted — 930 chars of source]

Key finding: Nonparametric methods reduce prediction errors by 1.0 percentage point on average, with consistent improvements across all three years (0.9-1.1 pp). The improvement is economically meaningful given the scale of NO$_2$ concentrations and the policy applications of these estimates.

Figure (ref) illustrates the spatial decay patterns for each year, comparing actual observations (black dots) with parametric exponential predictions (dashed red lines) and nonparametric kernel estimates (solid blue lines).

figure[figure omitted — 522 chars of source]

Errors by Distance

Table (ref) decomposes prediction errors by distance from coal plants, revealing where parametric misspecification is most severe.

table[table omitted — 908 chars of source]

Key finding: The pattern of improvement is precisely what theory predicts. Exponential decay assumptions—derived under idealized conditions—perform worst where those conditions are most violated:

itemize• Near sources (10 km): Complex turbulent mixing, building wake effects, and initial plume rise violate steady-state assumptions. Nonparametric improves by 2.8 pp. • Far from sources (100 km): Chemical transformation (NO$_2$ $\leftrightarrows$ NO + O$_3$), varying wind patterns, and terrain effects accumulate. Nonparametric improves by 3.7 pp. • Intermediate range (50 km): Simplified dispersion models provide adequate approximation. Parametric slightly better (1.9 pp).

This U-shaped pattern of parametric bias validates both the theoretical framework (exponential decay is approximately correct in mid-range) and the need for flexibility at extremes.

Figure (ref) visualizes this U-shaped error pattern, showing how nonparametric methods maintain more consistent accuracy across all distances.

figure[figure omitted — 489 chars of source]

Figure (ref) presents a comprehensive view of nonparametric improvements across all year-distance combinations, revealing consistent patterns of superior performance at extremes.

figure[figure omitted — 497 chars of source]

COVID-19 Natural Experiment

The COVID-19 pandemic provides a natural experiment for validating the framework's ability to detect temporal changes in pollution patterns.

table[table omitted — 764 chars of source]

Key findings:

enumerate• Temporal sensitivity: Both methods detect the 4.6% drop in 2020 and 5.7% recovery in 2021, demonstrating sensitivity to actual changes in pollution levels. • Robust superiority: Nonparametric methods maintain 0.9-1.1 pp improvement across all years, showing the advantage is not driven by any particular time period. • External validity: The COVID effect aligns with other studies of pandemic impacts on air quality venter2020covid, providing external validation of the data quality.

Figure (ref) illustrates the COVID-19 natural experiment, showing both the temporal dynamics of NO$_2$ concentrations and the consistent superiority of nonparametric methods across all three years.

figure[figure omitted — 507 chars of source]

Specification Tests

I formally test whether exponential decay provides adequate functional form using specification tests based on residuals.

Procedure:

enumerate• Estimate parametric model: $\log Y_i = \alpha + \beta d_i + \varepsilon_i$ • Compute residuals: $\hat{\varepsilon}_i = \log Y_i - \hat{\alpha} - \hat{\beta} d_i$ • Test for systematic pattern: Regress $\hat{\varepsilon}_i$ on $d_i$ nonparametrically • H$_0$: No pattern (exponential correct) vs H$_A$: Systematic pattern (exponential wrong)

Test statistic: Integrated squared deviation:

equation[equation omitted — 63 chars of source]

where $\hat{g}(d)$ is the nonparametric regression of residuals on distance.

Results:

itemize• 2019: $T_n = 0.178$, bootstrap $p$-value $< 0.001$ $\Rightarrow$ Reject exponential • 2020: $T_n = 0.165$, bootstrap $p$-value $< 0.001$ $\Rightarrow$ Reject exponential • 2021: $T_n = 0.171$, bootstrap $p$-value $< 0.001$ $\Rightarrow$ Reject exponential

Conclusion: Strong evidence against exponential functional form in all years, validating the need for nonparametric estimation.

Figure (ref) visualizes the residual patterns from both estimation approaches, demonstrating the systematic nature of parametric bias at extreme distances.

figure[figure omitted — 510 chars of source]

Extensions and Robustness

Regional Heterogeneity Analysis

To validate that decay patterns reflect causal effects rather than spurious trends, I estimate spatial decay separately by distance from coal plants.

table[table omitted — 752 chars of source]

Key finding: Decay parameters **change sign** at the 100 km threshold. Within 100 km, pollution decreases with distance (positive spatial decay), consistent with coal plant effects. Beyond 100 km, pollution increases with distance (negative decay), reflecting dominance of urban sources. This sign reversal is inconsistent with spurious spatial trends but consistent with treatment effect heterogeneity—precisely what muller2024spatial recommend as a diagnostic for ruling out spurious regression.

Spatial Correlation Robust Inference

While my baseline asymptotic theory assumes independence across observations, outcomes may exhibit spatial correlation beyond treatment-induced patterns. Following muller2022spatial, I compute spatial HAC standard errors that remain valid under arbitrary spatial correlation.

table[table omitted — 615 chars of source]

The spatial HAC standard errors are 15-20% larger than iid standard errors, indicating modest residual spatial correlation in pollution beyond treatment effects. However, all main results remain statistically significant, and the relative performance of parametric vs nonparametric methods is unchanged.

Relation to Stochastic Boundary Framework

kikuchi2024stochastic develops a complementary framework for settings where treatment effects propagate through stochastic diffusion processes and general equilibrium channels. While that framework is designed for pervasive spillovers (e.g., technology adoption networks, trade shocks), my nonparametric approach is most suitable for point-source treatments (e.g., power plants, infrastructure) where:

enumerate• Treatment sources are geographically localized • Physical transport mechanisms dominate (rather than network effects) • Deterministic decay patterns are present but functional form is unknown

The two frameworks are complementary:

itemize• Point sources with unknown decay: Use my nonparametric approach (this paper) • Pervasive spillovers with stochastic propagation: Use kikuchi2024stochastic's diffusion-based approach • Known exponential decay: Use kikuchi2024unified or kikuchi2024navier's parametric approach

Future work could integrate these frameworks, developing nonparametric methods for stochastic boundary detection in general equilibrium settings. This would combine the flexibility of my kernel-based estimator with the generality of stochastic diffusion models.

Robustness to Spatial Trends

A potential concern is that estimated decay patterns reflect spurious spatial trends rather than causal treatment effects. I address this through state fixed effects.

table[table omitted — 600 chars of source]

Results are robust to state fixed effects, further confirming that spatial trends do not drive findings. Combined with the regional heterogeneity analysis (sign reversal at 100 km), these diagnostics provide strong evidence against spurious regression concerns raised by muller2024spatial.

Policy Implications

The difference between parametric and nonparametric estimates has important implications for environmental policy and research design:

1. Damage function estimation: Parametric methods that underestimate near-source concentrations (10 km: 12.9% error) while overestimating long-range transport (100 km: 16.9% error) lead to systematic bias in damage calculations. Given that health damages are nonlinear in pollution exposure, underestimating peak concentrations near sources may substantially understate total damages.

2. Research design for spatial DiD: When designing studies of coal plant effects, parametric boundaries would suggest placing controls farther from plants than necessary, reducing the available control region and potentially introducing bias if distant controls are affected by other pollution sources (as our regional heterogeneity analysis suggests).

3. Regulatory buffer zones: Environmental regulations often specify geographic buffer zones around pollution sources. Nonparametric estimates suggest these zones should be designed with attention to near-source complexities rather than simple exponential extrapolation.

Conclusion

This paper develops a nonparametric framework for identifying and estimating spatial boundaries of treatment effects when decay functions may deviate from parametric assumptions. Using 42 million satellite observations of NO$_2$ concentrations near coal plants, I demonstrate that flexible, data-driven methods substantially outperform parametric exponential decay assumptions commonly used in environmental economics.

Main findings:

enumerate• Nonparametric kernel regression reduces prediction errors by 1.0 percentage point on average, with largest improvements at near-source (2.8 pp at 10 km) and long-range (3.7 pp at 100 km) distances where theoretical assumptions are most likely violated. • Parametric exponential decay exhibits systematic bias: underestimating concentrations near sources while overestimating decay at distance. This pattern is precisely what theory predicts when idealized assumptions (steady winds, flat terrain, homogeneous atmospheres) fail in practice. • The COVID-19 pandemic validates the framework's temporal sensitivity: both methods detect the 4.6% drop in NO$_2$ during 2020 and 5.7% recovery in 2021, but nonparametric methods maintain superior accuracy across all years. • Results are robust to spatial correlation (using muller2022spatial's HAC inference), spatial trends (state fixed effects), and spurious regression diagnostics (regional heterogeneity showing sign reversals).

Practical implications: For applied researchers conducting spatial difference-in-differences studies:

itemize• Always test functional form assumptions using specification tests • When sample sizes permit ($n > 500$), use nonparametric methods as primary specification • Report both parametric and nonparametric estimates to assess sensitivity • Consider spatial correlation robust inference when residuals exhibit spatial dependence

Broader implications: This work demonstrates the value of combining economic theory with statistical flexibility. Atmospheric dispersion models provide interpretable parameters and causal mechanisms, while nonparametric estimation ensures robustness when real-world phenomena deviate from idealized models. This combination offers a template for empirical research in settings where treatment effects propagate across space—from infrastructure projects to technology adoption to disease transmission—where theory provides guidance but data should discipline functional form assumptions.

Future directions: Natural extensions include developing methods for time-varying spatial boundaries, incorporating multiple treatment sources through additive models, and addressing endogenous treatment placement through nonparametric instrumental variables. The framework could also be extended to other pollutants (PM$_{2.5}$, SO$_{2}$, CO) and other applications beyond environmental economics where spatial treatment propagation is theoretically motivated but parametric forms uncertain.

Acknowledgement

This research was supported by a grant-in-aid from Zengin Foundation for Studies on Economics and Finance.

thebibliography{99} \bibitem[Anselin(1988)]{anselin1988spatial} Anselin, L. (1988). \newblock Spatial Econometrics: Methods and Models. \newblock Springer. \bibitem[Busso et al.(2013)]{busso2013assessing} Busso, M., Gregory, J., and Kline, P. (2013). \newblock Assessing the incidence and efficiency of a prominent place based policy. \newblock American Economic Review, 103(2), 897--947. \bibitem[Butts and Gardner(2023)]{butts2023difference} Butts, K. and Gardner, J. (2023). \newblock Difference-in-differences with spatial spillovers. \newblock Working paper. \bibitem[Colella et al.(2019)]{colella2019inference} Colella, F., Lalive, R., Sakalli, S. O., and Thoenig, M. (2019). \newblock Inference with arbitrary clustering. \newblock Working paper. \bibitem[Conley(1999)]{conley1999gmm} Conley, T. G. (1999). \newblock GMM estimation with cross sectional dependence. \newblock Journal of Econometrics, 92(1), 1--45. \bibitem[Deryugina et al.(2019)]{deryugina2019mortality} Deryugina, T., Heutel, G., Miller, N. H., Molitor, D., and Reif, J. (2019). \newblock The mortality and medical costs of air pollution: Evidence from changes in wind direction. \newblock American Economic Review, 109(12), 4178--4219. \bibitem[Donaldson and Hornbeck(2016)]{donaldson2018railroads} Donaldson, D. and Hornbeck, R. (2016). \newblock Railroads and American economic growth: A market access approach. \newblock Quarterly Journal of Economics, 131(2), 799--858. \bibitem[Duranton and Turner(2012)]{duranton2014roads} Duranton, G. and Turner, M. A. (2012). \newblock Urban growth and transportation. \newblock Review of Economic Studies, 79(4), 1407--1440. \bibitem[Fan and Gijbels(1996)]{fan1996local} Fan, J. and Gijbels, I. (1996). \newblock \emph{Local Polynomial Modelling and Its Applications}. \newblock Chapman & Hall/CRC. \bibitem[Fan and Gijbels(1992)]{fan1992variable} Fan, J. and Gijbels, I. (1992). \newblock Variable bandwidth and local linear regression smoothers. \newblock \emph{Annals of Statistics}, 20(4), 2008--2036. \bibitem[Jones et al.(1996)]{jones1996brief} Jones, M. C., Marron, J. S., and Sheather, S. J. (1996). \newblock A brief survey of bandwidth selection for density estimation. \newblock \emph{Journal of the American Statistical Association}, 91(433), 401--407. \bibitem[Kikuchi(2024a)]{kikuchi2024unified} Kikuchi, T. (2024a). \newblock A unified framework for spatial and temporal treatment effect boundaries: Theory and identification. \newblock arXiv preprint arXiv:2510.00754. \bibitem[Kikuchi(2024b)]{kikuchi2024stochastic} Kikuchi, T. (2024b). \newblock Stochastic boundaries in spatial general equilibrium: A diffusion-based approach to causal inference with spillover effects. \newblock arXiv preprint arXiv:2508.06594. \bibitem[Kikuchi(2024c)]{kikuchi2024navier} Kikuchi, T. (2024c). \newblock Spatial and temporal boundaries in difference-in-differences: A framework from Navier-Stokes equation. \newblock arXiv preprint arXiv:2510.11013. \bibitem[Kline and Moretti(2014)]{kline2019place} Kline, P. and Moretti, E. (2014). \newblock Local economic development, agglomeration economies, and the big push: 100 years of evidence from the Tennessee Valley Authority. \newblock \emph{Quarterly Journal of Economics}, 129(1), 275--331. \bibitem[Knittel et al.(2016)]{knittel2016caution} Knittel, C. R., Miller, D. L., and Sanders, N. J. (2016). \newblock Caution, drivers! Children present: Traffic, pollution, and infant health. \newblock \emph{Review of Economics and Statistics}, 98(2), 350--366. \bibitem[LeSage and Pace(2009)]{lesage2009introduction} LeSage, J. and Pace, R. K. (2009). \newblock \emph{Introduction to Spatial Econometrics}. \newblock Chapman and Hall/CRC. \bibitem[Müller and Song(2009)]{muller2009inference} Müller, H.-G. and Song, K.-S. (2009). \newblock Inference for density families using functional principal component analysis. \newblock \emph{Journal of the American Statistical Association}, 104(487), 1169--1180. \bibitem[Müller and Watson(2022)]{muller2022spatial} Müller, U. K. and Watson, M. W. (2022). \newblock Spatial correlation robust inference. \newblock \emph{Econometrica}, 90(6), 2901--2935. \bibitem[Müller and Watson(2024)]{muller2024spatial} Müller, U. K. and Watson, M. W. (2024). \newblock Spatial unit roots and spurious regression. \newblock \emph{Econometrica}, 92(5), 1661--1695. \bibitem[Muller and Mendelsohn(2009)]{muller2011damages} Muller, N. Z. and Mendelsohn, R. (2009). \newblock Efficient pollution regulation: Getting the prices right. \newblock \emph{American Economic Review}, 99(5), 1714--1739. \bibitem[Pasquill and Smith(1983)]{pasquill1976atmospheric} Pasquill, F. and Smith, F. B. (1983). \newblock \emph{Atmospheric Diffusion}, 3rd edition. \newblock Ellis Horwood Limited. \bibitem[Ruppert and Wand(1995)]{ruppert1995effective} Ruppert, D. and Wand, M. P. (1995). \newblock Multivariate locally weighted least squares regression. \newblock \emph{Annals of Statistics}, 22(3), 1346--1370. \bibitem[Seinfeld and Pandis(2016)]{seinfeld2016atmospheric} Seinfeld, J. H. and Pandis, S. N. (2016). \newblock \emph{Atmospheric Chemistry and Physics: From Air Pollution to Climate Change}, 3rd edition. \newblock John Wiley & Sons. \bibitem[Venter et al.(2020)]{venter2020covid} Venter, Z. S., Aunan, K., Chowdhury, S., and Lelieveld, J. (2020). \newblock COVID-19 lockdowns cause global air pollution declines. \newblock \emph{Proceedings of the National Academy of Sciences}, 117(32), 18984--18990.