Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
19,131 characters · 6 sections · 28 citation commands
Is Productivity Advantage of Cities Really Down To Mean and Variance?
Firms in denser areas are systematically more productive Melo2009, Combes2015EmpiricsAgglomerationEconomies, Duranton2020EconomicsUrbanDensity. This productivity advantage may reflect agglomeration economies or firm selection --- the concentration of firms in dense environments making them more productive, or stronger competition in denser areas culling weaker firms at higher rates, respectively. Understanding which of these mechanisms drives the productivity advantage of cities is important for evaluating place-based interventions like tax incentives, infrastructure investments, or regional development programs Neumark2015PlaceBasedPolicies.
Combes2012b disentangle these channels by assuming that the distributions of total factor productivity (TFP) across denser and less dense areas are the same except for shifts in mean and variance that reflect agglomeration and for tail truncation due to firm selection. Under this assumption, they decompose productivity differences along the two channels. If, however, the TFP distributions across areas differed in other ways (e.g. tail behavior or shape), this decomposition may be invalid and instead reflect those differences, complicating interpretations of sorting patterns Gaubert2018FirmSortingAgglomeration and evaluations of place‑based interventions.
Despite its central role, this assumption has not been directly tested. Empirical validation is challenging because productivity is not directly observed, but instead estimated with noise. Thus, naive comparisons of TFP distributions based on noisy TFP estimates may conflate (policy-irrelevant) noise with true heterogeneity Jochmans2019.
This paper empirically tests this assumption of distribution equality. Using Spanish administrative firm-level data, we nonparametrically test for differences in TFP distributions across denser and less dense areas, using methods that account for estimation noise in TFP. Across all sectors, we find that TFP distributions share the same shape, with differences fully explained by the parameters posited by Combes2012b.
Our results suggest that urban productivity advantages may be fully explained by agglomeration rather than stronger selection. Of the three Combes2012b parameters, only the mean and variance are needed to align the TFP distributions across areas, while the tail truncation parameter is not needed. This suggests that policies enhancing agglomeration, such as improving infrastructure or fostering knowledge spillovers, may be more effective than competition‑enhancing policies, such as reducing entry barriers or promoting firm turnover.
This paper makes three contributions. First, we empirically validate the distribution equality assumption in Combes2012b, complementing the partial results of Morozov2026InferenceExtremeQuantiles, who finds no evidence of left-tail truncation. Second, our methodology for testing distribution equality with noisy data is also applicable to other contexts, such as geographical differences in worker skills Roca2017. Last, we contribute to the literature on working with noisy estimates of unobserved heterogeneity Okui2019b, Jochmans2019, Morozov2026InferenceExtremeQuantiles by proposing a goodness-of-fit test based on noisy observations.
To evaluate the assumption of distribution equality for TFPs, we work in the setting of Combes2012b (CDGPR12 henceforth). Their approach proceeds in two steps: first estimating firm-specific TFPs, then comparing productivity distributions between high-density (above-median density, AMD) and low-density (below-median density, BMD) areas.
Following CDGPR12, we define firm-level TFP $\theta_i$ by assuming that firm $i$ in year $t$, sector $s$, and area $a\in\curl{\text{AMD}, \text{BMD}}$ produces value added $V_{i,t}$ according to a Cobb-Douglas production function:
where $\theta_i$ is firm (log) total factor productivity (TFP), $K_{it}$ and $L_{it}$ are capital and labor, and $U_{it}$ captures measurement error in $V_{it}$. Here $\beta_{1, s, a}$ and $\beta_{2, s, a}$ are sector- and area-specific factor shares, while $\beta_{0, t, s, a}$ is a time-specific intercept, allowing for productivity trends. Firms in sector $s$ draw $\theta_i$ from $F_{s, AMD}$ (AMD) and $F_{s, BMD}$ (BMD).
To separate agglomeration economies and selection, CDGPR12 assume that $F_{s, AMD}$ and $F_{s, BMD}$ differ only in location ($\mu$), scale ($\sigma$), and left tail truncation $(\xi)$:
Here $(\mu, \sigma)$ capture agglomeration economies, while $\xi$ reflects stronger competition in denser areas. CDGPR12 view (ref) as a continuum of moments and estimate $(\mu, \sigma, \xi)$ with GMM.
Our goal is to test assumption (ref). If it fails, then the CDGPR12 estimators of $(\mu, \sigma, \xi)$ may potentially be biased or produce quantities that are challenging to interpret. Formally, we test the following null of distribution equality up to mean and variance for every sector $s$:
where $\mu_{s, AMD}, \mu_{s, BMD}, \sigma_{s, AMD}, \sigma_{s, BMD}$ are AMD/BMD sector-specific location and scale parameters. This null is more restrictive than (and hence implies) CDGPR12's, as it excludes $\xi$. We omit $\xi$ based on Morozov2026InferenceExtremeQuantiles who nonparametrically finds no evidence of left tail truncation. Moreover, our results in section (ref) show that location-scale shifts are sufficient for distribution equality for all $s$.
Importantly, our analysis is fully agnostic about how $F_{s, AMD}$ and $F_{s, BMD}$ arise or how to interpret $(\mu, \sigma)$. We also make no functional form assumptions on $F_{s, AMD}$ and $F_{s, BMD}$.
Testing the null in (ref) is challenging for two reasons. First, the TFPs $\theta_i$ are latent and must be estimated from data using some $\hat{\theta}_i$, as is standard in the literature. Second, $\hat{\theta}_i$ are noisy measurements of $\theta_i$. As Jochmans2019 show, naively using $\hat{\theta}_i$ may lead to invalid inference. We now address these challenges in turn.
We use administrative balance sheet and income statement data from the Banco de España's CBI dataset BancoDeEspana2024MicrodataIndividualEnterprises from 2000 to 2019. CBI is representative, covering $\sim$80% of incorporated non-financial Spanish firms and matching the aggregate dynamics of output, employment and wages Almunia2018. All monetary variables are deflated using industry-specific deflators from EU-KLEMS.
Firms are classified as belonging to AMD and BMD areas using the experienced density measure from Roca2017. This measure averages population within a fixed radius around individuals, smoothing boundary irregularities. Urban areas are defined according to Spain’s Ministry of Housing, excluding the exclave cities Ceuta and Melilla. Firms are linked to urban areas via municipality-postal code concordances from Instituto Nacional de Estadística. Additional standard data cleaning steps are detailed in the Appendix.
Productivity estimates $\hat{\theta}_i$ are obtained in two steps, mirroring Combes2012b:
Importantly, for testing (ref), we retain only firms observed for at least 15 periods to test the null (ref). This restriction is done for two reasons:
The number of such firms varies between $101$ and $16649$, depending on the sector and area type (see table (ref) for exact counts). 15 years is a compromise between having sufficiently long time series to capture selection, control estimation error (in the sense discussed in section (ref)) and maintaining a sufficient cross-sectional sample size for testing.
Having constructed $\hat{\theta}_i$, we proceed to test the null (ref) of distribution equality up to mean and variance. To do so, we use a tailored nonparametric two-sample Kolmogorov-Smirnov (KS) test with debiasing. We tailor the test to handle two key challenges:
Our approach --- formally defined in the Appendix --- modifies the standard KS test in two ways to address these challenges:
The test is straightforward. For each sector, we ask whether the standardized distributions of TFP (adjusted for mean and variance) are identical across AMD and BMD areas. Any remaining differences would indicate that the assumption of distributional equality is violated. The debiased test is statistically valid even with noisy TFP estimates thanks to the restriction on the minimal number of observations per firm (see section (ref) and the Appendix).
Table (ref) summarizes our results. It provides the $p$-values of our bias-corrected KS test for each sector, along with the firm counts in AMD and BMD areas. Additionally, Figure (ref) depicts debiased distribution functions (CDFs) of TFP. In line with the null hypothesis (ref), we standardize the CDFs to have the same mean and variance in both areas to allow for comparison of differences beyond these parameters.
Our results provide strong empirical support for the distribution equality assumption (ref). For every sector, we fail to reject the null that TFP distributions differ only in mean and variance between AMD and BMD areas. All $p$-values are large, with the smallest equal to 0.499. This pattern holds for both sectors with lower values of $N$ (e.g. Arts) and more populous ones (such as Trade). Visually, the mean-variance standardized distribution functions in Figure (ref) are extremely similar, particularly in more populous sectors where the estimates are more stable.
Interpreted narrowly, this finding validates the core assumption of CDGPR12 in the data, justifying their decomposition framework (and related methods) in prior work.
More broadly, our results suggest an even stronger relationship between TFP distributions: just two parameters --- mean and variance --- suffice to align the distributions up to the resolution afforded by our data. This suggests that the productivity differences between denser and less dense areas are driven predominantly by mechanisms that operate through these channels, such as various aspects of agglomeration economies.
Importantly, our results are nonparametric in nature, requiring only correct specification of sectoral production functions. They allow for trends in productivity, impose no restrictions on the interpretation of mean-variance shifts, and make no assumptions on functional forms of the TFP distributions.
These findings narrow the range of plausible explanations for productivity differences between areas. Since agglomeration economies are the leading explanation for mean-variance shifts, future research should focus on mechanisms operating through these channels. Policies aimed at fostering agglomeration (improvements in local infrastructure, labor market thickness, or knowledge spillovers) may help reduce productivity gaps between dense and sparse areas. Conversely, policies targeting firm selection (e.g., through competition policy) may be less effective, given the absence of tail truncation in our results and the extreme quantile results of Morozov2026InferenceExtremeQuantiles.
In this paper, we empirically validate the key assumption of Combes2012b that TFP distributions differ between denser and less dense areas only in mean, variance, and tail truncation. Using Spanish firm-level data and econometric methods robust to noisy productivity estimates, we show that this assumption holds across all sectors. Moreover, the mean and variance alone are sufficient to account for differences in distributions. Our results support the use of models relying on this assumption in policy analyses of urban productivity gaps and the design of place-based interventions.
Our findings suggest two further areas of research. First, future work could explore the specific mechanisms driving mean-variance shifts in productivity distributions, such as local infrastructure, labor market thickness, or knowledge spillovers, to help policymakers design more effective place-based interventions. Second, our statistical approach --- combining HPJ debiasing with nonparametric distribution tests --- may provide a template for validating assumptions in other contexts where noisy estimates are used (e.g., worker skills in Roca2017 or mutual fund manager skills in Barras2022SkillScaleValuea).