EconBase
← All papers

GeMA: Learning Latent Manifold Frontiers for Benchmarking Complex Systems

Jia Ming Li, Anupriya, Daniel J. Graham

arXiv 17 Mar 2026 · Machine Learning

arXiv:2603.16729 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Benchmarking the performance of complex systems such as rail networks, renewable generation assets and national economies is central to transport planning, regulation and macroeconomic analysis. Classical frontier methods, notably Data Envelopment Analysis (DEA) and Stochastic Frontier Analysis (SFA), estimate an efficient frontier in the observed input-output space and define efficiency as distance to this frontier, but rely on restrictive assumptions on the production set and only indirectly address heterogeneity and scale effects. We propose Geometric Manifold Analysis (GeMA), a latent manifold frontier framework implemented via a productivity-manifold variational autoencoder (ProMan-VAE). Instead of specifying a frontier function in the observed space, GeMA represents the production set as the boundary of a low-dimensional manifold embedded in the joint input-output space. A split-head encoder learns latent variables that capture technological structure and operational inefficiency. Efficiency is evaluated with respect to the learned manifold, endogenous peer groups arise as clusters in latent technology space, a quotient construction supports scale-invariant benchmarking, and a local certification radius, derived from the decoder Jacobian and a Lipschitz bound, quantifies the geometric robustness of efficiency scores. We validate GeMA on synthetic data with non-convex frontiers, heterogeneous technologies and scale bias, and on four real-world case studies: global urban rail systems (COMET), British rail operators (ORR), national economies (Penn World Table) and a high-frequency wind-farm dataset. Across these domains GeMA behaves comparably to established methods when classical assumptions hold, and provides additional insight in settings with pronounced heterogeneity, non-convexity or size-related bias.

Citation extraction

44
references
67
in-text mentions
44
distinct cited
0
self-citations
6,402
main-text words

appendix boundary found by appendix_command · 52% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Keshvari, A. and Kuosmanen, T (2013) Stochastic non-convex envelopment of data: Applying isotonic regression to frontier estimation0.73732100%
2Breiman, L (2001) Random forests0.64422100%
3Goodfellow, I., Bengio, Y., and Courville, A (2016) Deep Learning0.64422100%
4Greene, W. H (1993) The econometric approach to efficiency analysis0.64422100%
5Sickles, R. C. and Zelenyuk, V (2019) Measurement of productivity and efficiency0.64422100%
6Aigner, D., Lovell, C. K., and Schmidt, P (1977) Formulation and estimation of stochastic frontier production function models0.64422100%
7Bengio, Y., Courville, A., and Vincent, P (2013) Representation learning: A review and new perspectives0.64422100%
8Bose, A. and Patel, G (2015) “neuraldea”–a framework using neural network to re-evaluate dea benchmarks0.64422100%
9Bronstein, M. M., Bruna, J., LeCun, Y., Szlam, A., and Vandergheynst… (2017) Geometric deep learning: going beyond euclidean data0.64422100%
10Bronstein, M. M., Bruna, J., Cohen, T., and Veličković, P (2021) Geometric Deep Learning: Grids, Groups, Graphs, Geodesics, and Gauges0.64422100%

Showing the top 10 of 44 scored citations.