Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
44,780 characters · 15 sections · 26 citation commands
GeMA: Learning Latent Manifold Frontiers for Benchmarking Complex Systems
\twocolumn[ \icmltitle{GeMA: Learning Latent Manifold Frontiers\\ for Benchmarking Complex Systems}
\icmlsetsymbol{equal}{*}
\icmlcorrespondingauthor{Jia Ming (Simon) Li}{[email removed]}
\icmlkeywords{Latent variable models, Deep generative models, Manifold learning, Frontier estimation, Efficiency analysis}
\vskip 0.3in ]
\printAffiliationsAndNotice{Transport Strategy Centre, Imperial College London, London, SW7 2AZ, United Kingdom}
Keywords: Latent variable models, Deep generative models, Manifold learning, Frontier estimation, Efficiency analysis.
Benchmarking the performance of complex systems is a widely used instrument in governance, regulation and strategic planning. Urban and national rail networks, renewable energy assets and national economies are routinely compared using efficiency scores derived from frontier methods. Conceptually, the task is to estimate an efficient frontier describing the best attainable outputs for given inputs, and to quantify how far each decision-making unit (DMU) lies from this frontier.
Two main families of methods have been particularly influential in this field. Stochastic Frontier Analysis (SFA) specifies a parametric production function with a composed error term that separates noise and inefficiency aigner1977formulation, meeusen1977efficiency, Greene1993, Sickles2019. Data Envelopment Analysis (DEA) constructs a piecewise-linear efficient frontier enveloping the data via linear programming charnes1978measuring, coelli2005introduction. Both approaches have been widely applied in transport, energy, health and macroeconomic applications oum1999survey, graham2008productivity, simar2015statistical and have become standard tools in regulatory benchmarking smith2005role, rostamzadeh2021application.
However, contemporary datasets pose several structural challenges. First, heterogeneity: DMUs may operate under markedly different technological or organisational regimes, so that a single global frontier is difficult to interpret and may systematically favour some groups over others. Secondly, non-convexity and non-linearity: indivisible investments, network effects and physical constraints can lead to production sets that are not well approximated by convex hulls or simple parametric forms keshvari2013stochastic. Thirdly, scale bias: in macroeconomic and infrastructure datasets, larger entities often receive higher efficiency scores simply because of their absolute size, even after controlling for standard inputs. Finally, there is a trust question: efficiency scores are commonly reported as point estimates, with limited information on their stability to small perturbations in data or model specification.
Recent work has sought to relax some of the structural assumptions in classical frontier analysis. Convex non-parametric least squares (CNLS) and related approaches reinterpret DEA as a convex regression problem and impose shape constraints via inequalities on regression coefficients daraio2006robust, kuosmanen2010data. Stochastic non-convex envelopment and order-$m$ frontiers explore non-convex production sets and robustness to extreme points simar2007improve, keshvari2013stochastic. In parallel, machine learning methods have been used to estimate production functions with greater flexibility, employing kernel methods, tree ensembles and neural networks Breiman2001_RF, chen2016xgboost, Goodfellow-et-al-2016. Deep learning has also been brought into efficiency analysis, both for predicting outputs and for re-evaluating DEA benchmarks bose2015neuraldea, guerrero2022combining, tsionas2023bayesian. These developments, whilst valuable, typically retain the view of the frontier as a function in the observed input--output space and do not fully exploit the potential of latent variable modelling.
In this paper, we take a complementary perspective. Building on the manifold hypothesis in representation learning bengio2013representation and ideas from geometric deep learning bronstein2017geometric, bronstein2021geometric, we model the production possibility set as a low-dimensional, non-linear manifold embedded in the joint input--output space, and treat the frontier as the Pareto boundary of the production set induced by a latent variable model. We introduce Geometric Manifold Analysis (GeMA), implemented by a productivity-manifold variational autoencoder ($\textsc{ProMan-VAE}$), which learns a latent technology coordinate and a separate inefficiency factor from data. This construction enables us to represent flexible, potentially non-convex frontiers, accommodate heterogeneous technologies as distinct regions of a shared manifold, define a quotient manifold that reduces certain scale effects, and attach a simple geometric robustness score to each efficiency estimate.
Building on these ideas, our contributions are as follows:
The remainder of the paper is organised as follows. Section (ref) briefly reviews related work in frontier analysis, latent variable models and geometric deep learning. Section (ref) introduces the $\textsc{GeMA}$ framework and the $\textsc{ProMan-VAE}$ architecture. Section (ref) presents synthetic experiments and case studies on wind farms, national rail operators and macroeconomic data. Section (ref) discusses implications and limitations, and Section (ref) concludes. Detailed derivations, additional experiments and data descriptions are provided in the appendices.
\paragraph{Frontier and efficiency analysis.} Classical efficiency analysis builds on the notion of a production set and its efficient frontier farrell1957measurement. Stochastic Frontier Analysis (SFA) specifies a parametric production function, often Cobb--Douglas or Translog, with a composed error term that separates statistical noise from a one-sided inefficiency component aigner1977formulation, meeusen1977efficiency, Greene1993, Sickles2019. This econometric approach supports statistical inference and panel extensions, but is sensitive to functional-form misspecification and typically relies on restrictive distributional assumptions. Data Envelopment Analysis (DEA) and related non-parametric methods construct a piecewise-linear frontier that envelops the data using linear programming charnes1978measuring, coelli2005introduction. Extensions relax disposability and returns-to-scale assumptions or consider non-convex hulls and order-$m$ frontiers deprins1984measuring, simar2007improve, keshvari2013stochastic. These methods have been widely applied in transport, energy and other infrastructure sectors oum1999survey, graham2008productivity, rostamzadeh2021application, but typically operate in the observed input--output space and impose convexity or parametric structure on the production set.
\paragraph{Machine learning and deep latent variable models.} Machine learning methods have increasingly been used to estimate production functions and efficiency scores with greater flexibility, employing kernel methods, tree ensembles and neural networks Breiman2001_RF,chen2016xgboost, Goodfellow-et-al-2016. Recent work has explored the use of deep learning architectures to model complex production relationships and inefficiency bose2015neuraldea, guerrero2022combining, tsionas2023bayesian. Deep generative models and variational autoencoders (VAEs) introduce explicit latent variables to capture low-dimensional structure underlying high-dimensional observations kingma2013auto, rezende2014stochastic, with disentangled representations aiming to separate latent factors with distinct semantic roles bengio2013representation, higgins2017beta. In the context of efficiency analysis, most deep learning approaches focus on prediction quality or on re-scoring DEA benchmarks, and do not explicitly define a production set or frontier in latent space. In contrast, $\textsc{GeMA}$ uses a VAE-style architecture to define a generative frontier model with explicit production-set semantics and an inefficiency factor.
\paragraph{Geometric and structured deep learning.} Geometric deep learning emphasises the importance of exploiting the underlying geometric structure of data---manifolds, graphs and more general structured domains bronstein2017geometric, bronstein2021geometric. The manifold hypothesis suggests that high-dimensional observations often concentrate near lower-dimensional manifolds; methods such as Isomap and locally linear embedding (LLE) roweis2000nonlinear,tenenbaum2000global and more recent approaches such as UMAP mcinnes2018umap provide tools for learning or visualising manifold structure. In parallel, latent manifold models have been proposed in scientific domains to interpret complex data geometry lopez2018deep, moon2019visualizing, nieh2021geometry, perich2025neural. In econometrics and operations research, geometric ideas have begun to appear in the analysis of efficient frontiers and multi-objective optimisation chatigny2024learning, felten2024toolkit, and there is growing interest in causal representation learning on learned manifolds scholkopf2021toward. $\textsc{GeMA}$ draws inspiration from this geometric perspective but uses relatively standard encoder--decoder architectures; the geometric structure enters through the interpretation of the learned latent space as a productivity manifold and through the quotient and certification constructions that we use for benchmarking and robustness diagnostics.
We now formalise Geometric Manifold Analysis (GeMA) and its $\textsc{ProMan-VAE}$ implementation. We first define a productivity manifold and the induced production set, then specify a generative model that disentangles latent technology from inefficiency. We subsequently introduce two geometric diagnostics: a quotient construction for scale-invariant benchmarking and a local certification radius for robustness assessment.
Let $\mathcal{X} \subset \mathbb{R}^d$ denote the input space and $\mathcal{Y} \subset \mathbb{R}^v$ the output space. Classical production theory assumes the existence of a production set $\mathcal{T} \subset \mathcal{X} \times \mathcal{Y}$ such that $(x,y) \in \mathcal{T}$ if and only if output $y$ is feasible with inputs $x$. The output-oriented efficient frontier is the Pareto-efficient boundary \[ \partial \mathcal{T} = \{(x,y) \in \mathcal{T}: \nexists \, y' \ge y \text{ with } (x,y') \in \mathcal{T}\}, \] with the inequality interpreted component-wise.
$\textsc{GeMA}$ defines the production set constructively via a low-dimensional manifold embedded in the joint input--output space. Let $z \in \mathbb{R}^K$ denote a latent technology coordinate and let \[ g_\theta : \mathcal{X} \times \mathbb{R}^K \to \mathbb{R}^v \] be a decoder network with parameters $\theta$. The productivity manifold is \[ \mathcal{M}_\theta \;=\; \{ (x, y) \in \mathcal{X} \times \mathbb{R}^v : \exists z \in \mathbb{R}^K \text{ with } y = g_\theta(x,z)\}. \] For each input--technology pair $(x,z)$ the decoder produces a point $(x, g_\theta(x,z))$ on the manifold. The production set induced by $g_\theta$ is then defined as \[ \mathcal{T}_\theta \;=\; \{ (x,y) \in \mathcal{X} \times \mathbb{R}^v : \exists z \in \mathbb{R}^K \text{ with } y \le g_\theta(x,z)\}, \] with $y \le g_\theta(x,z)$ interpreted component-wise. This captures the idea that any output vector not exceeding a frontier point in each component is feasible. The estimated frontier is the Pareto-efficient boundary $\partial \mathcal{T}_\theta$.
The latent coordinate $z$ may be viewed as a set of intrinsic structural parameters describing technology or operating conditions. Different regions of latent space encode different technological regimes or business models, and efficiency is assessed relative to the geometry of $\mathcal{T}_\theta$ at a given $(x,z)$.
We observe a dataset $\{(x_i,y_i)\}_{i=1}^N$ of decision-making units (DMUs). The $\textsc{ProMan-VAE}$ model treats each observed input--output pair as a noisy realisation of a two-stage process. In the first stage, a unit adopts a structural paradigm or technology $z_i$, which locates it on the productivity manifold $\mathcal{M}_\theta$. In the second stage, the corresponding frontier output is scaled down by an inefficiency factor and contaminated by random noise. The model is trained to invert this process: given an observation $(x_i,y_i)$, it infers a latent technology $z_i$ and an inefficiency $u_i$ that best explain the data.
Concretely, we introduce latent variables $z_i \in \mathbb{R}^K, u_i \in [0,\infty)$, where $z_i$ represents technology and $u_i$ represents inefficiency. We place a standard normal prior on technology and an exponential prior on inefficiency, $z_i \sim \mathcal{N}(0, I), u_i \sim \mathrm{Exp}(\lambda).$ Given $(x_i,z_i)$, the decoder network $g_\theta$ produces a (theoretical) frontier output $y_i^\ast = g_\theta(x_i,z_i).$ Actual output is modelled as \[ y_i = y_i^\ast \exp(-u_i) \varepsilon_i, \] where $\varepsilon_i$ is a multiplicative noise term. In practice we work in log-space and approximate the noise as additive Gaussian, yielding an SFA-style structural equation with the parametric frontier $f$ replaced by the manifold mapping $g_\theta(x,z)$. Full details of the log-space likelihood and variational objective are given in Appendix A.
\paragraph{Encoder and latent disentanglement.} To infer $(z_i,u_i)$ from data, $\textsc{ProMan-VAE}$ employs a split-head encoder $q_\phi(z_i,u_i \mid x_i,y_i)$ with parameters $\phi$. A shared multilayer perceptron processes the concatenated inputs and outputs and branches into two heads:
Sampling from $q_\phi(z_i,u_i \mid x_i,y_i)$ is implemented via the reparameterisation trick, allowing gradients to propagate through stochastic nodes during training.
To respect basic economic logic, we impose an approximate monotonicity constraint on $g_\theta$ by adding a penalty term $\mathcal{R}_{\mathrm{mono}}(\theta)$ to the loss, encouraging non-negative marginal products in each input dimension. In practice, we estimate partial derivatives of $g_\theta$ with respect to inputs by finite differences over a grid of reference points and penalise negative increments. Further details on the network architecture and regularisation are provided in Appendix B.
\paragraph{Variational objective and efficiency scores.} The model is trained by maximising an evidence lower bound (ELBO) on the log-likelihood of the outputs given the inputs. For a single observation $(x_i,y_i)$ the ELBO has the generic form
where $p(z_i,u_i) = p(z_i)p(u_i)$ is the prior and $p_\theta(y_i \mid x_i,z_i,u_i)$ is induced by the log-space structural equation. In our implementation, closed-form expressions are available for the Kullback--Leibler terms, and the reconstruction term is approximated by a per-output Huber or squared loss on log-outputs. Aggregating over the dataset and adding the monotonicity penalty yields the overall training objective. Derivations and implementation details are given in Appendix A.
Under this model, the inefficiency variable $u_i$ provides a natural scalar efficiency measure: the factor $\exp(-u_i)$ scales the frontier output downwards, so we define an output-oriented efficiency index $\mathrm{Eff}_i = \exp(-u_i) \in (0,1]$. In principle, one may also consider distance-based measures defined via the geometry of $\mathcal{T}_\theta$, such as the minimal distance in output space between $y_i$ and the frontier at input $x_i$. In this work we primarily use $\exp(-u_i)$, as it is directly learned by the model and has a simple interpretation as the proportion of frontier output.
The latent technology vectors $z_i$ provide a low-dimensional representation of structural heterogeneity. After training, we may project the $z_i$ into two or three dimensions using a manifold visualisation method such as UMAP and cluster them to identify endogenous peer groups. These clusters can be interpreted as regions of the latent manifold corresponding to distinct technological regimes or business models.
In macroeconomic and infrastructure applications, absolute size (for example measured by GDP, population or network length) can strongly influence efficiency scores. Larger countries or systems may appear efficient simply because they operate at a greater scale, even if their underlying technology is similar to that of smaller units. To address this, we introduce a simple equivalence relation that identifies scale-equivalent production points.
Given two points $(x,y)$ and $(x',y')$ in $\mathcal{T}_\theta$, we write \[ (x,y) \sim (x',y') \quad \text{if} \quad \exists \lambda > 0 \text{ such that } (x',y') = (\lambda x, \lambda y). \] Intuitively, two configurations are equivalent if they represent the same technology up to a common rescaling of all inputs and outputs. The set of equivalence classes $\mathcal{M}_\theta / \sim$ may be regarded as a quotient manifold in which absolute scale has been factored out.
In practice, we approximate this quotient mapping by normalising inputs and outputs and by relying on the latent technology variable $z_i$ learned by $\textsc{ProMan-VAE}$, which is encouraged to capture structural rather than purely scale-related variation. Benchmarking in the quotient space amounts to comparing units based on their position in latent technology space rather than on absolute magnitudes of inputs and outputs. In our experiments, we illustrate this idea using Penn World Table macroeconomic data, showing that a standard DEA-based efficiency index exhibits substantial correlation with country size, whereas an index based on $\textsc{GeMA}$ in the quotient space displays a much weaker association.
Efficiency scores computed from any model can be sensitive to small perturbations in the underlying data or to local irregularities of the estimated frontier. To provide a simple indication of local robustness, we define a certification radius based on the behaviour of the decoder near a given input.
Assume that for fixed latent technology $z$ the decoder mapping $x \mapsto g_\theta(x,z)$ is globally $L_\theta$-Lipschitz in $x$ with respect to the Euclidean norm. Let $J(x)$ denote the Jacobian matrix of partial derivatives of $g_\theta$ with respect to $x$ at $(x,z)$, and let $\sigma_{\min}(J(x))$ be its smallest singular value.
A large value of $R_{\mathrm{cert}}(x_i)$ suggests that the local mapping from inputs to frontier outputs is smooth and relatively well-conditioned, whereas a small value indicates that the decoder may have sharp bends or folds near $x_i$. To interpret this quantity, consider a perturbation $\delta x$ with $\lVert \delta x \rVert_2 \le r < R_{\mathrm{cert}}(x_i)$. By the Lipschitz property, the change in the frontier output is bounded by \[ \lVert g_\theta(x_i + \delta x, z_i) - g_\theta(x_i, z_i) \rVert_2 \le L_\theta r < \sigma_{\min}(J(x_i)). \] Thus, within a ball of radius $r$ in input space, the variation in frontier output is controlled by a bound that is strictly smaller than the smallest local amplification factor implied by the Jacobian. As $\sigma_{\min}(J(x_i))$ decreases towards zero, the certification radius shrinks, signalling that very small changes in inputs may induce large changes in outputs due to local geometric irregularities.
In our empirical analysis, we use $R_{\mathrm{cert}}(x_i)$ as a local robustness indicator for the efficiency score of unit $i$. High nominal efficiency that coincides with a very small certification radius can be interpreted as a “fragile” score in the sense that it relies on frontier geometry that is locally ill-conditioned. Such cases may warrant closer scrutiny in regulatory or policy applications. Practical details of Jacobian computation and input whitening are given in Appendix B.
We evaluate $\textsc{GeMA}$ on synthetic data and several real-world domains. The synthetic experiments probe whether $\textsc{GeMA}$ behaves sensibly in classical settings and where it brings structural advantages relative to established frontier estimators. The real-world studies focus on two domains of particular interest for the machine learning and regulatory communities: wind farm operations (WF) and national rail operators (ORR). Additional analyses on urban rail systems (COMET) and macroeconomic data (PWT) are reported in the appendix C.
Unless otherwise stated, all $\textsc{ProMan-VAE}$ models use the same encoder--decoder architecture with domain-specific input/output dimensions and modest hyperparameter tuning. Baselines include DEA with variable returns to scale (VRS), parametric SFA with a Translog specification, a free disposal hull (FDH) estimator, convex nonparametric least squares (CNLS) and a purely predictive machine learning baseline (random forest). Implementation details and hyperparameters are given in Appendix C.
The synthetic experiments examine three stylised settings in which key assumptions commonly imposed in efficiency analysis are selectively violated:
In all scenarios, outputs are generated from known production frontiers with multiplicative inefficiency and noise, following the structural form of the $\textsc{ProMan-VAE}$ model. We use $n=500$ DMUs and average results over $30$ Monte Carlo replications.
We compare methods using four metrics: frontier approximation error (Scenario A), inefficiency ranking quality (Scenarios A--C), cluster recovery via adjusted Rand index (Scenario B) and scale bias measured as the correlation between efficiency and size (Scenario C). Full data-generating processes and metric definitions are in Appendix B.
Table (ref) summarises the main results. When the data-generating process closely aligns with smooth parametric or low-dimensional predictive models (Scenario A), classical SFA and the machine learning predictor achieve the lowest frontier approximation errors and relatively high inefficiency rank correlations, and $\textsc{GeMA}$ performs comparably. In Scenarios B and C, where technological homogeneity and scale separability are violated, $\textsc{GeMA}$ achieves better recovery of latent technology groups (higher ARI) and substantially attenuates the correlation between estimated inefficiency and size when benchmarking is performed in the quotient space. This suggests that the advantages of $\textsc{GeMA}$ are structural rather than universal, and are most pronounced when heterogeneity and scale confounding are present.
We next examine the robustness of frontier-based performance assessments for British train operating companies (TOCs), using panel data compiled by the Office of Rail and Road. The aim is not to establish numerical dominance over classical frontier estimators, but to illustrate how $\textsc{GeMA}$ augments point assessments with robustness diagnostics that are directly relevant for regulatory benchmarking.
\paragraph{Data and setup.} The ORR data cover multiple TOCs observed over several years. Inputs include labour, route length, rolling stock and planned capacity; outputs include passenger-kilometres and train-kilometres, with additional quality indicators used in supplementary analyses. Route length also serves as a scale proxy. We treat each operator--year as one observation, apply log-transformations for numerical stability and estimate a $\textsc{ProMan-VAE}$ model with a low-dimensional technology space. For comparison, we estimate VRS DEA and parametric SFA models using the same inputs and outputs. Further details on variable construction and preprocessing are given in Appendix E.
\paragraph{Certification radii and fragile high scores.} For each operator--year observation, we compute a certification radius $R_{\mathrm{cert}}(x_i)$ as defined in Section (ref). Figure (ref) (left) shows the distribution of certification radii; most observations have moderate radii, suggesting locally well-conditioned frontier geometry for a large share of the sample, but the distribution exhibits a non-negligible left tail. Table (ref) reports representative percentiles.
A key use of this diagnostic is to highlight cases where high point scores may be sensitive to noise. Figure (ref) (right) plots the $\textsc{GeMA}$-based performance score against $R_{\mathrm{cert}}(x_i)$ and marks a “high-score / low-robustness” region (top decile of the score combined with bottom quartile of $R_{\mathrm{cert}}$). Observations in this region are not necessarily misclassified, but their rankings are more likely to be fragile with respect to measurement error or small input perturbations. From a regulatory perspective, such cases warrant closer scrutiny than similarly high-scoring observations supported by strong robustness guarantees.
\paragraph{Comparison with classical frontiers.} Overall rankings under $\textsc{GeMA}$ and SFA/DEA are moderately aligned at the aggregate level, but several TOCs exhibit substantial differences. Some operators that appear highly efficient under SFA have very small certification radii under $\textsc{GeMA}$, indicating scores that rely on locally ill-conditioned frontier geometry, while others have modest efficiency scores but large radii. The certification radius thus adds information that is absent from classical frontier estimators and can help regulators distinguish high scores that are well supported by the frontier geometry from those that are potentially fragile.
Our final main case concerns wind farm operations in China and highlights $\textsc{GeMA}$'s ability to recover non-linear physical frontiers in a physics-informed machine learning (PIML) spirit.
\paragraph{Data and model.} We use a publicly available wind power dataset from the State Grid Renewable Energy Generation Forecasting Competition, covering six wind farms with different turbine models, hub heights and rotor diameters. For each farm, we use two years of 15-minute SCADA measurements, including hub-height wind speed, wind direction, air temperature, air pressure, relative humidity and total active power output; wind direction is encoded via its sine and cosine. We also incorporate farm-level turbine configuration features derived from manufacturer specifications, such as swept rotor area, average hub height, average rotor diameter and number of turbines. All continuous variables are log-transformed or standardised as appropriate; details are in Appendix E.
We adapt $\textsc{GeMA}$ to the wind domain by treating each 15-minute timestamp as a DMU in a static frontier setting. The model takes as inputs the environmental variables and turbine configuration features and outputs a latent technology representation and a frontier prediction in log-transformed power space. Observations are modelled as $y_{\log} = y_{\log}^{\mathrm{frontier}} - u$, where $u \ge 0$ captures multiplicative operational losses and unexplained deviations from the physical frontier, including wake interactions, curtailment and grid-level constraints, rather than managerial inefficiency in the usual economic sense.
\paragraph{Predictive accuracy and efficiency levels.} Under a year-based train/validation/test split, $\textsc{GeMA}$ achieves an RMSE of about $0.72$ and an $R^2$ of about $0.82$ on the 2020 test set in log-transformed power space, indicating that the learned frontier is consistent with observed data. Averaging the efficiency ratio $\rho = p_{\mathrm{obs}} / p_{\mathrm{frontier}}$ over time, we find that the six farms operate at roughly $40$--$50\%$ of their learned frontier output once local wind conditions are controlled for, indicating broadly similar operational utilisation across farms.
\paragraph{Learned power curves and comparison to engineering models.}
Figure (ref) shows, for two representative wind farms, the empirical scatter of observed capacity factor versus hub-height wind speed, overlaid with the learned $\textsc{GeMA}$ frontier and a simple “toy” turbine curve constructed from manufacturer cut-in, rated and cut-out speeds. To facilitate visual comparison, we normalise the frontier capacity factor for each farm by its $95$th percentile so that the empirical plateau aligns approximately with capacity factor one.
Across all farms, the learned frontiers reproduce the characteristic three-stage structure of turbine power curves: near-zero output below cut-in, a steep non-linear ramp-up in the mid-range and a plateau around rated capacity. At high wind speeds the frontier capacity factor declines instead of remaining flat, particularly in farms with frequent high-wind events, consistent with early curtailment and grid-level constraints that are absent from the idealised toy curves. A simple parametric turbine model fitted to the learned $\textsc{GeMA}$ frontier yields cut-in and rated speeds that cluster around manufacturer values, despite the absence of hard physics constraints in the architecture or loss.
These results indicate that $\textsc{GeMA}$ can recover physically plausible non-linear frontiers from high-frequency operational data, while simultaneously providing an inefficiency factor that captures time-varying operational losses. This illustrates how latent manifold frontiers can be used as flexible approximations to engineering curves in a PIML framework, with potential applications in performance benchmarking, anomaly detection and planning under uncertainty.
For completeness, we briefly summarise two further applications; full details are given in Appendix G.
\paragraph{Urban rail systems (COMET).} We apply $\textsc{GeMA}$ to anonymised metro systems from the Community of Metros, covering networks from Asia-Pacific, Europe and the Americas. A two-dimensional latent technology space reveals four endogenous peer groups corresponding to large legacy systems, newer high-density networks and medium-sized balanced systems. Compared with a single global DEA frontier, $\textsc{GeMA}$ provides more graded within-group performance signals and separates structural heterogeneity (peer grouping) from performance differences.
\paragraph{Macroeconomic benchmarking (PWT).} Using Penn World Table data, we examine how benchmarking in latent technology space alters the interpretation of macroeconomic efficiency. A quotient-based efficiency index defined in the latent manifold attenuates the correlation between efficiency and population size relative to a standard DEA index, and reorders rankings within latent peer groups. This suggests that the quotient construction can mitigate scale-related biases while retaining a familiar frontier-based notion of efficiency.
Our experiments suggest that $\textsc{GeMA}$ is most informative when production technologies are heterogeneous, non-convex or confounded with scale, and when the goal is to obtain interpretable and robust benchmarking rather than to maximise predictive accuracy alone. In smooth, low-dimensional settings closely aligned with parametric specifications, classical SFA and convex envelopment methods perform very well and $\textsc{GeMA}$ behaves comparably but does not dominate.
Several limitations deserve emphasis. First, model complexity and computational cost are higher than for classical frontier models, as training a deep generative model with monotonicity regularisation and Jacobian-based certification requires GPUs and hyperparameter tuning. Second, while the latent technology and inefficiency variables have clear conceptual roles, the decomposition is not uniquely identifiable in a purely data-driven sense; in practice it is regularised by priors, shape constraints and the SFA-style structural equation, and individual latent dimensions should not be over-interpreted. Third, the quotient construction targets a specific notion of scale equivalence and the certification radius is a qualitative robustness indicator based on conservative Lipschitz bounds rather than a formal adversarial guarantee. These diagnostics should therefore be interpreted in conjunction with domain knowledge and sensitivity checks.
We introduced Geometric Manifold Analysis ($\textsc{GeMA}$), a latent manifold frontier framework that combines deep generative modelling with classical concepts from efficiency analysis. By modelling the production set as the image of a low-dimensional manifold in joint input--output space and augmenting it with a quotient construction and a certification radius, $\textsc{GeMA}$ provides tools for representing heterogeneous, non-convex technologies and for assessing the local robustness of efficiency scores. Synthetic experiments and case studies on national rail operators, wind farms and macroeconomic data indicate that $\textsc{GeMA}$ behaves sensibly relative to established frontier estimators and can offer additional insight in settings with pronounced heterogeneity, non-convexity or scale bias. Future work includes developing theoretical guarantees for latent manifold frontiers and extending the framework to dynamic “world-model” representations of evolving production systems.
This work proposes a latent-manifold framework for efficiency analysis that aims to provide fairer and more interpretable benchmarking of complex systems. In domains such as rail regulation and macroeconomic policy, the ability to distinguish structural heterogeneity from inefficiency and to flag fragile high-efficiency scores may help regulators and policymakers avoid misleading performance assessments. At the same time, the use of deep generative models introduces additional complexity and potential opacity; the latent variables should not be interpreted as causal without further analysis, and the robustness diagnostics we propose are qualitative rather than formal guarantees. We view $\textsc{GeMA}$ as a complementary tool that can support, but not replace, existing domain expertise and established frontier methods.