Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
39,988 characters · 8 sections · 53 citation commands
On Spatio-Temporal Stochastic Frontier Models
Understanding and analysing the efficiency of decision-making units has long been a cornerstone of different disciplines, including economics (fare1994productivity), agriculture (kumbhakar1995efficiency), and health studies (hollingsworth2008measurement). Traditionally, stochastic frontier analysis (SFA, Aigner1977, battese1995model) has served as a valuable tool for this purpose, providing information on both the level of production achieved and the potential for improvement. However, a crucial limitation of standard SFA lies in its neglect of the spatial context in which these units operate. Spatial dependence and spillover effects are increasingly recognised as significant factors that may influence production efficiency, which calls for analytical frameworks that incorporate these spatial dimensions.
This paper delves into the spatial stochastic frontier analysis (SSFA, fusco2013spatial) framework, an extension of SFA that explicitly accounts for spatial relationships and interactions. By acknowledging the spatial interconnections among decision-making units, SSFA provides a deeper and more detailed insight into efficiency differences. This enhanced framework allows us to disentangle the spatial aspects of inefficiency from pure unit-specific inefficiencies, revealing crucial insights into how location, proximity, and spatial interaction patterns influence production outcomes.\\ The rationale behind the use of SSFA is multifaceted. First, spatial dependence can arise from knowledge spillovers, shared infrastructure, or environmental features, leading to higher or lower efficiencies for neighbouring units depending on the nature of these interactions. Another point is that spatial autocorrelation in the error term could indicate the presence of unobserved spatial influences affecting production within the research region. Ignoring these factors might result in skewed assessments of technical efficiency and erroneous findings regarding the actual drivers of performance. The SSFA method has been applied to estimate efficiency in various domains such as agriculture (Vidolietal2016,Zulkarnain2020,Mittag2023), local governments (FuscoAllegrini20), hospitals (Cavalieri2020), food industry (Cardamone2020) or airports competition (Bergantino2021).
The extension of the SSFA to panel data is not simply an econometric and statistical exercise, but rather a response to a dynamic and evolutionary interpretation of economic firms, which operate in both time and space and undergo a continuous transition between equilibria. In fact, evolutionary approaches, which originated with darwin2016origin and emphasised the mechanism of natural selection and the dynamic struggle for survival, found some early contributions to economics in Schumpeter1934 before being systematised by nelson1985evolutionary and boschma1999evolutionary, Boschma2006.\\ For evolutionists, three key processes take place in the marketplace: selection, variation, and reproduction, where essentially, in the struggle for survival, the species with the greatest ability to adapt to new conditions will survive with the greatest ability to adapt to the environment. Similarly, in the market, firms compete dynamically to attract consumers, and the market is therefore a selection mechanism for firms. It shapes the opportunities and constraints on the growth, profitability, and likelihood of survival of firms, promotes "localized collective learning in a regional context" and "the spatial formation of newly emerging industries as an evolutionary process" (boschma1999evolutionary).\\ In this sense, dynamic efficiency (i.e. the ability to innovate) is much more important than static (or allocative) efficiency, which makes it necessary to study economic processes over time. But the selection principle just seen does not imply that "selection" necessarily goes "from worst to best". The "best" and "worst" are concepts that depend on the specific selection mechanisms, their history, but also by the space in which they live, and the distribution of the characteristics actually present in a given ecology or market at a given time. And so it is also the ecological niche - read as a territorial or neighbour constraint - that conditions, in the long run, the mode of production, the efficiency, and hence the survival of firms. Therefore, it is because of these premises that the study of the economic efficiency of firms in both time and space is crucial, because the former shows the evolution of the system, while the latter ignores this dynamic, depending on whether or not they belong to an ecological niche\footnote{Clearly understood as neighbourhoods, both physically, but also as part of production chains, networks, districts.} in the market.
This paper aims to contribute to the growing body of knowledge on spatial stochastic frontier methods by presenting a new methodology for estimating production and cost frontiers for time-varying spatial data focussing on the composite error part. We achieve this by providing comprehensive evidence through simulations, analysis of real-world datasets, and the development of specific software code designed to implement our approach. Incorporating time-varying elements into spatial frontier models allows for a more nuanced understanding of how efficiency and productivity evolve in response to changing conditions and external factors.\\ We contribute to the existing literature in two distinct ways. First, we address a significant methodological gap by focussing on the most crucial aspect of the specification of stochastic frontiers: the composite error term, which encompasses both the random error and inefficiency components. By refining the modelling of the composite error, we improve the precision and interpretability of frontier estimates, enabling a more accurate assessment of inefficiency levels. Second, we extend this estimation framework to accommodate both time-varying and time-invariant cases, providing a comprehensive analysis that accounts for different types of temporal dynamics in spatial data and offering a versatile framework that can be adapted to a wide range of applications.
The remainder of the paper is organised as follows. Section (ref) introduces the spatial stochastic frontier models with a particular focus on dynamic models. Section (ref) introduces the proposed methodology for estimation in space and time, while Sections (ref) and (ref) present, respectively, an application on simulated and real data. Section (ref) discusses the main findings and concludes.
Research on frontier efficiency methods that simultaneously account for spatial and temporal effects with panel data is limited and started in the 2000s (for a detailed overview, see Ayouba2023).\\ The main problem, in fact, as mentioned by sfa_book2004 was that "in classical time invariant SFA the fixed effects (the $u_i$) are intended to capture the variation between producers in time-invariant technical efficiency. Unfortunately, they also capture the effects of all phenomena (such as the regulatory environment, for example) that vary across producers but that are time invariant for each producer".
The primary methods suggested in the literature over the past two decades to address this consideration can be categorized based on three issues: (i) whether to consider efficiency as varying or constant over time; (ii) the method for estimating efficiency, thus regression-based with a two-stage approach or maximum likelihood; (iii) whether spatial dependence and spillover effects are only on the dependent variable, i.e. the output (spatial lag specification - SAR), on the regressors (spatially lagged $X$ specification - SLX), on the error term (spatial Error model specification - SEM), on the first two (spatial Durbin specification - SDM), or on all terms (General nesting spatial specification - GNS).
A first stream of literature proposes the SAR frontier model assuming a time-variant, as in Glass2013, Glass2014 and Kutlu2019 or time-invariant, as in Han2016, inefficiency term using a regression-based approach with a distribution-free approach to compute inefficiency in a second step.\\ Although distribution-free approaches have the advantage of not assuming a specific distribution for the inefficiency term, they are not robust to outliers. Therefore, Glass2016 proposed a SAR model for panel data integrating SAR and a half-normal stochastic frontier model proposed by Aigner1977 using a ML approach. Similarly, kutlu2020spatial also employed the SAR approach in their spatial stochastic frontier analysis. Tsukamoto2019 extended the SAR stochastic frontier specification considering also the determinants of technical inefficiency, as proposed by Battese95 using a ML approach.
A second group of scholars proposed a spatial Durbin specification as Adetutu15, Glass2016, Gude2018 and Ramajo2018 that integrates SLX and a half-normal stochastic frontier model; Galli2023 proposed a spatial Durbin stochastic frontier model for panel data introducing spillover effects in the determinants of technical efficiency (SDF-STE) using the ML approach.
Finally, a third stream of panel data literature has been proposed including spatial dependence only in the error term or in the inefficiency term (see e.g. Druska04,Areal12,tsionas2016spatial). In particular, Druska04 accounts for time-invariant fixed effects and uses a distribution-free approach, and Areal12, tsionas2016spatial uses a Bayesian approach.
In our opinion, grasping spatial lag in this manner is essential, despite the fact that the term "spatial" is frequently employed in this specific literature without a clear definition: within an evolutionary reading of economic processes introduced earlier, in fact, it is necessary to separate the effects of the firm from those of the neighbourhood in terms of productive efficiency in order to take into account different evolutionary paths, neighbourhood, and/or spillover effects; in other words, in stochastic efficiency models, it is more important to develop spatial error models within composite error (error + inefficiency part) (thus generically borrowing spatial error (SEM) approaches) than to only try to describe more precisely the frontier (as the spatial lag models, SAR), assuming that it is homogeneous in space. In fact, there is no practical value in incorporating a generic $W$ on the frontier and then comparing all heterogeneous units with each other on a single frontier, but rather in showing how much space affects the lower or higher efficiency of each individual production unit. Therefore, the composite error term lies at the heart of frontier models, and here the impact of the territory on the dynamic evolution of firms needs to be studied.
As introduced previously, this paper proposes an extension of the fusco2013spatial, Fusco20SSFA approach, which globally accounts for spatial and temporal effects in the term of inefficiency. In particular, coherently with the stochastic frontier literature (and specifically with Battese1992), this paper proposes two different versions of the model, namely: (i) a time-invariant model and (ii) a time-varying model. Notice that in both cases, the initial values for the maximum-likelihood estimation are derived from the algorithm. The two models will be discussed in detail in the next two sections.
In the proposed setting, we examine a production unit $i=1,...,N$ that, in a period $t=1,...,T$, uses $P$ inputs $\mathbf{x}_{it}=(x_{1it},...,x_{Pit})$, $\mathbf{x}_{it}\in \mathbb{R}_+^P$ to produce $Q$ outputs $\mathbf{y}_{it}=(y_{1it},...,y_{Qit})$, $\mathbf{y}_{it}\in \mathbb{R}_+^Q$. In this setting - and in the case of time-invariant - the generalisation of Fusco20SSFA for panel data, i.e. the Spatio-Temporal Stochastic Frontier Analysis (STSFA), can be written as:\footnote{Note that as in fusco2013spatial, Fusco20SSFA we consider the homoskedastic case and a Normal-Half-Normal distribution for the error term. Generalisations on these aspects can be developed in future research.}
where:
It is important to mention that in the proposed model, for the sake of simplicity, it has been assumed that the neighbourhood structure remains constant over time; in other terms, the spatial weight matrix $\mathbf{W}$ and the parameter $\rho$ have been set constant for all $t=1,...,T$.
The starting point for obtaining the general STSFA log-likelihood function, is to modify the density functions of $u$ and $v$ defined in Fusco20SSFA\footnote{In order to simplify formulas $(1-\rho \sum_{i}w_{i.})$ is replaced by $\delta(\rho)$ as in Fusco20SSFA.} accordingly with Battese1992:
where $f([\delta(\rho)]^{-1}\widetilde{u}_{i})$ is the same of SSFA as $[\delta(\rho)]^{-1}\widetilde{u}_{i}$ is independent of time and $f(\mathbf{v}_{i})$ now is time dependent so $\mathbf{v}_{i}=(v_1,...,v_T)$ .
Given the independence assumption, the joint density function of $[\delta(\rho)]^{-1}\widetilde{u}_{i}$ and $\mathbf{v}_{i}$ can be written as:
Since in the general form $\epsilon_{it} = v_{it} - s \cdot [\delta(\rho)]^{-1} \widetilde{u}_{i}$; therefore the new joint density function for $[\delta(\rho)]^{-1}\widetilde{u}_{i}$ and $\epsilon_{i}=(v_1-s\cdot[\delta(\rho)]^{-1} \widetilde{u}_{i},...,v_T-s\cdot[\delta(\rho)]^{-1} \widetilde{u}_{i})$ becomes:
and the marginal density function of $\epsilon_{i}$ is:\footnote{$\phi(\cdot)$ e $\Phi(\cdot)$ are the standard Normal density and distribution functions.}
Whereupon, the sought general "time-invariant" log-likelihood function for the sample of $N$ producers and $T$ periods is given by:
Finally, the "time-invariant" technical efficiency of the firm $i$ at the time period $t$ is given by:
where $\mu_i^*$ and $\sigma^*$ are the same in Eq.((ref)).
In the case of time-varying assumption, the generalisation of Fusco20SSFA for panel data, using the standard Battese1992 decay-model formulation, can be written as:
where $\eta_{it}=\exp \left\{-\eta(t-T) \right\}$ is a multiplicative factor, equal for all producers, that captures the dynamics of the degree of inefficiency. In particular, $\eta_{it}\geq 0$ decreases at an increasing rate if $\eta>0$, increases at an increasing rate if $\eta<0$, or remains constant if $\eta=0$.
Note that $u$ is once again independent of time, since the dynamics is given by $\eta_{it}$, so the density functions of $u$ and $v$ and the conjoint one are identical to Eqs. (ref), (ref) and (ref). Instead, what is changing in time-varying formulation is $\epsilon_{it} = v_{it} - s \cdot \eta_{it} \cdot [\delta(\rho)]^{-1} \widetilde{u}_{i}$, therefore the new joint density function for $[\delta(\rho)]^{-1}\widetilde{u}_{i}$ and $\epsilon_{i}=(v_1-s\cdot\eta_{i1}\cdot[\delta(\rho)]^{-1} \widetilde{u}_{i},...,v_T-s\cdot\eta_{iT}\cdot[\delta(\rho)]^{-1} \widetilde{u}_{i})$ is:
and $\eta_i$ represents the $T\times1$ vector of $\eta_{it}$'s associated with the time periods observed for the $i$-th firm.
The marginal density function of $\epsilon_{i}$ becomes:
The sought general "time-varying" log-likelihood function for the sample of $N$ producers and $T$ periods is:
Finally, the "time-varying" technical efficiency of the firm $i$ at the time period $t$ can be derived from the formula:
where $\mu_{\eta i*}$ and $\sigma_{\eta*}$ are the same in Eq.((ref)).
Note that if $\eta=0$, $\eta_i$ is equal to 1 and $\eta_i'\eta_i$ is equal to $T$, therefore, all equations collapse to their time-invariant versions.
In order to evaluate the small sample inferential properties of the estimators proposed in the previous section, we will first conduct a thorough analysis by presenting the results of a series of Monte Carlo experiments. These experiments are designed to rigorously test the performance and robustness of the estimators under various conditions. The simulation process ensures that each generated sample reflects the specified characteristics and distributional assumptions, providing a robust framework to analyse the performance and estimation accuracy of various econometric models applied to stochastic frontier analysis. More specifically, the common DGP used to simulate 1,000 samples consisting of $100$, $200$, and $400$ observations for $T=5$ years is a single-input-single-output stochastic production frontier for panel data. This model is structured as follows:
where:
It is important to highlight that for the purpose of emphasising the efficiency term in the simulation results, the regressor and the random term are kept constant throughout all 5 years. \\ Specifically, the frontier is estimated by the time-varying model for all combinations of $\rho=(0.05, 0.2, 0.4, 0.6, 0.8)$ and $\eta= (-0.10, -0.05, 0, 0.05, 0.10)$.\footnote{This sequence of values is chosen in order to increase or decrease the inefficiency by a maximum of 50% and the value $0$ is included to test the time-invariant specific case.} \\ Finally, the spatial weight matrix is a sparse matrix, as suggested by Lesage09, due to its low computational costs, where the average number of neighbours for each observation, identified by the nearest-neighbour method, is equal to $10\%$ of the number of observations $N$.
The parameters $\beta$, $\rho$ and $\eta$ are estimated according to the data generated by the rules and the selected DGP. The goodness of fit is assessed as usual by calculating the mean squared error (MSE) of the parameters and displaying the bias and standard deviation. The main findings are summarized as follows.\footnote{The elaborations have been carried out using the new functions of the SSFA R package (ssfa_r). The robustness check of the software is provided in (ref).} The estimation of the parameters demonstrates very low bias, standard deviation and MSE, consistently across all parameters (Table (ref)). This precision improves as the sample size grows (see Tables (ref) and (ref) for 100 and 200 sample size results). Furthermore, the kernel densities in Figures (ref) to (ref), illustrating that the distributions are closely aligned with the true values.
A well-known dataset used in the stochastic frontier literature to test our model is the Indonesian rice farm dataset published by Feng2012. This dataset is particularly valuable due to its comprehensive nature, capturing a wide range of variables that influence rice production in Indonesia. It includes panel data with $N=171$ spatial observations across $T=6$ waves, allowing an in-depth analysis of temporal and spatial variations in farm efficiency. Detailed information on inputs such as labour, land, and capital, as well as outputs, provides a rich foundation for evaluating the performance and efficiency of rice farms, making it an excellent resource for modelling and testing theories in the context of agricultural productivity and efficiency. More specifically, the inputs to rice production included in the data set are seed (kg), urea (kg), trisodium phosphate (TSP) (kg), labour (hours of work) and land (hectares). The output is measured in kilogrammes of rice. The data also include dummy variables. DP equals 1 if pesticides were used and $0$ otherwise. DV1 equals 1 if high-yield rice varieties were planted and DV2 equals 1 if mixed varieties were planted. The DSS equals $1$ if it was a wet season, $0$ otherwise. The spatial weights matrix $W$ is constructed as in Druska04: "Farms in the same village (out of six) are considered contiguous.".
Results of STSFA time-invariant and time-varying models and a comparison with Druska04\footnote{Results in "Table 3. Rice Regressions, Weighting Scheme M2 - Common Villages", on page 193 (Druska04).} are reported in Table (ref). It is important to note that from a methodological point of view, there are some main differences between our method and that of Druska04. In particular: (i) in D&H the model is time-invariant with fixed effects and the efficiency is estimated with a free-efficiency approach, instead in STSFA the model is time-invariant or time-varying and the efficiency is estimated with a ML approach; (ii) in D&H all the error term is spatially lagged, instead in STSFA only the inefficiency term is spatially lagged; (iii) in D&H the $\rho$ parameter is estimated for each $t$ and then the average is kept instead in STSFA $\rho$ is global.
Main findings are that $\beta$ coefficients are close to Druska04\footnote{Note that the robustness check of the time-invariant TSFA's intercept, $\sigma_u^2$ and $\sigma_v^2$ has been done by comparison with that of Horrace1996 (TSFA has been included in the Table (ref) as baseline model). The only difference is the absence of some dummies.} but the use of fixed effects to account for inefficiency in the error term in the first step leads to some bias in H&D's $\beta$ coefficients; our global $\rho$ is equal to 0.7103 (in a time-invariant model) and 0.7305 (in time-varying model) greater than D&H due to the fact that all the error is lagged and perhaps the parameter can be lowered by the absence of spatial autocorrelation in the random term; very interesting the $\eta$ is positive in SFA and negative in time-varying STSFA. This might imply that disregarding the spatial dependence in the inefficiency term can lead to a biased $\eta$. Consequently, while the overall efficiency of farms increases over time as a result of positive externalities, individual farm efficiencies tend to decline over time.
Economic processes are increasingly viewed as dynamic and constantly evolving and adapting, often influenced by spatial interactions and externalities, such as technology spillovers, local innovation, and resource sharing among neighbouring regions. \\ The introduction of the STSFA family of estimators marks a significant advancement in evolutionary economics by incorporating both spatial and temporal dependencies into the stochastic frontier model and focussing on the composite error term. Traditional models, in fact, often fail to account for these factors, leading to biased efficiency estimates. In particular, neglecting temporal dynamics can cause inefficiencies to be underestimated or overestimated. By integrating spatial and temporal dynamics, the STSFA model offers a more nuanced and accurate analysis of efficiency, which is crucial to understanding complex production processes. The STSFA model complements other methods proposed in the literature, but differs from them in that it focusses on modelling the error and inefficiency part, which we believe is precisely the core of frontier models; according to our perspective, in fact, it is more beneficial to understand the reasons and locations of systematic deviations from efficient behaviour by firms rather than focussing on the frontier itself.\\ This deeper insight is especially useful for policymakers who aim to create strategies that boost productivity. The STSFA model provides a sophisticated framework for identifying regions where policy interventions can have the greatest impact. By capturing the interaction between spatial and temporal factors, the model enables the identification of areas that would benefit the most from targeted policy measures. This information is vital for designing policies that not only improve efficiency, but also promote sustainable development.\\ To support the estimation of the STSFA model, the SSFA R package has been updated, offering a robust tool for researchers and facilitating a deeper exploration of efficiency in different sectors and geographic areas. \\ Finally, further applications of the model across various sectors and geographical areas would be valuable in generalising the findings, also exploring non-linear spatial dependencies and other forms of interactions between units.