EconBase
← Back to paper

Assignment at the Frontier: Identifying the Frontier Structural Function and Bounding Mean Deviations

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

94,881 characters · 10 sections · 48 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Assignment at the Frontier: Identifying the Frontier Structural Function and Bounding Mean Deviations

abstractThis paper analyzes a model in which an outcome equals a frontier function of inputs minus a nonnegative unobserved deviation. The inputs may be endogenous (statistically dependent on the deviation). If zero lies in the support of the deviation given the inputs---an assumption we term assignment at the frontier---then the frontier is identified by the supremum of the outcome given those inputs, obviating the need for instruments. We then consider estimation with random error that is mean-independent of the inputs. Motivated by the assignment at the frontier assumption, we regularize estimation by requiring the fitted distribution of the deviation to maintain a minimum probability mass in a neighborhood of zero. Finally, we derive a lower bound on mean deviation, using only variance and skewness, that is robust to scarcity of data near the frontier. We apply our methods to estimate a frontier production function and mean inefficiency. Keywords: Assignment at the frontier, frontier structural function, deviations, stochastic frontier analysis, production function, identification, bound, regularization.

Introduction

This paper analyzes a model in which an outcome equals a function of inputs minus a nonnegative unobserved deviation. The function represents a frontier structural function (FSF) relating inputs to the outcome; the nonnegative unobservable is the deviation from the FSF.\footnote{The FSF is analogous to the average structural function (ASF) defined by blundell2003endogeneity.} The assumption that deviations are nonnegative is motivated by applications where they represent inefficiencies or other wedges, such as regulations, markups, or taxes. The inputs may be endogenous (statistically dependent on the deviation). If zero lies in the support of the deviation given inputs---an assumption we term assignment at the frontier---then the FSF is identified by the supremum of the outcome at those inputs, even when inputs are correlated with the deviation. Like random assignment, assignment at the frontier obviates the need for instrumental variables; unlike random assignment, inputs need not be exogenous.

For example, consider a production function where output equals a linear function of labor minus (nonnegative) inefficiency. When firms choose labor optimally based on productivity, labor is endogenous, and hence correlated with inefficiency. As a result, the conditional mean of output given labor does not identify the marginal productivity of labor, because as labor varies, changes in mean output conflate the effect of labor with changes in average inefficiency (so ordinary least squares is biased). However, when inefficiency enters separably,\footnote{Section (ref) discusses when the frontier identifies the production and mean production functions under additive separability and input rescaling, and when input-inefficiency interactions prevent identification.} the conditional supremum of output given labor does identify the marginal productivity: as labor varies, changes in frontier output are due only to the effect of labor, since frontier firms have zero inefficiency given labor. The source of identification differs. Identification of the mean production function typically comes from mean independence of productivity and a set of exogenous variables. In contrast, identification of the frontier production function comes from the presence of firms arbitrarily close to efficiency, given inputs.

The identifying assumption that zero lies in the support of the deviation given inputs is equivalent to the infimum of the deviation equaling zero. Assignment at the frontier imposes no further restriction on the joint distribution of the deviation and inputs. In particular, the assumption allows the density of the deviation to vary with inputs at all points of its support. Figure (ref) illustrates this, showing how the FSF corresponds to the maximum of the outcome’s support at each input value (that is, the upper envelope of the outcome distribution), and that the distribution of the deviation varies with the input. Thus, inputs need not be exogenous, and neither the mean nor any quantile is necessarily a constant downward shift from the FSF.

figure[figure omitted — 1,828 chars of source]

As Figure (ref) suggests, the FSF is identified by the supremum of the outcome given inputs, making a nonparametric Data Envelopment Analysis (DEA) frontier a natural estimator charnes1978measuring,farrell1957measurement. DEA, however, ignores random errors. In our setting, we observe inputs and outcomes; we do not observe any auxiliary indicator or proxy for proximity to the frontier. As a result, a sample supremum of the outcome given inputs may reflect a large positive error rather than a near-frontier observation. We therefore model outcomes using stochastic frontier analysis (SFA) to separate deviations from random error. Unlike much of the SFA literature, we allow the deviation to be statistically dependent on inputs. Without such dependence, the frontier reduces to a constant upward shift of the conditional mean, implicitly imposing exogeneity. The random error is assumed mean-independent of inputs.

In a cross-sectional setting, identification strategies rely on support restrictions (deviations on the nonnegative reals and errors on the reals) and often impose symmetry on the random error. Recent work shows that when both the frontier and deviations depend on the same observables, finite-sample estimation can have difficulty separating the two parmeter2024inference. We therefore focus on a panel setting with time-invariant deviations, which exploits within- and between-firm variation for identification rather than symmetry or support restrictions on the random error.

We implement this strategy using a method of moments estimator olson1980monte adapted to the panel setting. First, the outcome is nonparametrically regressed on inputs to estimate the conditional mean outcome. This regression includes any endogeneity bias arising from dependence between inputs and the deviation. Second, we use within–between decompositions to form bias-corrected powers of the residuals and, following approaches for variance estimation under heteroskedasticity hall1989variance, fan1998efficient, flexibly regress these bias-corrected powers of the residuals on inputs to estimate conditional central moments of the deviation. Third, at each input value we specify a parametric conditional distribution for the deviation and estimate its parameters by method of moments, subject to a near-frontier mass constraint (a regularization that enforces a minimum probability mass near the frontier). Finally, we compute the conditional mean deviation implied by these estimated parameters and add it to the estimated conditional mean outcome to obtain an estimate of the frontier at a given input value. This estimation strategy allows the conditional distribution of the deviation to vary flexibly with inputs.\footnote{For large-sample properties and inference of related estimators, see simar2017nonparametric and parmeter2024inference.}

The regularization imposed by the near-frontier mass constraint addresses the fact that estimated central moments may be only weakly informative about near-frontier probability mass: different parameter values can fit the estimated central moments similarly while implying different mean deviations. Motivated by the assignment at the frontier assumption, we impose a near-frontier mass constraint requiring the fitted distribution to place at least a minimum probability mass in a neighborhood of zero, with the threshold decreasing with the local effective sample size. This regularization may introduce bias when the true near-frontier probability mass falls short of the imposed lower bound, but it reduces sensitivity of the implied mean deviation to estimation noise in the central moments. In the simulations, the unconstrained estimator can have explosive mean squared error (MSE). In the empirical application, it produces implausibly large mean inefficiency estimates.

Often, research interest lies not in the frontier itself, but in deviations from it: thus, firm inefficiency in the regulated utility sector knittel2002alternative; regulatory tax measurement in the housing market BenMosheGenesoveRegulation,glaeser2005manhattan; cross-country differences in factor use as indicators of misallocation hsieh2009misallocation; and firm markups loecker2012markups,de2020rise,hall1988relation. In DEA, empirical work usually reports efficiency scores and peer benchmarks rather than characterizing the frontier banker1984some,fare1994productivity; in SFA, the primary objective is usually to measure firm-level inefficiency, with the frontier estimated to separate it from random error jondrow1982estimation,battese1995model. When the assignment at the frontier assumption holds and the frontier is estimable, deviations can be recovered from differences between the estimated frontier and observed outcomes.

However, assignment at the frontier need not hold. For example, if firms that choose inputs suboptimally are also always inefficient in applying those inputs, then the upper envelope of observed outputs will lie strictly below the frontier production function, touching it only at cost-minimizing input vectors chosen by efficient firms.\footnote{For strictly quasi-concave and monotone production functions, cost minimization under unconstrained input choice and given prices implies a unique input vector. Thus, identification of the frontier production function over some region of inputs requires: (a) variation in input prices, (b) adjustment costs for inputs, or (c) firms that are efficient in applying inputs making suboptimal input choices. These are the same sources for input variation as in mean regression.} More generally, firms that choose inputs suboptimally but apply those inputs efficiently may exist, but may be rare at any given input vector. Near-frontier observations may therefore be scarce, especially at input vectors far from those of cost-minimizing efficient firms.

Scarcity of near-frontier data motivates our focus on bounding the mean deviation. We derive a lower bound on the mean deviation using only variance and skewness. Because it depends only on the second and third central moments, it can remain informative even when there are no observations near the frontier. This contrasts with frontier-based methods, which must rely on extrapolation based on distributional assumptions. Regularization is attractive when a parametric specification is a reasonable approximation and the near-frontier mass constraint is plausible. The lower bound is a robust complement: it imposes no parametric distributional assumptions and remains valid even when scarcity causes the near-frontier mass constraint to fail.

We apply our results to estimate a production function, which is usually estimated using conditional mean regression. To address the main challenge of endogeneity, the now standard approach uses a control function with a proxy variable to capture unobserved productivity olleypakes,levinsohn2003estimating. In contrast, our frontier approach identifies the frontier production function by exploiting two assumptions: inefficiency is nonnegative and zero lies in its support given inputs. This identification strategy requires neither proxy variables nor instrumental variables. Our approach allows the distribution of inefficiency to be statistically dependent on inputs, so inputs need not be exogenous. Allowing this dependence is central to our approach: if inefficiency were mean-independent of inputs, the frontier would differ from the conditional mean only by a constant shift.

Our analysis builds on the SFA literature kumbhakar2022stochastic, which originally assumed a linear frontier production function and mutually independent inefficiency, random error, and inputs aigner1977formulation,meeusen1977efficiency. Since then, SFA models have incorporated more flexible frontier functions fan1996semiparametric and distributions reifschneider1991systematic,parmeter2019combining.

Given the centrality of endogeneity in econometrics, growing attention has focused on it in frontier estimation amsler2016endogeneity,karakaplan2017endogeneity,prokhorov2021estimation,centorrino2023maximum. A central contribution of our paper is to show that common solutions, such as instrumental variables, are often superfluous. In fact, the frontier is identified by nonnegative inefficiency and the assignment at the frontier assumption, provided that the distribution of inefficiency is allowed to vary with inputs. Some SFA models permit dependence between inefficiency and inputs, often to model the determinants of inefficiency. However, to our knowledge, the literature has not recognized that allowing for this dependence is what accounts for endogeneity in frontier estimation.

A potential outcomes framework can elucidate why endogeneity is not a concern under assignment at the frontier and how this contrasts with random assignment. Let $\tau \in [0,1]$ index unobserved firm types, with potential outcomes $Y_0(\tau)$ under control and $Y_1(\tau)$ under treatment, each strictly decreasing in $\tau$ and with finite suprema at the frontier type $\tau=0$. The observed outcome is $Y = D \cdot Y_1(\tau) + (1-D) \cdot Y_0(\tau)$, where $D$ indicates treatment. Under random assignment, $\tau$ and $D$ are independent, and the average treatment effect is identified by the difference in conditional means, $E[Y \mid D = 1] - E[Y \mid D = 0]$. Alternatively, under assignment at the frontier, $\tau=0$ is in the support under both treatment and control, and the frontier treatment effect is identified by the difference in conditional suprema, $\sup(Y \mid D = 1) - \sup(Y \mid D = 0)$. Random assignment requires the distribution of types to be the same across treatment and control. Assignment at the frontier requires that the potential outcomes have finite suprema and that every neighborhood of the frontier type has strictly positive probability under both groups. The average treatment effect and the frontier treatment effect may differ substantially, and the relevant question determines which is of primary interest.\footnote{For policies aimed at the general population, the average effect is often the relevant object. For policies concerned with best practice or maximum potential, where a natural frontier exists, the frontier effect may be more relevant.}

Our approach can be compared to the identification at infinity literature, which exploits tail behavior for identification---using large values of covariates or instruments lewbel2007endogenous,chamberlain1986asymptotic or large values of the outcome itself dhaultfoeuille2013another. That work assumes the identified structural relationship in these extremes extends to average observations through a stable functional form and is usually interested in average effects. By contrast, our interest is in the frontier (e.g., the frontier production function) and the economically meaningful deviations from it (e.g., firm inefficiency). Our identification does not rely on extreme inputs or extreme unconditional outcomes; instead, it relies on extreme outcomes conditional on inputs.

Monte Carlo simulations assess finite-sample performance of our moment-based estimator of mean deviation and of the skewness-based lower bound. Two issues arise in small samples: the estimated lower bound can exceed the estimated mean, and the estimated mean can be unstable even when the lower bound remains accurate. The near-frontier mass constraint substantially reduces MSE. It rarely binds when the underlying distribution has high near-frontier mass, indicating it does not mechanically override moment information.

For illustrative purposes, we apply the methods to plant-level data from Colombian manufacturing in the food products industry.\footnote{These data have been widely used in production applications eslava2004effects,fernandes2007tradepolicy. We use the version provided in the gnrprod R package for gandhi2020identification.} Estimated mean inefficiency is correlated with inputs, whereas standard OLS and SFA rule out such dependence by assumption. In the absence of regularization, unconstrained central moment matching can produce fitted inefficiency distributions with little mass near the frontier and implausibly large mean inefficiency estimates. The regularized estimator yields results that capture the overall shape of the empirical distribution while satisfying the near-frontier mass constraint.

The remainder of the paper is organized as follows. Section (ref) presents the structural model and defines the FSF and mean deviation. Section (ref) introduces the assignment at the frontier assumption and identifies the FSF. Section (ref) provides conditions under which the FSF recovers the ASF and structural function. Section (ref) introduces random errors. Section (ref) derives lower bounds for mean deviation using variance, skewness, and higher central moments via the Stieltjes moment problem. Section (ref) presents semiparametric method of moments estimation with regularization using the near-frontier mass constraint in a panel data setting with time-invariant deviation. Section (ref) presents simulations. Section (ref) presents the empirical application to production. Section (ref) concludes.

The Model

Consider the structural model

align[align omitted — 67 chars of source]

where $y$ is an observed scalar outcome, $x$ is a vector of observed inputs that may be arbitrarily dependent on the vector of unobservables $\omega$, and $\widetilde g(\cdot)$ is an unknown structural function.

Define the FSF as the supremum of the structural function given $x$, and assume this supremum is finite:

align[align omitted — 85 chars of source]

The supremum need not be attained. When it is, it may be achieved by multiple $\omega$-types for a given $x$, and this set of optimal types can vary across $x$. For example, in a production setting with $\omega$ representing managerial efficiency, different managers could be equally optimal for a given set of inputs, and the most effective manager type could change with the firm's input mix. The substantive assumption, however, is that the supremum itself is finite. For example, a firm’s maximum possible output must be constrained by some technological barrier, given current know-how.

Rewrite the structural model (ref) as the frontier model

align[align omitted — 79 chars of source]

where $u := g(x)  - \widetilde g(x,\omega) \geq 0$ is the unobserved deviation from the FSF ($u=0$ if and only if $\omega \in \arg\max_{\omega'} \widetilde g(x,\omega') $). This framework allows an unrestricted joint distribution of $(u,x)$, and hence endogenous $x$, requiring only $u \geq 0$.

In addition to the FSF in (ref), we consider the mean deviation:

align*[align* omitted — 49 chars of source]

For example, in a production setting, for given inputs the FSF equals the frontier production function and the mean deviation equals mean inefficiency.

Assignment at the Frontier and Identification

We assume that zero is in the support of the deviation, given inputs.

asn[Assignment at the Frontier] \begin{align*} 0 \in Support(u \mid x). \end{align*}

In the context of the structural model (ref), Assumption (ref) requires that at least one type $\omega \in \arg\max_{\omega'} \widetilde g(x,\omega')$ lies in the support. This is the formal counterpart to the condition in the potential outcomes framework that the frontier type $\tau=0$ lies in the support. Beyond requiring support at zero (equivalently, $\inf(u\mid x)=0$ since $u\ge 0$), the assumption imposes no restriction on how the distribution of $u$ varies with $x$.\footnote{If the infimum deviation is a constant $u_{\min}>0$ common to all inputs, the FSF is identified only up to a parallel shift. This could arise, for example, if deviations reflect regulatory taxes with a uniform minimum level. If instead the infimum varies with inputs, not even the shape of the FSF is identified.} For example, if $u$ measures inefficiency, the assumption states that given inputs $x$, there exist firms arbitrarily close to full efficiency ($u=0$), though the density near zero may vary with $x$.

Violations of the assignment at the frontier assumption may be empirically observable. Economic theory can impose shape constraints on the frontier (e.g., monotonicity, concavity, or homogeneity), so an estimated frontier that violates any of these (e.g., a downward sloping supply curve or increasing returns where only constant or decreasing returns are admissible) suggests the frontier is not being observed in that region.

The FSF is identified by the supremum of the outcome given inputs.

propAssume that (ref)--(ref) and Assumption (ref) hold. Then $g(x)$ is identified.
proofBy Assumption (ref) ($0\in\text{Support}(u\mid x)$) and (ref) ($u\ge 0$), it follows that $\inf(u\mid x)=0$. Using (ref) ($y=g(x)-u$), we have \begin{align*} g(x) &= g(x) - \inf(u \mid x) = \sup\big(g(x)-u \mid x\big) = \sup(y \mid x). \qedhere \end{align*}

Thus the assignment at the frontier assumption obviates the need for instrumental variables or other corrective measures for endogeneity.

The above identification argument follows a structure similar to that of mean regression under conditional mean independence: If $y=g(x)-\varepsilon$ and $E[\varepsilon \mid x]=0$, then

align*[align* omitted — 99 chars of source]

This same logic also applies when identification relies on instrumental variables.\footnote{Suppose $y=g(x)-\varepsilon$ and there exists exogenous variable $z$ such that $E[\varepsilon \mid z] = 0$. Then $g(x)$ is identified from $E[g(x) \mid z] = E[y \mid z]$, under a completeness condition (i.e., the mapping $g \mapsto E[g(x) \mid z]$ is injective) newey2003instrumental.}

However, the assignment at the frontier and mean independence assumptions are conceptually different. Assignment at the frontier assumes that the support of $u \mid x$ includes zero, so identification does not require exogenous inputs, instruments, or any orthogonality conditions. Moreover, it is a pointwise support condition: at each input value $x$ where identification is claimed, $0 \in \text{Support}(u \mid x)$. This is an economically substantive restriction. It requires that $g(x)$ be attainable arbitrarily closely at that specific input vector, rather than being ruled out there by inefficiency, distortions, or other frictions. Mean independence, by contrast, assumes an orthogonality condition, here that $E[\varepsilon \mid x]=0$, and places no restrictions on the support of $\varepsilon \mid x$, including its behavior in neighborhoods of zero. Its identifying content comes from cross-$x$ restrictions on conditional means, rather than from a pointwise support condition.

Relating the FSF to the ASF and Structural Function

In general, unless strong restrictions are imposed on the structural function $\widetilde g(\cdot)$ in (ref), the FSF does not recover the ASF or the structural function.\footnote{The ASF is a central object of interest in the control function literature, defined as $ASF(x) := E_{\omega}[\widetilde g(x,\omega)]$. It represents the average outcome for a given $x$ if the endogeneity problem were absent, as it integrates over the marginal distribution of the unobservable $\omega$ rather than the conditional distribution used to compute $E[y \mid x]$.} Nevertheless, there are cases where it does. Specifically, the FSF recovers the ASF under additive separability, and it recovers the full structural function when inefficiency acts solely by rescaling inputs. However, in other specifications, such as models with interactions between unobservables and inputs, the FSF generally fails to identify either the ASF or the structural function. The following three examples in a production-function setting illustrate these different identification cases.

First, consider the model in which inefficiency is additively separable: a type-$\omega$ firm has production function $g(x) - \omega$, where $\omega \geq 0$. In this case, the FSF identifies the function $g(\cdot)$ by observing the most efficient firms on the frontier (where $\omega=0$). However, the mean regression function, $E[y \mid x] =g(x) - E[\omega \mid x]$, fails to identify $g(x)$ if input choices are endogenous (that is $E[\omega\mid x]$ depends on $x$). Because the ASF is $g(x) - E[\omega]$, the FSF therefore identifies the ASF as well.\footnote{Formally, for $\widetilde g(x,\omega) = g_1(x) - g_2(\omega)$ with $g_2(\omega)\geq g_2(0)$, the FSF is $g_1(x) - g_2(0)$ and the ASF is $g_1(x) - E[g_2(\omega)]$. Their difference, $c = E[g_2(\omega)] - g_2(0)$ is identified by $E[FSF(x)- y]$, provided $FSF(x)$ is identified over the support of $x$.}

Second, consider a model in which inefficiency rescales inputs: a type-$\omega$ firm has production function $g(x e^{-\omega})$, where $x e^{-\omega}$ are inefficiency-adjusted inputs and $\omega$ is a vector with nonnegative components. The structural function is thus $\widetilde g(x, \omega) = g(x e^{-\omega})$. This means the output of a firm with inputs $x$ and inefficiency $\omega$ equals that of an efficient ($\omega=0$) firm with inputs rescaled to $x e^{-\omega}$. In this case, the FSF identifies the entire structural function. Once the function $g(\cdot)$ is obtained from the frontier, the structural outcome for any $(x, \omega)$ pair is known simply by evaluating the frontier function at the rescaled inputs.\footnote{Formally, the FSF identifies the function $g(\cdot)$ because $FSF(x) = \widetilde g(x,0) = g(xe^{-0}) = g(x)$. The structural function for any $(\omega,x)$ is $\widetilde g(x, \omega) = g(xe^{-\omega})$. Therefore, the structural function is simply the identified frontier function evaluated at the rescaled input vector, i.e., $\widetilde g(x, \omega) = FSF(xe^{-\omega})$.} However, the ASF is not identified in general: a scalar outcome $y$ is not sufficient to uniquely determine the multidimensional inefficiency vector $\omega$ (and thus not its distribution). An exception is the single-input case with a strictly monotonic frontier $g$, which allows $\omega$ to be recovered from the observed data as $\omega = \ln(x) - \ln(g^{-1}(y))$.

Finally, consider a model in which inefficiency interacts with inputs: a type-$\omega$ firm has production function $y=\beta x-\alpha x\omega-\omega$, where $x,\beta,\alpha,\omega\ge 0$. The marginal productivity of such a firm is $\beta-\alpha\omega$. However, consider the alternative model in which a type-$\omega^*$ firm has production function $y=\beta x-\omega^*$, where $\omega^*:=(\alpha x+1)\omega$. The marginal productivity of such a firm is $\beta$. The two models are observationally equivalent in $(x,y)$, but differ in that the first assumes that $\omega$ is invariant to a change in $x$, whereas the second assumes that it is $\omega^*$ that is invariant. The models share the same $FSF(x)=\beta x$ but differ in the ASF. While the frontier is identified by observing efficient firms with $\omega=0$, the interaction parameter $\alpha$ is not.\footnote{More generally, for any alternative parameter $\alpha^* \ge 0$, the model $y=\beta x-\alpha^*x\omega^*-\omega^*$ with $\omega^*:=\frac{\alpha x+1}{\alpha^*x+1}\omega$ has interaction $\alpha^*$ yet generates the same $(x,y)$ distribution, with mean marginal productivity $\beta-\alpha^*E[\omega^*\mid x]$ rather than $\beta-\alpha E[\omega\mid x]$. Consequently, $\alpha$ is not identified.}

Random Errors

Consider the frontier model (ref)--(ref) with random errors in a panel setting with time-invariant deviations, so that for any given firm

align[align omitted — 124 chars of source]

where $u$ is a time-invariant deviation and $v_{t}$ is a random error.

Assume that the errors are mean independent of the inputs and that errors and deviation are mutually independent conditional on inputs.

asn\ \begin{enumerate}[(i)] • $E[v_t \mid x_1,\ldots,x_T] = 0$ for all $t$; • $u \mathrel{\perp\!\!\!\perp} (v_1,\ldots,v_T) \mid (x_1,\ldots,x_T)$; • $(v_1,\ldots,v_T)\mid(x_1,\ldots,x_T)$ are mutually independent and identically distributed. \end{enumerate}

The independence in Assumption (ref)(ii) holds only conditional on $(x_{1},\ldots,x_{T})$, allowing dependence between $u$ and $\{v_t\}_{t=1}^T$ through the (potentially endogenous) inputs. Conditional on $(x_1,\ldots,x_T)$, the distribution of $y_t$ is the convolution of the distribution of $g(x_t)-u$ with that of $v_t$.

In a cross-sectional setting (i.e., $T=1$), identification of $g$ relies on differences in the supports of $u$ on $[0,\infty)$ and $v$ on $(-\infty,\infty)$, often exploiting asymmetry of $u$ and symmetry of $v$.\footnote{schwarz2010consistent show identification when $u$ has gaps in its support and $v \sim N(0,\sigma^2)$ with unknown $\sigma^2>0$. delaigle2016methodology show identification when $u$ is symmetric and indecomposable and $v$ is symmetric. bertrand2019flexible show identification when $u$ is bounded and $v \sim N(0,\sigma)$. florens2020estimation show identification when $\text{Support}(u) = [0,\infty)$ and $\text{Support}(v) = (-\infty,\infty)$ with $v$ symmetric.} In contrast, panel data (i.e., $T \geq 2$) lets us separate the central moments of $u$ and $v_t$ using within- and between-firm variation, through within-firm differencing and between-firm averaging. For the procedures below, we therefore only require identification of the second through fourth conditional central moments (Appendix (ref)).

Bounding Mean Deviations

When data near the frontier are scarce, direct estimation of the frontier is infeasible from the data alone; instead, any estimate must rely on extrapolation based on distributional assumptions. This undermines precise estimation of the mean deviation based on the difference between an estimated frontier and the observed outcome. However, the mean deviation can still be bounded from below by a method that is robust to scarcity of data near the frontier. In this section, we use the Stieltjes moment problem to derive a family of lower bounds on the mean of a nonnegative random variable using only central moments. One bound in this family that uses only variance and skewness is especially useful in practice, as these moments are often reliably estimable even when data are scarce near the frontier, thereby avoiding strong assumptions about the distribution of the deviation near zero. These bounds provide no meaningful information about $g(x)$, as they impose no constraints on its derivatives.

Moment problems seek necessary and sufficient conditions for a sequence of raw moments $(m_j)_{j=1}^\infty$, where $m_j = E[u^j]$, to correspond to the distribution of a random variable $u$ schmudgen2017moment.\footnote{Recent work uses moment problems in fixed-effects logit models davezies2025identification,dobronyi2021identification.} The Stieltjes moment problem adds the restriction that the support of $u$ is contained in the nonnegative real line. Let $\Delta_n^{(0)}$ and $\Delta_n^{(1)}$ denote the $n\times n$ primary and shifted Hankel matrices: {

align*[align* omitted — 432 chars of source]

}A well-known result states that if $\text{Support}(u)\subseteq[0,\infty)$ and \[ \det\bigl(\Delta_n^{(0)}\bigr) \geq 0,\quad \det\bigl(\Delta_n^{(1)}\bigr) \geq 0 \quad\text{for all } n \in \mathbb{N}, \] then there exists at least one distribution on $[0,\infty)$ with those moments. If any determinant is negative, no such distribution exists.

In Appendix (ref), we show that nonnegativity of the primary Hankel determinants implies well-known inequalities, such as nonnegative variance and that kurtosis exceeds the squared skewness by at least one. These restrictions depend only on central moments and not $E[u]$, so they do not produce bounds on $E[u]$.

Next, we turn to the nonnegative determinants of the shifted Hankel matrices. These produce inequalities in the raw moments $m_j$ that, once re-expressed in terms of $E[u]$ and central moments, yield bounds on $E[u]$. The first three are:

align[align omitted — 180 chars of source]

Expressing raw moments in terms of $E[u]$ and central moments $\mu_j=E[(u - E[u])^j]$, each condition $\det\bigl(\Delta_n^{(1)}\bigr) \geq 0$ becomes an $n$th degree polynomial inequality in $E[u]$ with coefficients that are polynomials in $\mu_2,\ldots,\mu_{2n-1}$. Equation (ref) states only that $E[u]\geq0$. We now state a skewness-based lower bound for $E[u]$, obtained by solving (ref) for $E[u]$.

prop\begin{align} E[u] \geq \frac{-\mu_3 + \sqrt{\mu_3^2 + 4\mu_2^3}}{2\mu_2} =\frac{\sigma}{2}\Bigl(-\gamma + \sqrt{\gamma^2 + 4}\Bigr). \end{align} where $\sigma = \sqrt{\mu_2}$ denotes the standard deviation and $\gamma=\mu_3/\mu_2^{3/2}$ denotes the skewness.
proofExpanding (ref) gives a quadratic in $E[u]$: $0 \le \mu_2(E[u])^2+\mu_3E[u]-\mu_2^2$. Since $\mu_2>0$ and $E[u]\ge 0$ by (ref), this holds if and only if $E[u]$ is at least the larger root. \qedhere

The right-hand side of (ref) is nonnegative and strictly decreasing in $\gamma$, converging to $0$ as $\gamma\to\infty$.

If one is prepared to assume that $u$ is negatively skewed, then the proposition implies that the standard deviation provides a lower bound on the mean deviation. This is useful for inferring bounds on mean inefficiency from reported empirical results on heterogeneity in firm cost or productivity.\footnote{For example, syverson2004market reports a 0.34 standard deviation in TFP of plants in high construction markets, with a pictured probability density distribution that shows no substantial right-skewness (pages 1184 and 1185). The bound then allows us to reasonably conclude that mean inefficiency is at least 0.34. In contrast, foster2008reallocation, who report a 0.22 mean standard deviation of TFP across a large set of industries, provide no information on higher moments that might allow us to bound mean inefficiency.} Specifically, if $\gamma\leq 0$ then

align*[align* omitted — 96 chars of source]

Expanding (ref) gives a cubic inequality in $E[u]$.\footnote{The full expansion is

align*[align* omitted — 209 chars of source]

} If $\mu_3\leq 0$ and $\mu_5\leq 0$ then $$ E[u]\geq\sqrt{\frac{\mu_4}{\mu_2}}=\sigma \times \sqrt{\kappa}, $$ where $\kappa= \mu_4/\mu_2^2$ denotes the kurtosis.

Higher-order determinants produce polynomial inequalities in $E[u]$ that may yield tighter bounds than the skewness-based bound in (ref). While we have clear intuition for variance and skewness, higher-order moments, such as the fifth central moment, are less transparent in their interpretation. Moreover, they are challenging to estimate accurately due to their sensitivity to outliers, which is why higher-order moments are rarely employed in practice.

Estimation

Estimating a frontier is challenging. Under certain tail behaviors, extreme value theory implies logarithmic convergence rates. For example, goldenshluger2004bndry assume normally distributed errors and show that the best-possible convergence rate for frontier estimation is logarithmic. Nonparametric methods for estimating the densities of $u$ and $v$ via deconvolution of (ref) have not been widely adopted in practice, as they require tuning parameters, involve substantial computational complexity, and converge slowly. Consequently, researchers often impose parametric assumptions to simplify estimation and improve convergence rates.

When $x$ is continuous, estimation becomes even more demanding, as it often requires smoothing in $x$ or restricting how the distributions of $u$ or $v$ vary with $x$. These restrictions can be problematic, especially when one wishes to allow for endogeneity. These challenges motivate two broad estimation strategies in the SFA literature: maximum likelihood aigner1977formulation and moment-based estimators olson1980monte. We adopt the latter in a panel-data setting, nonparametrically estimating the conditional mean outcome and second through fourth conditional central moments of deviations and errors. Appendix (ref) collects moment identities and sample formulas used for the panel estimator, and Appendix (ref) discusses related estimators under alternative data structures.

In earlier versions of the paper, we also implemented a cross-sectional variant (Appendix (ref)). In Monte Carlo simulations and in the empirical application, the cross-sectional moment-matching objective often exhibited multiple well-fitting local optima. With only second through fourth central moments, tail thickness can trade off between the deviation and the random error, so different decompositions can fit the data similarly while attributing heavy tails to different components. This is consistent with parmeter2024inference, who show in cross-sectional Monte Carlo simulations that specification tests can have empirical rejection rates that deviate substantially from nominal size in finite samples, even under correct parametric specification and when optimization is initialized at the true parameter values.

Consider the model (ref)--(ref) with panel data $\{y_{it},x_{it}\}$ and time-invariant deviations:

align[align omitted — 145 chars of source]

Suppose Assumption (ref) holds and let $x_i=(x_{i1},\ldots,x_{iT_i})'$. Our identification strategy relies on decomposing the variation in $y_{it}$ into “within” and “between” components to separate the central moments of the random error $v_{it}$ from the time-invariant deviation $u_i$. Define the population centered residuals as $\varepsilon_{it} = y_{it} - E[y_{it} \mid x_i]=-(u_i-E[u_i\mid x_i])+v_{it}$. We decompose these residuals into:

align[align omitted — 293 chars of source]

where $\bar v_i=T_i^{-1}\sum_{t=1}^{T_i} v_{it}$ is the time-averaged error. Equation (ref) shows that the within variation depends only on $v$, allowing identification of the conditional error central moments $\mu_{k,v}(x_i)$.\footnote{For a random variable $w$, let $\mu_{k,w}(z)=E[(w-E[w\mid z])^k\mid z]$ denote the $k$th central moment conditional on $z$.} Equation (ref) contains $u$ and the averaged $v$, so that once the central moments of $v$ are identified, they can be subtracted from the central moments of $\bar \varepsilon_i$ to recover the central moments of $u$.

Estimation proceeds in three stages. First, we estimate the conditional expectation $E[y_{it} \mid x_i]$ using a flexible nonparametric regression (where $ E[v_{it}\mid x_i]=0$ by Assumption (ref)(i)). This yields the sample residuals $\hat \varepsilon_{it}$, which we decompose into sample within residuals $\hat \varepsilon_{it}^w$ and sample between residuals $\bar{\hat \varepsilon}_i$ (analogous to (ref) and (ref)).

Second, we estimate the conditional central moments of the unobserved components. Using Assumption (ref)(iii), we relate the empirical central moments of the within residuals $\hat \varepsilon_{it}^w$ to the error central moments and solve for $\widehat\mu_{k,v}( x_i)$. Then, using Assumption (ref)(ii), we construct bias-corrected powers of the between residuals $\widehat u_i^k$. We construct these by taking the powers $(\bar{\hat \varepsilon}_i)^k$ and removing the contribution of the error term $\bar v_i$ using the estimates $\widehat\mu_{k,v}( x_i)$. Following the literature on variance estimation under heteroskedasticity hall1989variance,fan1998efficient, we regress $\widehat u_i^k$ on $ x_i$ to obtain smoothed estimates of the deviation central moments $\widehat\mu_{k,u}( x_i)$.\footnote{The explicit formulas for the moment conditions and bias-corrected powers are provided in Appendix (ref) BenMosheGenesoveMoments.}

In the final stage, we specify a parametric family for the deviation distribution $u\mid x$, and estimate parameters by the method of moments. We match model-implied central moments to the second-stage estimates, both unconstrained and subject to the near-frontier mass constraint described next.

Assumption (ref) implies that every neighborhood $[0,\delta]$ has positive conditional probability mass:

align[align omitted — 118 chars of source]

The assumption does not require a density at zero, nor does it control how quickly $F_{u\mid x}(\delta)$ shrinks as $\delta\downarrow 0$ (see Appendix (ref)). To prevent our estimation from admitting distributions with arbitrarily thin left tails near zero, we regularize by imposing a finite-sample minimum near-frontier mass constraint. For chosen constants $c>0$ and $m_0>0$, we estimate $\widehat \theta(x)$ by

align[align omitted — 311 chars of source]

where $\widehat \sigma_u(x)=(\widehat \mu_{2,u}(x))^{1/2},$ $\mu_k(x;\theta)$ denotes the $k$th conditional central moment of $u\mid x$ under $\theta$, and $F_{u\mid x;\theta}$ denotes the corresponding conditional CDF.\footnote{Equivalently, the constraint can be written as $Q_{u\mid x;\theta}(m_0/\widehat n_{\mathrm{eff}}(x))\le c\cdot \widehat\sigma_u(x)$, where $Q$ is the conditional quantile function.} We then compute $\widehat E[u \mid x]= E[u\mid x;\widehat\theta(x)]$ and estimate the FSF by $\widehat g(x_{it}) = \widehat E[y_{it} \mid x_i] + \widehat E[u \mid x_i]$.

Central moments are translation-invariant: $u$ and $u+\alpha$ have identical central moments for any constant $\alpha$. In the population, nonnegativity and assignment at the frontier ($u\ge 0$ and $0\in \text{Support}(u\mid x)$) distinguish $u$ and $u+\alpha$ for $\alpha>0$, because $u$ has mass arbitrarily near zero whereas $u+\alpha$ does not. In finite samples, however, when data near the frontier are sparse, this distinction can be weak: neither a sample from $u$ nor a sample from $u+\alpha$ may contain observations near zero, even though their means differ by $\alpha$.\footnote{For example, let $u \sim q\cdot\mathrm{Beta}(a,b)$, where $(a,b)$ determine the shape and $q$ scales the support. When observations near zero are sparse, different values of $q$ can produce distributions with similar shapes over the region where data are observed, while implying very different support widths and hence very different means.} As a result, the moment-matching objective (ref) can admit parameter values that fit the estimated central moments similarly while implying very different means. The estimator can therefore fit the observed shape well while remaining weakly informative about the location of that shape relative to zero. The constraint (ref) restores this finite-sample distinction by requiring mass near zero, preventing the fitted distribution from drifting away from the frontier. Indeed, as illustrated in the empirical application (Figure (ref)), unconstrained and constrained fits can have similar overall shape while implying very different means, because the unconstrained fit places much less mass near the frontier.\footnote{The estimator also admits a quasi-Bayesian interpretation. Moment-based estimators can be viewed as maximizing a quasi-likelihood derived from the moment distance chernozhukov2003mcmc. The constraint is equivalent to specifying a prior with restricted support that assigns zero mass to the set of parameter values violating the boundary restriction.} Thus, (ref) should be understood as finite-sample regularization, not an exogeneity assumption; it does not assume that a conditional quantile of the deviation is constant across inputs.\footnote{Some approaches estimate the frontier as an extreme conditional quantile of the outcome given inputs daouia2007nonparametric,daouia2010frontier. cazals2016nonparametric develop a nonparametric instrumental variable version to deal with endogeneity.}

Using $c\cdot \widehat\sigma_u(x)$ defines the neighborhood in local standard-deviation units, so the restriction adapts to the local dispersion of $u\mid x$. The effective sample size at $x_i$ is $\widehat n_{\mathrm{eff}}(x_i):=1/\sum_{j=1}^n w_{ij}^2$ kish1965survey, where $w_{ij}=K(\|x_i-x_j\|/h)\big/\sum_{r=1}^n K(\|x_i-x_r\|/h)$, $K(\cdot)$ is a nonnegative kernel function with weights normalized to sum to one, and $h$ is a bandwidth. This definition implies $\widehat n_{\mathrm{eff}}(x_i)=n$ under uniform weights ($w_{ij}=1/n$), and $\widehat n_{\mathrm{eff}}(x_i)$ decreases as weights become more concentrated on observations near $x_i$.

The threshold $m_0/\widehat n_{\mathrm{eff}}(x)$ scales with the amount of information available for learning near-zero behavior at $x$. The constraint can be rewritten as a conditional moment inequality,

align*[align* omitted — 118 chars of source]

so that $m_0$ is the required expected number of effective observations within the neighborhood $[0,c\cdot \widehat\sigma_u(x)]$. For example, if $\widehat n_{\mathrm{eff}}=25$ and $m_0=1$ then $m_0/\widehat n_{\mathrm{eff}}=0.04$, so the constraint requires at least 4% probability mass within $[0,c\cdot \widehat\sigma_u]$, corresponding, on average, to one effective observation in that neighborhood. If $\widehat n_{\mathrm{eff}}=100$ and $m_0=1$ then $m_0/\widehat n_{\mathrm{eff}}=0.01$, yielding a weaker restriction.

Relative to cross-sectional strategies that rely on asymmetry of $u$ and symmetry of $v$ for identification, our approach exploits the panel structure: within-variation identifies the random error, while between-variation identifies the deviation. Consequently, we do not impose any symmetry restriction on $v$; provided the relevant moments exist, its distribution is otherwise left unspecified. The distribution of $u$ may be left-skewed, right-skewed, or symmetric, unlike standard SFA specifications (e.g., half-normal, truncated normal, exponential, or gamma), which impose right skewness greene2008econometric.

Monte Carlo Simulations

We report Monte Carlo simulations to assess the finite-sample performance of the lower bound and point estimators of $E[u]$. Two finite-sample issues arise. First, in small samples the estimated lower bound can exceed the estimated mean, so that $\widehat{LB}>\widehat{E}[u]$. Second, the estimator of $E[u]$ can be highly inaccurate even when $\widehat{LB}$ remains accurate. Motivated by the ill-posed mapping from central moments to mean deviation, we study whether imposing a near-frontier mass constraint can regularize this mapping and improve estimation of $E[u]$.

We simulate a panel with time-invariant deviations, \[ y_{it} = g - u_i + v_{it}, \qquad i = 1,\ldots,n,\; t = 1,\ldots,T, \] where the random errors are normally distributed, $\{v_{it}\}_{i=1,t=1}^{n,T} \overset{\text{iid}}{\sim} \mathrm{N}(0, \sigma_v^2)$. The deviations are generated via a rejection sampling mechanism: we draw candidate deviations $\tilde{u} \sim q\cdot\mathrm{Beta}(a,b)$ and retain them with probability $1-p$ if they fall below the $f$-th quantile and with probability 1 otherwise, repeating until $n$ deviations $\{u_i\}_{i=1}^n$ are obtained. We set $g=5$, $q=4$, and $T=8$, and the scarcity parameters to $f=0.05$ and $p=0.95$. The shape parameters $(a,b)$ vary over an equally spaced $80\times 80$ grid in $(\log a,\log b)\in[-2,2]^2$. For each configuration $(a,b)$ on the grid and each $n\in\{25,250,2\,500\}$, we generate 150 independent Monte Carlo replications of panel datasets, each of size $nT$. For each design, we set $\sigma_v$ equal to the standard deviation of $u$ from the rejection sampling mechanism.

We treat the outcomes $\{y_{it}\}_{i=1,t=1}^{n,T}$ as the only observed variables; the deviations $\{u_i\}_{i=1}^n$ and errors $\{v_{it}\}_{i=1,t=1}^{n,T}$ are treated as unobserved. For each simulated dataset, we estimate the deviation central moments $\mu_{2,u}$, $\mu_{3,u}$, $\mu_{4,u}$, the lower bound $LB$, and the mean deviation $E[u]$ using the three-step procedure in Section (ref). First we compute residuals and decompose them into within- and between-firm components:

align*[align* omitted — 287 chars of source]

Second, we estimate the central moments of $u$ using the closed-form estimators in (ref)--(ref) in Appendix (ref). We then estimate the lower bound by the sample analog of (ref),

align*[align* omitted — 123 chars of source]

where $\widehat{\sigma}_u = (\widehat{\mu}_{2,u})^{1/2}$ and $\widehat{\gamma}_u = {\widehat{\mu}_{3,u}}/{\widehat{\sigma}_u^3}$.

Third, to estimate the mean deviation $\widehat{E}[u]$, we fit the parameters of a $q \cdot \mathrm{Beta}(a,b)$ distribution. We estimate $(\widehat a,\widehat b,\widehat q)$ by matching the estimated central moments to the model-implied central moments while imposing the near-frontier mass constraint. With $m_0=1$ and $c\in \{0.5,1.00, \infty\}$, we solve

align*[align* omitted — 263 chars of source]

where $F_{\mathrm{Beta}(a,b)}$ is the CDF of a $\mathrm{Beta}(a,b)$ random variable on $[0,1]$. For a fixed neighborhood width $c\cdot \widehat\sigma_u$, the constraint requires at least $m_0$ expected observations in this neighborhood. The unconstrained estimator corresponds to $c=\infty$. We then compute the implied mean, $\widehat{E}[u] = \widehat q \cdot \widehat a / (\widehat a + \widehat b)$.

We partition the parameter grid into four regions defined by the qualitative shape of the distribution of $u$. The high near-frontier mass region ($a<1$) corresponds to densities that are high near zero. The unimodal right-skewed ($a\geq 1, b\geq 1, b \geq a$) and unimodal left-skewed ($a\geq 1, b\geq 1, b < a$) regions correspond to bell-shaped densities that are right-skewed (or symmetric) and left-skewed, respectively. The low near-frontier mass region ($a\ge 1,b<1$) corresponds to densities with little mass near zero, rising toward the upper end of the support.

Figure (ref) and Table (ref) compare the unconstrained estimator $\widehat{E}[u]$ with its lower bound estimator $\widehat{LB}$. The top row of Figure (ref) shows, for each $(a,b)$, the fraction of the 150 Monte Carlo replications in which $\widehat{LB}>\widehat{E}[u]$, a violation of the population inequality (ref). Table (ref) shows the fractions averaged across grid points within each region: for $n=25$, 5.70% in high near-frontier mass, 3.49% in low near-frontier mass, 0.74% in unimodal right-skewed, and 1.19% in unimodal left-skewed. As $n$ increases to $250$, these violations virtually disappear.

figure[figure omitted — 2,432 chars of source]
table[table omitted — 1,875 chars of source]

The middle and bottom rows of Figure (ref) show, for each $(a,b)$, the median relative absolute errors ${|\widehat{E}[u]-E[u]|}/{E[u]}$ and ${|\widehat{LB}-LB|}/{LB}$. Table (ref), averaging across grid points within each region, shows that the median relative absolute error of $\widehat{E}[u]$ is large at $n=25$ and varies substantially by region: 27.5% in high near-frontier mass, 38.0% in unimodal right-skewed, 56.6% in unimodal left-skewed, and 64.4% in low near-frontier mass. In contrast, the lower-bound estimator is more accurate and stable across regions (20.5%, 15.2%, 15.3%, and 20.3%). As $n$ increases, the median error of $\widehat{LB}$ declines in all regions (to between 1.7% and 3.7% at $n=2{,}500$), while the median error of $\widehat{E}[u]$ remains large in parts of the grid: at $n=2{,}500$, it is 4.0% in the high near-frontier mass region but 60.3% in the low near-frontier mass region. These results suggest that $\widehat{LB}$, which relies only on skewness and variance, offers a robust alternative to point estimation of $E[u]$.

The lower bound remains stable because it does not require recovering the support width or detailed distributional shape near zero; it depends only on the second and third moments. Next, we examine the near-frontier mass constraint as regularization for mapping estimated second through fourth central moments to $E[u]$. Figure (ref) shows $\log_{10}(\mathrm{MSE})$ of $\widehat{E}[u]$ over the parameter grid for the unconstrained estimator and the two constrained estimators $(m_0,c)=(1,1)$ and $(m_0,c)=(1,0.5)$. Table (ref) reports median bias, median MSE, and mean MSE by region.

For $n=25$, the unconstrained mapping performs reasonably over much of the grid but can generate very large errors in some regions, especially the unimodal left-skewed region. This instability reflects noisy estimation of the central moments. When the distribution is roughly symmetric and tightly concentrated at an interior mode, skewness is close to zero and the second and fourth moments constrain the shape of the distribution much more than its location relative to zero. Small sampling errors in the estimated moments can then be matched by fitted distributions with substantially different support widths $q$, translating into large dispersion in the implied mean. This is visible in Figure (ref), where the unconstrained estimator has very high MSE in parts of the grid, and in Table (ref), where it has much larger mean MSE than median MSE in several regions.

figure[figure omitted — 2,272 chars of source]
table[table omitted — 2,594 chars of source]

As the sample size increases and the moments are estimated more precisely, these failures become concentrated in regions where the moment-to-mean mapping is ill-conditioned: specifically, where the distribution of $u$ places almost all probability far from the frontier. The mean $E[u]$ depends on both the shape of the distribution and the support width $q$. When most probability is concentrated away from zero, the second through fourth moments mainly reflect the shape of the upper tail and provide limited information about $q$. The parametric beta family does not resolve this ambiguity because the induced scarcity mechanism implies that the realized distribution of $u$ is not exactly beta. As a result, the mapping can match the observed moments by shifting the fitted distribution along its support, increasing $q$ while preserving its shape, with little penalty, since there are no observations near the frontier to discipline the fit. In Figure (ref), this appears as regions where the unconstrained estimator has very large MSE even at $n=2{,}500$.

The unconstrained estimator has the smallest bias in magnitude across regions and sample sizes (Table (ref), Median Bias column; see also Table (ref), Median $|\text{Bias}|$ column). However, its MSE can be orders of magnitude larger. Regularization mitigates these failures by requiring the fitted distribution to place a minimal probability mass in a neighborhood of zero. The constraint binds mainly where near-frontier mass is low and rarely where it is high (see Table (ref) in Appendix (ref)). For example, under $(m_0,c)=(1,1)$ at $n=2{,}500$, the bind share ranges from 0 in high near-frontier mass to 0.466 in low near-frontier mass. In regions where the constraint rarely binds, the constrained and unconstrained estimators have nearly identical median bias and median MSE. Where it binds, it reduces the extreme failures of the unconstrained mapping and sharply lowers mean MSE. At the same time, the constrained estimators generally exhibit more negative bias, especially in regions with little near-frontier mass. Thus the near-frontier mass constraint stabilizes point estimation at the cost of bias in regions where the true near-frontier mass falls short of the imposed threshold. The lower bound, which does not require recovering support width or distributional shape near zero, remains stable even where point estimation is weakly informative.

Empirical Application to Production

Consider the log-production model

align[align omitted — 95 chars of source]

where $y$ is log output, $k$ is log capital, $l$ is log labor, and $g(\cdot)$ is strictly increasing in each input.\footnote{If $g(\cdot)$ is CES, then one of $\omega_k$, $\omega_l$ or $\omega_y$ is redundant.} The unobservables $\omega_k$ and $\omega_l$ represent inefficient applications of capital and labor, respectively, while $\omega_y$ is a Hicks-neutral shock.\footnote{The Hicks-neutral shock $\omega_y$ can represent the effect of an unobserved input.} These three inefficiencies all have support $[0, \infty)$. The total inefficiency in production is $u = g(k, l ) - [g(k - \omega_k, l - \omega_l) - \omega_y] \geq 0$, with $u=0$ if and only if the firm is efficient. The random error $v$ represents measurement error in $y$, including unpredictable random productivity shocks, and is assumed to be mean independent of $(k,l)$.

Endogenous inputs are a pervasive challenge in estimating production functions. The most widely used approaches for estimating the conditional mean production function combine panel data with control functions, using a proxy variable (e.g., investment or intermediate inputs) that is monotone in unobserved productivity and exploiting timing and information-set assumptions to control for simultaneity olleypakes,levinsohn2003estimating,ackerberg2015identification. In contrast, our frontier approach replaces proxy-variable and timing assumptions with bounded-above productivity ($u\geq 0$) and assignment at the frontier ($0 \in \textup{Support}(u)$) assumptions, which together identify the frontier relationship without instruments or proxies. Classical SFA methods also use the idea of a frontier, but typically proceed by imposing parametric structure on the frontier and distributional assumptions on inefficiency and random error; accommodating endogeneity is generally thought to require additional restrictions (e.g., controls or instruments).

To illustrate our approach, we use plant-level panel data and treat inefficiencies as time-invariant, as in (ref)--(ref):

align*[align* omitted — 133 chars of source]

In model (ref) described above, inefficiency enters as $g(k - \omega_k, l - \omega_l)$ so under a flexible $g(\cdot)$, the implied output loss generally depends on contemporaneous $(k,l)$. This would make the inefficiency component time-varying. Accordingly, in the application we restrict inefficiency to be Hicks-neutral ($\omega_k=\omega_l=0$). Following Mundlak1978, we further restrict the dependence between inputs and unobservables in Assumption (ref) to operate through $\bar x_i = T_i^{-1}\sum_{t=1}^{T_i} x_{it}$ (the time-averaged inputs): (i) $E[v_{it}\mid x_{it},\bar x_i]=0$; (ii) $ \mu_{k,v}(x_{it},\bar x_i) = \mu_{k,v}(\bar x_i)$, for $k\in\{2,3,4\}$; (iii) $u_i \mid (x_{i1},\ldots,x_{iT_i}) \overset{d}{=} u_i \mid \bar x_i $; (iv) $u_i \mathrel{\perp\!\!\!\perp} (v_{i1},\ldots,v_{iT_i}) \mid \bar x_i.$

These restrictions reduce the conditioning set from the full input history $(x_{i1},\ldots,x_{iT_i})$ to the low-dimensional summary $\bar x_i$, making nonparametric estimation of the conditional moments feasible. In principle, the identification strategy applies with conditioning on the full history; the Mundlak restriction is a practical dimension reduction. In the repeated measurement model $y_{it}=g(x_i)-u_i+v_{it}$ of Appendix (ref), neither restriction is needed: because inputs do not vary within plant, inefficiency $g(k_i,l_i)-g(k_{i}-\omega_{k},l_{i}-\omega_{l})+\omega_{y}$ is time-invariant without imposing $\omega_k=\omega_l=0$, and $\bar x_i = x_i$ by construction, so the Mundlak assumption imposes no additional restriction beyond conditioning on $x_i$ itself.

Our data are plant-level observations from the Colombian manufacturing census for 1981--1991. We focus on the food products industry (ISIC 311), using the colombian dataset distributed with the gnrprod R package.\footnote{See gandhi2020identification for discussion of this dataset and its use in production-function applications.} This application is intended to illustrate the method in a standard production setting, rather than to deliver a definitive structural account of Colombian manufacturing.

We set output to log real gross output ($y=\texttt{RGO}$), capital to log real capital stock ($k=\texttt{K}$), and labor to log labor input measured in employee-years ($l=\texttt{L}$), where all variables are expressed in logs and output and capital are deflated to constant prices. After excluding observations with fewer than 10 employee-years and then restricting the sample to plants observed for at least 8 years to ensure sufficient within-plant variation, we obtain an unbalanced panel of $n=408$ plants and $4,306$ plant-year observations, with an average duration of $10.6$ years per plant; $24$ plants are observed for $8$ years, $29$ for $9$ years, $52$ for $10$ years, and $303$ plants are present for the full $11$ years. In the sample, the mean log output is $9.21$ (s.d. $1.73$), with right-skewness of $0.20$ and kurtosis of $5.28$. Mean log labor is $4.05$ (s.d. $1.21$) and mean log capital is $7.37$ (s.d. $1.92$).

For mean regression in such a panel setting, a fixed effects (within) estimator would be natural, since differencing removes the time-invariant inefficiency term and thereby eliminates bias from correlation between inefficiency and inputs. Our interest, however, lies in the frontier relationship. Under assignment at the frontier, correlation between inefficiency and inputs is not a concern for identification of the frontier. Fixed effects therefore addresses a problem that is absent under our identifying assumptions, while discarding the between-plant differences that are precisely the most informative about which plants operate closest to the frontier.

In production applications, fixed effects estimation also has a long record of producing implausible estimates, particularly for the capital coefficient, which can be attenuated toward zero griliches1998production. One reason for this could be the limited within-plant variation in capital relative to between-plant variation; in our data, within-plant variance accounts for about 5% of total capital variance and 9% for labor. Moreover, consistency of the fixed effects estimator requires a strict exogeneity condition that productivity shocks are mean independent of the entire history of inputs. This rules out feedback from past productivity shocks to future input choices. By contrast, we maintain a weaker mean independence restriction that conditions only on contemporaneous inputs and their time averages, allowing for more general forms of dynamic adjustment in input choices.

We implement the three-stage estimator developed in Section (ref) with the Mundlak assumptions above as follows. First, we nonparametrically estimate the conditional expectation $E[y_{it}\mid x_{it},\bar x_i]$ using a tensor-product thin-plate regression spline basis. To separate within-plant and between-plant variation, the specification includes main effects and interactions in the time-averaged inputs $\bar{x}_i$, as well as main effects and interactions in the demeaned inputs $x_{it}-\bar{x}_i$. We also include interactions between the within and between components. This provides a flexible first-stage approximation, while nesting the additive structure implied by the maintained model. We set each marginal basis dimension on the order of $n^{1/2},$ where $n$ is the number of plant-year observations, yielding a total of 132 basis functions across terms, and estimate the resulting high-dimensional linear model by ridge regression. The fitted values yield centered residuals $\hat \varepsilon_{it}$, which we decompose into within-plant residuals $\hat \varepsilon_{it}^w$ and plant-level residuals $\bar{\hat \varepsilon}_i$ (the sample analogs of (ref) and (ref)).

Second, we estimate conditional central moments using the within-between decomposition. Using empirical central moments of $\hat \varepsilon_{it}^w$, we obtain plant-level error central moments $\hat\mu_{k,v}(\bar x_i)$ using (ref)--(ref), and smooth these as functions of $\bar x_i$ using ridge regression on the spline basis in $\bar x_i$. We then construct adjusted plant-level powers $\hat u_i^k$ using (ref)--(ref) in Appendix (ref) and obtain smoothed estimates of $\hat\mu_{k,u}(\bar x_i)$ for $k\in\{2,3,4\}$ by estimating these adjusted powers as functions of $\bar x_i$ by ridge regression.

Finally, we estimate the inefficiency distribution by matching the estimated central moments to model-implied central moments. For each plant, we consider two parametric families for $u\mid \bar x$: the scaled beta (as in our simulations) and truncated normal distributions (a generalization of the half-normal used in the original SFA literature, and often used in that literature now). We compute $\hat\theta(\bar x)$ by minimizing the distance between estimated and model-implied central moments subject to the near-frontier mass restriction in (ref). We set the constraint parameters to $(m_0,c)=(1,0.5)$, which ensures that the expected number of plants within $0.5$ standard deviations of the frontier is at least one. In the input-dependent case, $\widehat n_{\mathrm{eff}}(\bar x)$ is computed using kernel weights (bandwidth $h=0.20$), yielding an average effective sample size of about $30$, and a standard deviation of $15$ across plants. Consequently, the required probability mass $p_0(\bar x) = 1/\widehat n_{\mathrm{eff}}(\bar x)$ adapts to the local data density, averaging approximately $0.03$, which provides stronger regularization in regions where the effective sample size is small.

We first consider the case where the distributions of $u_{i}$ and $v_{it}$ are independent of inputs. This corresponds to the Corrected Ordinary Least Squares (COLS) estimator, where the conditional central moments reduce to constants (and $g(x)=E[y\mid x]+E[u]$). Denote the estimate of the negative plant-level mean centered residual by $\hat{\eta}_i = -\bar{\hat{\varepsilon}}_i = \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{(u_i - E[u]) - \bar{v}_i}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{(u_i - E[u]) - \bar{v}_i}{\tmpbox} $ (an estimate of centered inefficiency contaminated by plant-averaged error). Figure (ref) plots the kernel density of the shifted estimate $\hat{\eta}_i-\min_j(\hat{\eta}_j)$, which sets the minimum value to zero to visualize the support relative to zero.

We estimate the central moments of the inefficiency $u$ by subtracting the estimated error central moments from the central moments of the between residuals (see (ref)--(ref) in Appendix (ref)). The resulting central moment estimates are $\widehat\mu_{2,u} \approx 0.59$, $\widehat\mu_{3,u} \approx 0$, and $\widehat\mu_{4,u} \approx 1.09$. These imply a skewness of $0$ and kurtosis of $3.13$. The small skewness and kurtosis close to 3 are consistent with an approximately symmetric distribution.

We fit a scaled beta distribution $q\cdot\mathrm{Beta}(a,b)$ and a truncated normal distribution $TN(\mu,\sigma^2)$ on $[0,\infty)$ to these estimated central moments. The unconstrained fitted densities are shown as dashed lines in Figure (ref). These fitted densities reproduce the overall shape of the empirical kernel density (shifted so minimum value is zero) but imply implausibly large mean inefficiency: $\widehat E[u]\approx 20.87$ (beta) and $\widehat E[u]\approx 4.80$ (truncated normal).

This motivates the near-frontier mass restriction (ref) as a regularizer. The solid lines in the figures impose $(m_0,c)=(1,0.5)$, ensuring that the expected number of plants within approximately $ \widehat\sigma_u/2 \approx 0.38$ log points of the frontier is at least one. Equivalently, with $\widehat n_{\mathrm{eff}}=n=408$ this requires at least $p_0=m_0/n=1/408\approx 0.002$ probability mass within $0.5$ standard deviations of zero, ruling out fits that place essentially no mass near the frontier. This results in $\widehat E[u]\approx 2.30$ (beta) and $\widehat E[u]\approx 2.53$ (truncated normal).

figure[figure omitted — 419 chars of source]

We now turn to estimation allowing the distributions of $u$ and $v$ to depend on inputs through the Mundlak controls $\bar x_i$. Following the within--between decomposition described above, we obtain plant-level estimates of the conditional central moments of both components. Across plants, the estimated error second central moment has mean $0.15$ (s.d.\ $0.08$), the third central moment has mean $0.02$ (s.d.\ $0.01$), and the fourth central moment has mean $0.21$ (s.d.\ $0.11$). The estimated inefficiency second central moment has mean $0.61$ (s.d.\ $0.24$), the third central moment has mean $0.00$ (s.d.\ $0.01$), and the fourth central moment has mean $1.08$ (s.d.\ $0.29$).

We report the coefficients from regressions of estimated mean inefficiency on plant-level mean inputs (Table (ref), Panel A). The coefficient on labor is $-0.39$ (beta) and $-0.39$ (truncated normal), and the coefficient on capital is $0.20$ (beta) and $0.14$ (truncated normal). The positive coefficient on capital suggests that estimated mean inefficiency is correlated with capital, in contrast with the control-function approach of olleypakes, which assumes that the proxy variable (e.g., investment) is strictly increasing in productivity conditional on capital. Under this strict monotonicity, persistently more productive plants will have greater capital.

table[table omitted — 1,697 chars of source]

We next compare production-function elasticity estimates across procedures (Table (ref), Panel B). For the mean production function, OLS (within) regresses demeaned output on demeaned inputs, yielding elasticities of $0.27$ on labor and $0.20$ on capital, while the Levinsohn--Petrin estimator yields $0.12$ on labor and $0.15$ on capital. For the frontier, the parametric SFA model with truncated normal time-invariant inefficiency yields elasticities of $0.32$ on labor and $0.25$ on capital.\footnote{Under the OLS and SFA assumptions that inefficiency and error are mean independent of inputs ($E[u_i\mid x_{i1},\ldots,x_{iT}]=E[u_i]$ and $E[v_{it}\mid x_{i1},\ldots,x_{iT}]=0$), we have $E[y_{it}\mid x_{i1},\ldots,x_{iT}]=g(x_{it})-E[u_i]$, so the frontier differs from the conditional mean by a constant shift, and one would expect identical slope coefficients across mean and frontier regressions. The estimates differ in Table (ref) because OLS identifies $\beta$ from within-plant variation only, whereas the panel SFA estimator exploits total variation and estimates parameters by maximum likelihood. Using comparable estimators, cross-sectional SFA yields slope coefficients and intercepts nearly identical to cross-sectional OLS, implying negligible estimated inefficiency. This is consistent with diagnostic tests suggesting that standard SFA is unable to distinguish inefficiency from error in this setting.} Using our moment-based approach, we estimate $\widehat g(\bar x_i)=\widehat E[y\mid \bar x_i]+\widehat E[u_i\mid \bar x_i]$ and obtain frontier elasticities by regressing $\widehat g$ on labor and capital. The estimated frontier elasticities are $0.21$ (beta) and $0.20$ (truncated normal) for labor, and $0.61$ (beta) and $0.56$ (truncated normal) for capital. The MM frontier estimates imply decreasing returns to scale, with coefficient sums of 0.82 (beta) and 0.76 (truncated normal), which are substantially larger than the corresponding OLS and Levinsohn–Petrin estimates. The elasticity--cost share ratio (E/C Ratio) compares the relative labor--capital elasticity to the corresponding relative input usage; under neoclassical price-taking behavior with unrestricted input choice, this ratio equals one. In our estimates, the implied ratio ranges from $1.15$ (MM (beta)) and $1.22$ (MM (truncated normal)) to $4.62$ (OLS).

Finally, we report the unconditional lower bound and mean inefficiency estimates. The estimated lower bound on mean inefficiency is $0.76$. Our MM estimates of mean inefficiency (averaged over plants) are $1.43$ (beta) and $1.67$ (truncated normal). For comparison, the SFA model (assuming time-invariant inefficiency) yields a higher mean inefficiency of $2.75$, while COLS yields $2.30$ (beta) and $2.53$ (truncated normal). Mean inefficiency estimates are sensitive to whether one allows $(u_i,v_{it})$ to depend on inputs. In our data, the MM estimates that allow such dependence are lower than input-independent benchmarks, consistent with the possibility that input-independent methods overestimate inefficiency by attributing heterogeneity or input--error correlations to technical inefficiency.

To visualize the estimated structural frontier $g(x)$ against the data, Figure (ref) plots slices of the production surface. Each slice shows output against one input, restricting the other input to lie within a bandwidth around its sample median. The plotted curves represent smoothed estimates of the frontier, conditional mean, and 95% quantile regression. For example, Panel (a) plots log output against log capital for plants with log labor near the sample median. The slices show a broadly concave frontier consistent with diminishing marginal returns to capital. The estimated frontier lies near the 95% conditional percentile and is slightly more concave than the mean estimates.

\enlargethispage{\baselineskip}

figure[figure omitted — 797 chars of source]

\FloatBarrier

Conclusion

This paper identifies the frontier structural function from the supremum of the outcome given inputs, assuming a nonnegative deviation and assignment at the frontier (that zero lies in the deviation’s support given inputs). This identification holds even when inputs are endogenous, thereby obviating the need for instrumental variables. We then allow for random error that is mean independent of inputs. Using only observed inputs and outcomes, we develop estimators of the frontier structural function and mean deviation based on conditional central moments. We regularize the moment-matching step by imposing a near-frontier mass constraint that the fitted deviation distribution retain a minimum amount of probability mass near the frontier. We also derive a skewness-based lower bound on the mean deviation that is robust to scarcity of data near the frontier. Monte Carlo simulations show that the near-frontier mass constraint substantially improves finite-sample performance by reducing mean squared error, while the skewness-based lower bound remains accurate even when near-frontier observations are scarce.

In an application to the Colombian food products industry, estimated mean inefficiency is correlated with inputs. There are a number of possible explanations for this pattern, including that plants choose inputs based on productivity, that firms face different input prices according to productivity, or that inefficiency in input choice is correlated with inefficiency in input application; distinguishing among these explanations would require additional information or stronger structure. The regularized estimator yields fitted distributions that capture the overall shape of the empirical distribution while satisfying the near-frontier mass constraint, and the skewness-based lower bound provides an optimization-free alternative to these point estimates.

\begingroup \setstretch{1.30} \endgroup