EconBase
← Back to paper

A Classifier-Lasso Approach for Estimating Production Functions with Latent Group Structures

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

68,483 characters · 11 sections · 70 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

A Classifier-Lasso Approach for Estimating Production Functions with Latent Group Structures

abstractI present a new estimation procedure for production functions with latent group structures. I consider production functions that are heterogeneous across groups but time-homogeneous within groups, and where the group membership of the firms is unknown. My estimation procedure is fully data-driven and embeds recent identification strategies from the production function literature into the classifier-Lasso. Simulation experiments demonstrate that firms are assigned to their correct latent group with probability close to one. I apply my estimation procedure to a panel of Chilean firms and find sizable differences in the estimates compared to the standard approach of classification by industry. JEL Classification: C01, C13, C30, C55\\ Keywords: Production Functions, Latent Group Structures, Classifier Lasso, Penalized GMM

\thispagestyle{empty}

\onehalfspacing \setcounter{page}{1}

Introduction

Production functions are one of the central concepts in economic theory, with origins dating back at least to the mid-18th century (see h1997 for the history of production functions). It is all the more remarkable that the identification of production functions is still an econometric challenge.\footnote{\textcites{gnr2020}{s2020} are two very recent examples that solve this identification problem for gross output production functions.} The main challenge is an endogeneity problem first mentioned by ma1944. It arises when firms make input decisions based on their productivity, while productivity is usually unobservable to the researcher. Since the seminal work of \textcites{op1996}{lp2003}, all recent identification strategies, such as those of \textcites{bb2000}{acf2015}{gnr2020}, use panel data along with assumptions about firm behavior to solve the endogeneity problem.

Although these identification strategies are frequently applied in practice, a remaining challenge is latent firm heterogeneity in the data. For instance, firms may differ in their production technologies, which can be thought of model parameter heterogeneity. Therefore, it is common practice to classify firms based on prior information, like an industry classification, and then estimate separate production functions. I will refer to this approach as ex-ante classification throughout the paper. However, as pointed out by gm1999, even in a narrowly defined industry where firms produce nearly identical products, these firms may use their inputs very differently, which in turn may indicate different production technologies. On the other hand, firms in different industries may nevertheless use very similar production technologies so that neglecting this similarity reduces the precision of the estimation.

To overcome these limitations, I propose a new estimation procedure. Using examples of three identification strategies, I show how to embed them into the classifier-Lasso (C-Lasso) of ssp2016. My estimation procedure has no additional data requirements, except that the panel must be sufficiently long, and can even be used in situations where prior information is not available at all. I demonstrate good finite-sample performance of my estimator in simulation experiments. Even in moderately long panels, my estimation procedure is able to assign all firms to the correct latent group with probabilities approaching one. The properties of my estimator are asymptotically identical to an infeasible estimator that knows and exploits the latent group structure. Further, because the iterative estimation algorithm proposed by ssp2016 is computationally intensive or even infeasible for panels with many firms, I present an extension to the algorithm that overcomes this problem. Finally, I apply my estimation procedure to a panel of Chilean firms and find that an ex-ante classification by industry does not match the data-driven classification of my estimator. In particular, I classify firms from five industries into three latent groups, where each of the five industries is composed of firms from all three latent groups. A simple comparison of the mean squared residuals, reveals that the estimation based on the data-driven classification fits the data better than the ex-post classification, although the latter estimates more parameters. Moreover, estimating separate production functions by industry is not only less efficient, but also yields significantly different estimates and thus leads to different conclusions.

Besides the C-Lasso of ssp2016, there are at least two alternative approaches to deal with latent group structures. The first approach is a finite mixture specification, which models the probability of belonging to a latent group as a function of the explanatory variables. Some applications of parametric and semi-parametric finite mixtures are \textcites{s2005}{ks2009}{bc2014}. The second approach is a modification of the k-means clustering algorithm. Examples are \textcites{ln2012}{bm2015}{sw2015}. Of the two alternative approaches, k-means clustering is probably the most popular. swj2019 name and discuss three major differences between C-Lasso and k-means clustering: i) C-Lasso has an additional tuning parameter, ii) k-means forces group assignment, and iii) k-means is computationally more demanding. There are pros and cons for i) and ii), so the main argument for C-Lasso is iii), which is particularly relevant for applications with many firms due to the large number of potential cluster partitions that have to be tried by k-means. Both alternative approaches have been recently adapted to the estimation of production functions. kss2017 use a finite mixture specification to extend the identification strategy of gnr2020. They also provide empirical evidence of additional firm heterogeneity within narrowly defined industries in a panel of Japanese firms. css2019 extend the k-means clustering algorithm to account for multidimensional heterogeneity and reassess the rise of markups reported by dleu2020. They find that the level and growth of markups are lower when additional firm heterogeneity within industries is taken into account. I contribute to this literature by presenting an alternative estimation procedure based on the C-Lasso and providing additional empirical evidence that classification by industry is not sufficient to fully account for firm heterogeneity.

The paper is organized as follows. I introduce production functions with latent group structures in Section (ref). I show how these production functions can be estimated in Section (ref). I analyze the performance of my estimation procedure in Section (ref). I provide an empirical illustration using Chilean panel data in Section (ref). Finally, I give some concluding remarks in Section (ref).

Throughout this paper, I follow conventional notation: scalars are represented in standard type, vectors and matrices in boldface, and all vectors are column vectors. Further, $\lVert \cdot \rVert$ denotes the Euclidean norm.

Production Functions with Latent Group Structures

Structural Panel Model

I consider panel data of $N$ firms observed for $T + 1$ periods $\{\xi_{it} \colon i \in \{1, \ldots, N\}, t \in \{0, \ldots, T\}\}$, where $\xi_{it} \coloneqq (Y_{it}, \mathbf{X}_{it}, \mathbf{Z}_{it})$ is a collection of variables for firm $i$ at time $t$. $Y_{it}$ is the output, $\mathbf{X}_{it}$ is a vector of inputs, like capital $K_{it}$, labor $L_{it}$, or intermediate inputs $M_{it}$, and $\mathbf{Z}_{it}$ is a vector of additional firm characteristics, like input and output prices or a firm's export status. I follow standard convention and denote the corresponding natural log values in lowercase letters, i.e. $y_{it}$ is the natural logarithm of $Y_{it}$.

To allow for heterogeneous production technologies, I follow \textcites{ln2012}{bm2015}{ssp2016} and assume a time-homogeneous latent group structure where each firm belongs to exactly one of $J^{0}$ groups. I collect all firms that are members of a latent group $j \in \{1, \ldots, J^{0}\}$ in a set $G_{j}^{0}$, where $\bigcup_{j = 1}^{J^{0}} G_{j}^{0} = \{1, \ldots, N\}$ and $G_{j}^{0} \cap G_{j^{\prime}}^{0} = \varnothing$ for all $j \neq j^{\prime}$. I assume that $J^{0}$ is known to the researcher, but the group membership of each firm is not. All firms in $G_{j}^{0}$ have the same well-defined production function

equation[equation omitted — 135 chars of source]

where $P(\mathbf{x}_{it})$ is a basis system of transformed input variables, $\boldsymbol{\beta}_{i}^{0} \coloneqq \sum_{j = 1}^{J^{0}} \operatorname{\mathbf{1}}\{i \in G_{j}^{0}\} \boldsymbol{\alpha}_{j}^{0}$ are the corresponding model parameters that follow a general group pattern, $\omega_{it}$ is a predictable productivity shock, unobservable to the researcher, and $\epsilon_{it}$ is an unanticipated productivity shock or a measurement error. Further, I assume that $\lVert \boldsymbol{\alpha}_{j}^{0} - \boldsymbol{\alpha}_{j^{\prime}}^{0} \rVert > 0$ for all $j \neq j^{\prime}$, i.e. firms in different latent groups use different production technologies.

Besides the unknown latent group membership of the firms, the estimation of (ref) poses further econometric challenges. Some important issues are related to the measurements of output and inputs, the functional form of the production function, and endogeneity problems (see abbp2007 and a2019 for more details).\footnote{It is known since ma1944 that when firms choose their inputs optimally, at least to some extent, $\omega_{it}$ and $\mathbf{x}_{it}$ are correlated, leading to inconsistent estimators of output elasticities (see gm1999).} In particular, these problems led to very different estimation strategies that can be readily applied if the researcher knows the latent group membership of the firms.\footnote{For instance, it is common practice to classify firms by an industry classification.} Traditional identification strategies rely on instrumental variables or fixed effects. Both strategies are not very popular in recent work, because the former requires additional excluded instruments, like output and input prices, and the latter requires the restrictive assumption $\omega_{it} = \omega_{i}$.\footnote{These traditional strategies can already be applied using estimators presented in Section 2 and 3 of ssp2016.} Modern identification strategies are based on modeling firm behavior in a dynamic environment. The key assumptions are related to the timing of firms' input decisions and the information available to firms at the time of these decisions. There are three popular estimation strategies: control function / proxy variable approach (op1996, lp2003, and acf2015), dynamic panel estimator (ab1991, bb1998, and bb2000), and first-order condition approach (m1987 and gnr2020). All these approaches derive moment conditions from the underlying structural model and use these moments to estimate the model parameters by the generalized method of moments (GMM).

Moment Conditions for Three Modern Identification Strategies

Because there are numerous modifications and extensions of modern estimation strategies, I focus on the three most popular baseline strategies. In particular, I briefly summarize the control function approach of acf2015, the dynamic panel estimator of bb2000, and the first-order condition approach of gnr2020.

To keep the estimation strategies comprehensible and the notation simple, I restrict myself to the case of three inputs ($k_{it}$, $l_{it}$, and $m_{it}$), and make the following assumptions. i) Firms are time-homogeneous in their cd1928-type production technology, i.e. $P(\mathbf{x}_{it}) = (1, k_{it}, l_{it}, m_{it})$ and $\boldsymbol{\beta}_{i} = \boldsymbol{\beta}$. ii) Each firm chooses its period $t$ inputs based on available information $\mathcal{I}_{it}$, where $\omega_{it} \in \mathcal{I}_{it}$ is known before and $\epsilon_{it} \notin \mathcal{I}_{it}$ is realized after each firm's input decisions. iii) Firms predict their future productivity by $\omega_{it} = h_{\omega}^{0}(\omega_{it - 1}) + \eta_{it}$, where $\eta_{it}$ is unknown before $t$. iv) Input choices for capital and labor have dynamic implications, e.g. capital is accumulated by past investment decisions and labor is partially predetermined by hiring and firing costs.

In the following, I denote the $P$-dimensional vector with model parameters as $\boldsymbol{\theta}^{0}$. Furthermore, I refer to $\mathbf{g}(\xi_{i}^{t}, \boldsymbol{\theta}^{0})$ as $P^{\prime}$-dimensional vector of moments, where $\xi_{i}^{t} \coloneqq \{\xi_{it^\prime} \colon t^{\prime} \in \{0, \ldots, t\}\}$, so that $P^{\prime} \geq P$ moment conditions of the form $\operatorname{\mathbb{E}}[\mathbf{g}(\xi_{i}^{t}, \boldsymbol{\theta}^{0})] = 0$ can be used to estimate $\boldsymbol{\theta}^{0}$ by GMM.

exampleacf2015 propose a control function approach for a specific value-added production function \begin{equation*} y_{it} = \beta_{0}^{0} + \beta_{1}^{0} k_{it} + \beta_{2}^{0} l_{it} + \omega_{it} + \epsilon_{it} = \beta_{3}^{0} m_{it} + \omega_{it} + \epsilon_{it} \, . \end{equation*} This is a reasonable specification if there is perfect complementarity between $m_{it}$ and the other two inputs $k_{it}$ and $l_{it}$, i.e. \begin{equation*} y_{it} = \min(\beta_{0}^{0} + \beta_{1}^{0} k_{it} + \beta_{2}^{0} l_{it}, \beta_{3}^{0} m_{it}) + \omega_{it} + \epsilon_{it} \, . \end{equation*} Their identification strategy requires two additional assumptions. First, firms choose $m_{it} = h_{m}^{0}(k_{it}, l_{it}, \omega_{it})$. Second, $h_{m}^{0}(k_{it}, l_{it}, \omega_{it})$ must be a strictly monotonically increasing function in $\omega_{it}$, so that the only unobservable in this function can be expressed as $\omega_{it} = h_{m}^{0\langle-1\rangle}(k_{it}, l_{it}, m_{it})$, where $h_{m}^{0\langle-1\rangle}(\cdot)$ denotes the inverse function of $h_{m}^{0}(\cdot)$. This yields the following system of two equations: \begin{align} y_{it} =& \; \beta_{0}^{0} + \beta_{1}^{0} k_{it} + \beta_{2}^{0} l_{it} + h_{m}^{0\langle-1\rangle}(k_{it}, l_{it}, m_{it}) + \epsilon_{it} = h^{0}(k_{it}, l_{it}, m_{it}) + \epsilon_{it} \, , \nonumber \\ y_{it} =& \; \beta_{0}^{0} + \beta_{1}^{0} k_{it} + \beta_{2}^{0} l_{it} + h_{\omega}^{0}(h^{0}(k_{it - 1}, l_{it - 1}, m_{it - 1}) - \beta_{0}^{0} - \beta_{1}^{0} k_{it - 1} - \beta_{2}^{0} l_{it - 1}) + v_{it} \nonumber \, , \end{align} where $h^{0}(\cdot)$ and $h_{\omega}^{0}(\cdot)$ are unknown functions, usually approximated by polynomials, and $v_{it} \coloneqq \eta_{it} + \epsilon_{it}$. The model parameters can then be identified from the following moment conditions: $\operatorname{\mathbb{E}}[\epsilon_{it} \mid \mathcal{I}_{it}] = 0$ and $\operatorname{\mathbb{E}}[v_{it} \mid \mathcal{I}_{it - 1}] = 0$. For instance, if $h^{0}(k_{it}, l_{it}, m_{it}) = \alpha_{0}^{0} + \alpha_{1}^{0} k_{it} + \alpha_{2}^{0} l_{it} + \alpha_{3}^{0} m_{it}$ and $h_{\omega}(\omega_{it - 1}) = \delta^{0} \omega_{it - 1}$ then $\boldsymbol{\theta}^{0} = (\alpha_{0}^{0}, \alpha_{1}^{0}, \alpha_{2}^{0}, \alpha_{3}^{0}, \beta_{0}^{0}, \beta_{1}^{0}, \beta_{2}^{0}, \delta^{0})$ and $\mathbf{g}(\xi_{i}^{t}, \boldsymbol{\theta}^{0}) = (\epsilon_{it}, \epsilon_{it} k_{it}, \epsilon_{it} l_{it}, \epsilon_{it} m_{it}, v_{it}, v_{it} k_{it}, v_{it} k_{it - 1}, v_{it} l_{it - 1}, v_{it} m_{it - 1})$.
exampleThe dynamic panel estimator of bb2000 can be used to estimate the model parameters by imposing $h_{\omega}(\omega_{it - 1}) = \delta^{0} \omega_{it - 1}$. To see this, consider the production function \begin{equation*} y_{it} = \beta_{0}^{0} + \beta_{1}^{0} k_{it} + \beta_{2}^{0} l_{it} + \beta_{3}^{0} m_{it} + \omega_{it} + \epsilon_{it} \end{equation*} and its reformulation \begin{equation*} \Delta_{\delta}^{0}(y_{it}) = (1 - \delta^{0}) \beta_{0}^{0} + \beta_{1}^{0} \Delta_{\delta}^{0}(k_{it}) + \beta_{2}^{0} \Delta_{\delta}^{0}(l_{it}) + \beta_{3}^{0} \Delta_{\delta}^{0}(m_{it}) + w_{it} \, , \end{equation*} where $\Delta_{\delta}^{0}(x_{it}) \coloneqq x_{it} - \delta^{0} x_{it - 1}$ and $w_{it} \coloneqq \eta_{it} + \Delta_{\delta}^{0} \epsilon_{it}$. Under the additional assumption that $m_{it}$ has dynamic implications, e.g. due to adjustment costs or due to input market frictions, the model parameters can be identified from the moment conditions $\operatorname{\mathbb{E}}[w_{it} \mid \mathcal{I}_{it - 1}]$. For instance, the identification strategy of s2020 implies $\boldsymbol{\theta}^{0} = (\beta_{0}^{0}, \beta_{1}^{0}, \beta_{2}^{0}, \beta_{3}^{0}, \delta^{0})$ and $\mathbf{g}(\xi_{i}^{t}, \boldsymbol{\theta}^{0}) = (w_{it}, w_{it} k_{it}, w_{it} l_{it}, w_{it} k_{it - 1}, w_{it} l_{it - 1}, w_{it} m_{it - 1})$.
examplegnr2020 propose an estimation strategy for gross output productions functions, e.g. \begin{equation*} y_{it} = \beta_{0}^{0} + \beta_{1}^{0} k_{it} + \beta_{2}^{0} l_{it} + \beta_{3}^{0} m_{it} + \omega_{it} + \epsilon_{it} \, . \end{equation*} In their baseline identification strategy, they assume that firms maximize their expected future profits in a perfectly competitive environment. The idea is to identify $\beta_{3}^{0}$ from firms' optimal intermediate input decisions and then to exploit the dynamic structure of the model to identify the remaining model parameters. Under the additional assumption that $M_{it}$ is a flexible input with a linear cost function $P_{t}^{M} M_{it}$, firms choose $M_{it}$ as the solution to \begin{equation*} \max_{M_{it} \in \mathbb{R}^{>0}} \operatorname{\mathbb{E}}[P_{t}^{Y} \exp(\beta_{0}^{0} + \beta_{1}^{0} k_{it} + \beta_{2}^{0} l_{it} + \beta_{3}^{0} m_{it} + \omega_{it} + \epsilon_{it}) - P_{t}^{M} M_{it} \mid \mathcal{I}_{it}] \, , \end{equation*} where $P_{t}^{Y}$ and $P_{t}^{M}$ are common output and intermediate input prices, respectively. Reformulating the first-order condition \begin{equation*} P_{it}^{Y} \exp(\beta_{0}^{0} + \beta_{1}^{0} k_{it} + \beta_{2}^{0} l_{it} + \beta_{3}^{0} m_{it} + \omega_{it}) \beta_{3}^{0} \mathcal{E}^{0} = P_{it}^{M} M_{it} \end{equation*} and exploiting the dynamic structure yields the following system of two equations: \begin{align*} s_{it} =& \; \log(\beta_{3}^{0} \mathcal{E}^{0}) - \epsilon_{it} \, , \\ y_{it}^{r} =& \; \beta_{0}^{0} + \beta_{1}^{0} k_{it} + \beta_{2}^{0} l_{it} + h_{\omega}^{0}(y_{it - 1}^{r} - \beta_{1}^{0} k_{it - 1} - \beta_{2}^{0} l_{it - 1}) + \eta_{it} \, , \end{align*} where $s_{it} \coloneqq \log((P_{it}^{M} M_{it}) / (P_{it}^{Y} Y_{it}))$ is the logarithm of intermediate inputs expenditures relative to revenues, $\mathcal{E}^{0} \coloneqq \operatorname{\mathbb{E}}[\exp(\epsilon_{it}) \mid \mathcal{I}_{it}] = \operatorname{\mathbb{E}}[\exp(\epsilon_{it})]$ is a positive constant, and $y_{it}^{r} \coloneqq y_{it} - \beta_{3}^{0} m_{it} - \epsilon_{it}$. One difference to the other two strategies is that the researcher additionally needs data on the ratio of prices $P_{t}^{M} / P_{t}^{Y}$. The model parameters can then be identified from the following moment conditions: $\operatorname{\mathbb{E}}[\epsilon_{it} \mid \mathcal{I}_{it}] = 0$, $\operatorname{\mathbb{E}}[\exp(\epsilon_{it})] = \mathcal{E}^{0}$, and $\operatorname{\mathbb{E}}[\eta_{it} \mid \mathcal{I}_{it - 1}] = 0$. For instance, if $h_{\omega}(\omega_{it - 1}) = \delta^{0} \omega_{it - 1}$ then $\boldsymbol{\theta}^{0} = (\beta_{3}^{0}, \mathcal{E}^{0}, \beta_{0}^{0}, \beta_{1}^{0}, \beta_{2}^{0}, \delta^{0})$ and $\mathbf{g}(\xi_{i}^{t}, \boldsymbol{\theta}^{0}) = (\epsilon_{it}, \exp(\epsilon_{it}) - \mathcal{E}^{0}, \eta_{it}, \eta_{it} k_{it}, \eta_{it} l_{it}, \eta_{it} y_{it - 1}^{r})$.

Identifying Latent Group Structures

Penalized Generalized Method of Moments Estimator

I suggest to use the penalized GMM (PGMM) estimator of ssp2016 to estimate production functions with $J$ latent groups. The authors propose a novel Lasso penalization technique that achieves classification by shrinking firm-specific towards group-specific model parameters. Given $J$ and a strictly positive tuning parameter $\lambda$, the estimation procedure consists of three subsequent steps. First, the firm- and group-specific model parameters are estimated. Second, based on these estimates, each firm is assigned to a latent group. Third, group-specific model parameters are re-estimated based on the assigned latent group membership. In the following, I explain the three subsequent steps in more detail.

In the first step, the firm- and group-specific model parameters are estimated by minimizing

equation[equation omitted — 440 chars of source]

where $\mathbf{W}_{i}$ is a positive definite $P^{\prime} \times P^{\prime}$ weighting matrix, e.g. $\mathbf{W}_{i} = \operatorname{\mathbf{1}}_{P^{\prime}}$, and

equation[equation omitted — 186 chars of source]

Examples for $\boldsymbol{\pi}_{i}$ and $\mathbf{g}(\xi_{i}^{t}, \boldsymbol{\pi}_{i})$ are provided in Section (ref), e.g. $\boldsymbol{\pi}_{i} = (\beta_{3i}, \mathcal{E}_{i}, \beta_{0i}, \beta_{1i}, \beta_{2i}, \delta_{i})$ and $\mathbf{g}(\xi_{i}^{t}, \boldsymbol{\pi}_{i}) = (\epsilon_{it}, \exp(\epsilon_{it}) - \mathcal{E}_{i}, \eta_{it}, \eta_{it} k_{it}, \eta_{it} l_{it}, \eta_{it} y_{it - 1}^{r})$ for Example 3. The first term in (ref) is a sum of firm-specific GMM objective functions, and the second term is a penalty term controlled by $\lambda$. The former is minimal when all firm-specific moments are close to zero, whereas the latter is minimal when each $\boldsymbol{\pi}_{i}$ is close to any $\boldsymbol{\theta}_{j}$. Thus, while the GMM part ensures that the model fits the data, the penalty term forces each firm-specific model parameter to be close to any group-specific parameter.

In the second step, each firm is assigned to a latent group using a classification rule proposed by ssp2016:

equation[equation omitted — 357 chars of source]

where $(\hat{\boldsymbol{\pi}}_{1}, \ldots, \hat{\boldsymbol{\pi}}_{N})$ and $(\hat{\boldsymbol{\theta}}_{1}, \ldots, \hat{\boldsymbol{\theta}}_{J})$ are the PGMM estimates from the first step. Each firm is assigned to the latent group to which it is most similar, where similarity here is defined as the euclidean distance between the firm- and group-specific PGMM estimates.

In the third step, the group-specific model parameters are re-estimated by minimizing

equation[equation omitted — 463 chars of source]

where $\mathbf{W}_{j}$ is a positive definite $P^{\prime} \times P^{\prime}$ weighting matrix, e.g. $\mathbf{W}_{j} = \operatorname{\mathbf{1}}_{P^{\prime}}$, and $\widehat{G}_{1}, \ldots, \widehat{G}_{J}$ are the estimated latent groups from the second step. Thus, the third step is simply a standard GMM estimation done separately for each of the estimated latent groups. The corresponding Post-Lasso estimates are denoted as $(\tilde{\boldsymbol{\theta}}_{1}, \ldots, \tilde{\boldsymbol{\theta}}_{J})$.

remark(Group assignment). The classification rule in the second step ensures that each firm is assigned to a latent group. It is also possible to use a stricter rule that may leave some of the firms unclassified \begin{equation} \widehat{G}_{j} \coloneqq \{i \in \{1, \ldots, N\} \colon \min( \{\lVert \hat{\boldsymbol{\pi}}_{i} - \hat{\boldsymbol{\theta}}_{j^{\prime}} \rVert \colon j^\prime \in \{1, \ldots, J\}) = \lVert \hat{\boldsymbol{\pi}}_{i} - \hat{\boldsymbol{\theta}}_{j} \rVert \leq \varepsilon\} \; for \; j \in \{1, \ldots, J\} \, , \end{equation} where $\varepsilon$ is small positive constant. A firm is only assigned to its most similar group if the euclidean distance is sufficiently small. Stricter rules can help to deal with outlier firms and thus improve the performance of the estimation procedure. For clarification, unclassified firms are then excluded in the third step.

ssp2016 develop a limiting theory for PGMM estimators that can be applied to dynamic linear panel models such as those that are typically estimated by ab1991 estimators. Under asymptotics where $N, T \rightarrow \infty$, but not necessary at the same rate, and under the assumptions that $J = J^{0}$ and $\lambda \in \{T^{- a} \colon a \in (0, 0.5)\}$, the authors show that their PGMM estimator is able to assign all firms belonging to a latent group to the same group with probability approaching one.\footnote{Even if all firms are correctly classified, it does not follow that $\widehat{G}_{j} = G_{j}^{0}$. The classification rule only ensures that $\widehat{G}_{j} \in \{G_{1}^{0}, \ldots, G_{J^{0}}^{0}\}$, i.e. all firms in $G_{j}^{0}$ are assigned to the same latent group.} This classification consistency allows them to show that their Post-Lasso estimator is asymptotically equivalent to an infeasible estimator that knows and exploits the true latent group structure. Due to the close relationship between dynamic panel estimators and the estimation strategies for production functions, as pointed out by acf2015, I conjecture that the properties derived by ssp2016 also apply to my proposed nonlinear estimation procedure.\footnote{Although the production function itself is linear in parameters, the assumption about the evolution of productivity causes the entire model to become nonlinear. A formal proof for my estimation procedure will be added later. My conjecture is further supported by simulation experiments in Section (ref).} Thus, given that $(\tilde{\boldsymbol{\theta}}_{1}, \dots, \tilde{\boldsymbol{\theta}}_{J^{0}})$ is a suitable permutation of Post-Lasso estimators, the asymptotic distribution of $\tilde{\boldsymbol{\theta}}_{j}$ can be approximated by $\operatorname{\mathcal{N}}(\boldsymbol{\theta}_{j}^{0}, \mathbf{V}_{j})$ for all $j \in \{1, \ldots, J^{0}\}$, where $\mathbf{V}_{j}$ is a $P \times P$ covariance matrix.

remark(GMM inference). The weighting matrices $\mathbf{W}_{i}$ and $\mathbf{W}_{j}$ affect the performance of my estimation procedure only when the number of moment conditions is larger than the number of model parameters, i.e. $P^{\prime} > P$. Optimal weighting matrices that yield the most efficient GMM estimators are derived in \textcites{h1982}{hs1982}. Since these optimal weighting matrices depend on the model parameters, they have to be estimated. Thus, hs1982 suggest a two-step procedure where the optimal weighting matrix is constructed from first-step estimates. For $P^{\prime} > P$, $\mathbf{W}_{i}$ and $\mathbf{W}_{j}$ can also be constructed using two-step procedures, where in the first step $\mathbf{W}_{i} = \mathbf{W}_{j} = \operatorname{\mathbf{1}}_{P^{\prime}}$. Appropriate estimators for $\mathbf{V}_{1}, \ldots, \mathbf{V}_{J^{0}}$ depend on the identification strategy chosen and the underlying structural model assumptions. Examples of common estimators are heteroskedasticity consistent estimators ala w1980 or heteroskedasticity and autocorrelation consistent estimators ala nw1987. Alternatively, inference can be based on panel bootstrap procedures like k2008.
remark(Unbalanced panels). The estimation procedure can also be applied to unbalanced panels. Assuming that the observations of each firm are consecutive, i.e. $\{\xi_{it} \colon i \in \{1, \ldots, N\}, t \in \{t_{i}, \ldots, T_{i}\}, 0 \leq t_{i} < T_{i} \leq T\}$, adapting the estimator requires only replacing (ref) with \begin{equation} \overline{\mathbf{g}}(\xi_{i}^{T_{i}}, \boldsymbol{\theta}) = \frac{1}{T_{i} - t_{i}} \sum_{t = t_{i} + 1}^{T_{i}} \mathbf{g}(\xi_{i}^{t}, \boldsymbol{\theta}) \, . \end{equation} The asymptotic theory developed by ssp2016 applies only to balanced panels, but has recently been extended to unbalanced panels by swj2019. The crucial difference between the theories is that $\min(\{T_{i} - t_{i} \colon i \in {1, \ldots, N}\})$ has to be sufficiently large to ensure classification consistency in unbalanced panels.

Determining the Number of Latent Groups

The asymptotic theory assumes that the true number of latent groups is known. Since this is very unlikely in practice, I follow ssp2016 and suggest to estimate $J^{0}$ by minimizing a BIC-type information criterion

equation[equation omitted — 227 chars of source]

where $J$ is a guess for the true number of latent groups, $\tilde{r}_{it}(J, \lambda)$ are Post-Lasso residuals of an estimation given $\lambda$ and $J$, e.g. $\tilde{r}_{it}(J, \lambda) = \tilde{\eta}_{it}(J, \lambda) + \tilde{\epsilon}_{it}(J, \lambda)$ for Example 3, and $p(N, T)$ is a penalty term that satisfies $p(N, T) \rightarrow 0$ and $NT p(N, T) \rightarrow \infty$ as $N, T \rightarrow \infty$. ssp2016 suggest two penalty terms: $p(N, T) = 2 / 3 \, (NT)^{- 0.5}$ and $p(N, T) = 0.25 \log(\log(T)) / T$. Further examples of suitable penalty terms can be found in \textcites{bn2002}{ln2012}{bm2015}. Thus, given $\lambda$ and $p(N, T)$, the number of latent groups can be estimated as

equation[equation omitted — 177 chars of source]

where $\overline{J}$ is a known upper bound on the true number of latent groups.

remark(Joint determination). As pointed out by ssp2016, the information criterion can also be used to jointly determine $\lambda$ and $J$, \begin{equation} \widehat{J}_{p} \in \underset{\lambda \in \mathcal{L}}{\operatorname{\arg\,\min\;}} IC_{p}(\widehat{J}_{p}(\lambda), \lambda) \, , \end{equation} where $\mathcal{L} \coloneqq \{T^{- a} \colon a \in (0, 0.5)\}$. In practice, a grid of candidate values between 0 and 0.5 can be used to keep the number of estimates tractable, e.g. $a \in \{0.05, 0.1, 0.15, 0.2, 0.25, 0.3, 0.35, 0.4, 0.45\}$.

Estimation Algorithm

The PGMM objective function is not jointly convex in $(\boldsymbol{\pi}_{1}, \ldots, \boldsymbol{\pi}_{N})$ and $(\boldsymbol{\theta}_{1}, \ldots, \boldsymbol{\theta}_{J})$ and therefore very costly to minimize. To reduce the computational costs, ssp2016 suggest to split the optimization problem into a sequence of $J$ convex subproblems that are solved sequentially until convergence. The objective function of the $j$-th subproblem is

equation[equation omitted — 539 chars of source]

where $\zeta_{i}^{\langle j \rangle} \coloneqq \prod_{j^{\prime} \neq j}^{J} \lVert \boldsymbol{\pi}_{i}^{\langle j^{\prime} \rangle} - \boldsymbol{\theta}_{j^{\prime}} \rVert$ is the fixed part of the additive-multiplicative penalty term. The algorithm can be sketched as follows.

algorithm[algorithm omitted — 2,112 chars of source]

Although convexity significantly reduces the computational cost, the optimization problem is still challenging because the solution of each subproblem in Step 1 involves $(N + 1) \times P$ parameters. For instance, in my empirical illustration, I use a panel of $N = 571$ firms and a estimation strategy with $P = 6$ model parameters, implying a total of $3{,}432$ parameters per subproblem. To further reduce the computational costs, I suggest an Alternate Convex Search (see htw2015, chapter 5.9) that exploits the separable structure of (ref). Instead of minimizing jointly over $(\boldsymbol{\pi}_{1}^{\langle j \rangle}, \ldots, \boldsymbol{\pi}_{N}^{\langle j \rangle})$ and $\boldsymbol{\theta}_{j}$, I can alternate between solving two optimization problems: i) minimization over $(\boldsymbol{\pi}_{1}^{\langle j \rangle}, \ldots, \boldsymbol{\pi}_{N}^{\langle j \rangle})$ holding $\boldsymbol{\theta}_{j}$ fixed and ii) minimization over $\boldsymbol{\theta}_{j}$ holding $(\boldsymbol{\pi}_{1}^{\langle j \rangle}, \ldots, \boldsymbol{\pi}_{N}^{\langle j \rangle})$ fixed. Minimization i) is still over $N \times P$ parameters. However, holding $\boldsymbol{\theta}_{j}$ fixed, the optimization problem can be separated into $N$ independent subproblems involving only $P$ parameters each. Minimization ii) involves only $P$ parameters and is just a Fermat-Weber location problem for which there are numerous efficient algorithms, e.g. the fixed-point algorithm of w1937. The algorithm can be sketched as follows.

algorithm[algorithm omitted — 1,837 chars of source]
remark(Computation). First, although each minimization problem in Step 2 is convex, the corresponding objective function is not differentiable due to the Euclidean norm in the penalty function. For these types of minimization problems, there are special optimization algorithms, such as those presented in htw2015. Second, because the $N$ subproblems are independent of each other, they can also be easily parallelized.

Simulation Experiments

To analyze the classification accuracy and statistical properties of my proposed estimation procedure, I extend the representative firm model of s2001 to incorporate latent group structures. The model has the advantage that the input decision problem of each firm can be solved analytically. Similar firm models have been used, for instance, by \textcites{vb2007}{acf2015}{cwdl2016}.

I simulate samples of $N = 200$ firms. Each firm is observed for $T$ time periods and belongs to one of $J^{0} = 3$ latent groups. Each group consists of $N_{j}$ firms. All firms are time-homogeneous within a group, but differ in their parameter configuration and in their relative occurrence between groups. At the beginning of period $t$, a firm $i$ with rational expectations has the following input decision problem:

eqnarray[eqnarray omitted — 486 chars of source]

where $Y_{it}$, $K_{it}$, $M_{it}$, $I_{it}$ are output, capital, intermediate input, and investment, $\mathcal{I}_{it}$ is a set of information available at the beginning of period $t$, $\omega_{it}$ is an anticipated productivity shock, and $\epsilon_{it}$ is an unanticipated productivity shock realized after each firm's input decision. Furthermore, $\log(\phi_{i}^{- 1}) \sim \operatorname{\text{iid.}\;} \operatorname{\mathcal{N}}(0, 1)$, $\eta_{it} \sim \operatorname{\text{iid.}\;} \operatorname{\mathcal{N}}(0, \sigma_{\eta_{i}}^{2})$, $\epsilon_{it} \sim \operatorname{\text{iid.}\;} \operatorname{\mathcal{N}}(0, \sigma_{\epsilon_{i}}^{2})$, $\omega_{i0} \sim \operatorname{\text{iid.}\;} \operatorname{\mathcal{N}}(\alpha_{i} / (1 - \delta_{i}), \sigma_{\eta_{i}}^{2} / (1 - \delta_{i}^2))$, and $K_{i0} \coloneqq \mathbf{0}_{N}$. The model parameters are described and defined in Table (ref).

table[table omitted — 1,162 chars of source]

I briefly summarize the core features of the model. Firms maximize their expected future profits with respect to their intermediate input and investment choices. Future profits are discounted and all firms face firm-specific but time constant quadratic capital adjustment costs, as in acf2015. Within a latent group, all firms share the same time-homogeneous Cobb-Douglas production function with constant returns to scale. The current capital stock is accumulated through a dynamic process determined by depreciation and past investment decisions. Productivity is additively separable and can be decomposed into a persistent component and an ex-post shock. The persistent component can be predicted by an AR(1) process.

Because $M_{it}$ is a fully flexible input, i.e. there are no adjustment costs or other dynamic implications, firm $i$'s optimal choice at time $t$ follows immediately from the first-order condition, i.e.

equation*[equation* omitted — 125 chars of source]

where $\mathcal{E}_{i} = \exp(\sigma_{\epsilon_{i}}^{2} / 2)$. Using the Euler equation for investment along with forward substitution yields a fully deterministic optimal investment decision

equation*[equation* omitted — 386 chars of source]

To ensure that all firms are sampled from their steady state distribution, I extend each time span by $1{,}000$ initial periods that are excluded from the final sample. For clarification, the final sample has $N(T + 1)$ observations. The additional time period per firm is used to generate lagged values of the output and input variables so that the final sample used for the estimation has $NT$ observations. Because the optimal investment decision is a convergent series, I approximate it by the first $1{,}001$ terms of the sum.\footnote{More specifically, I split the series into two parts:

equation*[equation* omitted — 126 chars of source]

where $c_{i}$ is the factor in front of the series and $a_{i\tau}$ are the terms of the series. I assume that the first $1{,}001$ terms are sufficient to approximate the optimal level of investment, i.e. all remaining terms are negligible small.}

The length of the panel is the key determinant for the performance of my estimation procedure. Thus, I focus on samples with different time spans $T \in \{15, 25, 50\}$. The estimation procedure is based on the identification strategy of gnr2020. The corresponding moment conditions are adjusted to the functional forms of the data generating process, i.e. Cobb-Douglas production function, AR(1) process for the persistent component of productivity, and $s_{it} = m_{it} - y_{it}$. I do not exploit the constant returns to scale restriction that would allow me to recover both elasticities from the share equation only. The analysis consists of two parts. In the first part, I assume that $J^{0}$ is unknown and study the accuracy of different estimators for it. In the second part, I assume that $J = J^{0} = 3$ is known and analyze the finite sample behavior of my estimation procedure. All reported results are based on 100 simulated samples and $\lambda = T^{- 0.25}$.\footnote{In a preliminary analysis, I tried different values for $\lambda \in \{T^{- a} \colon a \in \{0.05, 0.1, 0.15, 0.2, 0.25, 0.3, 0.35, 0.4, 0.45\}\}$ and found that the estimation procedure is very robust to different choices of $\lambda$.}

In the first part, I analyze the performance of $\widehat{J}_{p}(\lambda)$ defined in (ref) with $\overline{J} = 5$. I consider different penalty terms: $p_{1, r}(N, T) \coloneqq r \, (NT)^{- 0.5}$ and $p_{2, r}(N, T) \coloneqq r \log(\log(T)) / T$, where $r \in \{0.25, 0.5, 0.75, 1\}$ is a finite sample adjustment factor.\footnote{I also tried some of the penalty terms suggested by bn2002, but I did not find any better alternatives.} The analysis is based on the following quantities: expected value and probabilities to select exactly or at least $J^{0} = 3$ latent groups. Table (ref) summarizes the results.

table[table omitted — 1,773 chars of source]

For $T \in \{25, 50\}$, all estimators perfectly predict $J^{0}$. Only for $T = 15$, I find noticeable differences in the performance. While the estimator with $p_{1, r}(N, T)$ slightly overestimates the number of latent groups for small $r$, the estimator with $p_{2, r}(N, T)$ underestimates it for large $r$. Given a suitable adjustment factor, estimators based on both specifications are able to perfectly predict $J^{0}$. Overall, my simulation experiments suggest a larger adjustment factor for $p_{1, r}(N, T)$ and a smaller factor for $p_{2, r}(N, T)$, e.g. $r = 1$ for the former and $r = 0.25$ for the latter.

In the second part, I take $J = J^{0}$ as given and compare the finite sample performance of my Post-Lasso estimator (Post-Lasso) with an infeasible estimator that knows and exploits the true latent group structure (Infeasible). For the comparison, I consider the following quantities: bias, standard deviation, root mean squared error, ratio of standard error and standard deviation, and coverage rates of confidence intervals with 95% nominal level. The statistics are computed separately for each firm. The bias, standard deviation, and root mean squared error are all relative to the true parameter value in percent. To keep the analysis concise, I focus on the output elasticities, i.e. $\boldsymbol{\gamma} \coloneqq (\gamma_{1}, \ldots, \gamma_{N})$ and $\boldsymbol{\beta} \coloneqq (\beta_{1}, \ldots, \beta_{N})$, and report aggregate rather than firm-specific quantities, e.g. relative bias as $100 / N \sum_{i = 1}^{N} (\bar{\hat{\gamma}}_{i} - \gamma_{i}) / \gamma_{i}$, where $\bar{\hat{\gamma}}_{j}$ is the average estimate of $\gamma_{i}$ over all simulated samples. Since the performance of Post-Lasso depends on the classification accuracy in the first step of the estimation procedure, I additionally report the fraction of correctly classified firms. Table (ref) summarizes the results.

table[table omitted — 2,120 chars of source]

I start with the first step classification accuracy. The fractions of correctly classified firms are always close to one, even for $T = 15$. The classification accuracy improves as $T$ increases. For $T = 50$, the fraction of correctly classified firms is one. The excellent classification accuracy is also reflected in the finite sample performance. Even for $T = 15$, the performance of Post-Lasso is nearly identical to that of the benchmark estimator. In general, I find biases smaller than 1%, ratios of standard errors and standard deviation near one, and coverage rates close to their nominal values. Misclassification in the first step mainly results in larger dispersions. Consequently, the largest differences in performance are apparent in the standard deviation and the root mean squared error.

Empirical Illustration

It is standard practice to estimate separate production functions for each industry. In most studies, industries are delineated by an industry classification, like ISIC, SIC, NACE, or NAICS.\footnote{Some recent and influential examples are \textcites{adkpvr2020}{dleu2020}. The former use a 2-digit SIC and the latter use a 2- and 4-digit NAICS classification to define an industry.} The underlying assumption here is that firms operating in the same economic environment, like an industry, use the same production technology. Although ex-ante classification by industry is very convenient and intuitive, the question arises whether this classification is sufficient to fully account for latent firm heterogeneity.

I use a balanced subsample of the Chilean panel data set used by gnr2020.\footnote{Previous studies that also use the Chilean data set include \textcites{p2002}{lp2003}.} The data is provided by Instituto Nacional de Estadística de Chile and includes all manufacturing plants with more than ten employees in the five largest industrial sectors, i.e. food products (311), textiles (321), apparel (322), wood products (331), and fabricated metals (381), that were in operation between 1979 and 1996.\footnote{The data is taken from the \href{https://www.journals.uchicago.edu/doi/suppl/10.1086/707736}{replication package} provided by the authors. \href{https://unstats.un.org/unsd/classifications/Econ/ISIC}{ISIC Rev. 2} industry classification.} In addition to the output and input variables $(Y_{it}, K_{it}, L_{it}, M_{it}, S_{it})$, the sample contains further firm characteristics. I know whether a firm is an exporter, is an importer of intermediate goods, pays above median industry wages, or is an advertiser. Further, $Y_{it}$ is measured as deflated revenue, $K_{it}$ is the capital stock constructed by the perpetual inventory method, $L_{it}$ is a weighted sum of unskilled and skilled number of workers, $M_{it}$ is measured as the sum of several intermediate input expenditures, e.g. raw materials, energy, and services, and $S_{it}$ is intermediate input expenditure relative to revenue.\footnote{Further details about the construction of the sample are provided in gnr2020.}

I consider the following time-homogeneous Cobb-Douglas production function

equation[equation omitted — 165 chars of source]

where $i \in \{1, \ldots, 571\}$, $t \in \{0, \ldots, 17\}$, $\omega_{it} = \delta_{0i} + \delta_{1i} \omega_{it - 1} + \eta_{it}$ is a persistent productivity shock following an AR(1) process, $\eta_{it}$ is an unexpected innovation, and $\epsilon_{it}$ is an ex-post productivity shock. I assume that the model parameters follow an unknown general group pattern. To deal with this latent firm heterogeneity, I consider two classification strategies: i) ex-ante classification by industry and ii) data-driven classification using PGMM estimation. Estimates of both strategies are based on the moment conditions of gnr2020, i.e. the moment conditions defined in Example 3.

For the PGMM estimation, I need to specify two additional parameters: $\lambda$ and $J$. To determine both parameters jointly, I use information criteria with two different penalty terms: $p(N, T) = (NT)^{- 0.5} \approx 0.0101$ and $p(N, T) = 0.25 \log(\log(T)) / T \approx 0.0153$. I compute both criteria for all combinations of $\lambda \in \{T^{- a} \colon a \in \{0.2, 0.25, 0.3, 0.35, 0.4\}\}$ and $J \in \{1, \ldots, 8\}$, where $J = 1$ refers to the absence of latent firm heterogeneity, i.e. firm homogeneity. Figure (ref) visualizes the results.

figure[figure omitted — 593 chars of source]

Both information criteria suggest $\lambda = 17^{-0.4} \approx 0.3220$ and $J = 3$ for the PGMM estimation. Thus, there is strong evidence for latent firm heterogeneity, as indicated by the significantly improved fit of the structural model for $J > 1$.

So far, I found that the estimated number of latent groups is smaller than the number of industries. However, it is unclear to which extent the data-driven classification matches the ex-ante classification. The chord diagram in Figure (ref) visualizes the matching.

figure[figure omitted — 703 chars of source]

First, I find firms from all three latent groups in each of the five industries. Second, I find that firms from different industries split very unequally among the latent groups. For instance, firms in industries 311 and 381 split mainly between two latent groups, whereas firms in the other industries split more evenly between all three latent groups. Furthermore, about three quarters of the firms in industry 311 are assigned to Group 2. Thus, the matching analysis provides evidence that an ex-ante classification by industry is not sufficient to fully account for latent firm heterogeneity, which is in line with kss2017 who report substantial firm heterogeneity even in narrowly defined industries.

Finally, I analyze the extent to which estimates based on ex-ante and data-driven classification differ. I start with the model fit and compare the mean squared residuals (MSR) for both classification approaches. The residuals are defined as $\hat{\eta}_{it} + \hat{\epsilon}_{it}$. The MSR for the industry classification ($\approx 0.0703$) is larger than the MSR for the data-driven classification ($\approx 0.0528$). Thus, the data-driven classification yields a better model fit and does so even with fewer parameters. Next, I analyze the differences in the estimated output elasticities and how these differences translate into heterogeneity in total factor productivity (TFP).\footnote{In addition to output elasticities, TFP is another quantity of interest in some empirical studies. These studies are particularly interested in getting a better understanding of persistent TFP differences that are frequently reported (see bd2000 and s2011 for comprehensive overviews).} I follow op1996 and estimate TFP in levels as $\exp(y_{it} - \hat{\beta}_{1i} k_{it} - \hat{\beta}_{2i} l_{it} - \hat{\beta}_{3i} m_{it})$. I analyze the heterogeneity in TFP through excluded firm characteristics using a pseudo-poisson estimator with the following conditional mean specification

equation[equation omitted — 184 chars of source]

where the parameters $\pi_{1i}$, $\pi_{2i}$, and $\pi_{3i}$ follow the group pattern implied by the ex-ante or data-driven classification, $\text{trade}_{it} \coloneqq \max(\text{exporter}_{it}, \text{importer}_{it})$, $\text{highwage}_{it}$, and $\text{advertiser}_{it}$ are indicator variables equal to one, if firm $i$ at time $t$ engages in international trade, either through exporting and/or importing, pays above median wages, and is an advertiser, respectively, and $\alpha_{i}$ and $\gamma_{t}$ are firm and year fixed effects. Table (ref) summarizes the results.

sidewaystable[!htbp] \begin{threeparttable} \caption{Estimation Results: Output Elasticities and Heterogeneity in TFP} \begin{tabular}{@*{1}{l}*{8}{c}@} \toprule &\multicolumn{5}{c}{Ex-Ante Classification}&\multicolumn{3}{c}{Data-Driven Classification}\\ \cmidrule(lr){2-6}\cmidrule(lr){7-9} &311&321&322&331&381&Group 1&Group 2&Group 3\\ \midrule \multicolumn{9}{l}{Output Elasticities:}\\[0.5em] $\quad$Capital & 0.163 & 0.096 & 0.149 & 0.095 & 0.186 & 0.158 & 0.137 & 0.199 \\ & (0.012) & (0.027) & (0.033) & (0.025) & (0.034) & (0.014) & (0.010) & (0.023) \\ $\quad$Labor & 0.182 & 0.346 & 0.249 & 0.284 & 0.472 & 0.299 & 0.155 & 0.503 \\ & (0.018) & (0.037) & (0.046) & (0.037) & (0.056) & (0.022) & (0.015) & (0.041) \\ $\quad$Intermediates & 0.674 & 0.494 & 0.550 & 0.583 & 0.439 & 0.557 & 0.720 & 0.386 \\ & (0.003) & (0.007) & (0.007) & (0.007) & (0.006) & (0.003) & (0.002) & (0.005) \\ \multicolumn{9}{l}{Heterogeneity in Total Factor Productivity:}\\[0.5em] $\quad$Trade & 0.027 & 0.047 & 0.043 & -0.036 & 0.028 & 0.042 & 0.023 & 0.070\\ & (0.016) & (0.026) & (0.032) & (0.039) & (0.025) & (0.024) & (0.011) & (0.036)\\ $\quad$Wages > median & 0.042 & 0.053 & 0.085 & 0.070 & 0.070 & 0.043 & 0.048 & 0.076\\ & (0.008) & (0.024) & (0.033) & (0.033) & (0.027) & (0.016) & (0.009) & (0.039)\\ $\quad$Advertiser & -0.008 & 0.008 & -0.060 & -0.003 & -0.027 & -0.031 & -0.003 & -0.022\\ & (0.009) & (0.023) & (0.035) & (0.019) & (0.022) & (0.015) & (0.008) & (0.028)\\ \bottomrule \end{tabular} \begin{tablenotes} • Note: Output Elasticities are estimated based on the moment conditions of gnr2020; w1980-type standard errors in parentheses; Heterogeneity in TFP is analyzed through excluded firm characteristics using a pseudo poisson estimator with conditional mean specification (ref); TFP is estimated in levels as $\exp(y_{it} - \hat{\beta}_{1i} k_{it} - \hat{\beta}_{2i} l_{it} - \hat{\beta}_{3i} m_{it})$; pseudo-poisson estimates can be interpreted as semi-elasticities and are relative to firms that do not export or import, pay below median wages and do not advertise; for instance, a firm in industry 311 engaged in international trade is $100 (\exp(0.027) - 1)\% \approx 2.7368\%$ more productive than a firm with the same characteristics that is not engaged. • Source: Data taken from the replication package of gnr2020. \end{tablenotes} \end{threeparttable}

I find sizable differences, relative to the magnitude of the standard errors, between the estimated output elasticities for the ex-ante and data-driven classification. There are also some similarities between the estimates for Industry 311 and Group 2 and the estimates for Industry 381 and Group 3. These similarities might be explained by the large overlap in the matching of the firms. The estimated capital elasticities for the data-driven classification are substantially larger than those for the ex-ante classification whereas the ranges of the estimated labor and intermediate input elasticities are wider. The differences in the estimated output elasticities also lead to different conclusions about the heterogeneity in TFP. For Industry 331, I find negative but insignificant TFP differences between firms that engage in international trade and those who do not. The differences in all other industries are positive, but only for industry 321 significantly different from zero.\footnote{This is surprising, as there is wide agreement that firms that engage in international trade, such as exporters, are more productive. Two reasons mentioned by \textcites{bw1997}{bj1999} are self selection into exporting, e.g. because entering foreign markets is costly, and learning by exporting, e.g. knowledge spillovers.} The remaining conclusions about heterogeneity in TFP remain qualitatively the same for both classification approaches. For instance, I find that firms with above median wages are significantly more productive, while firms that advertise are not significantly different from those that do not advertise.

Concluding Remarks

I propose a fully data-driven estimation procedure for production functions with latent group structures. My approach combines recent identification strategies with the classifier-Lasso of ssp2016. Simulation experiments confirm that my estimation procedure is well suited to deal with latent firm heterogeneity in sufficiently long panels. The practical relevance is illustrated with a panel of Chilean firms.

Future research could relax the time-homogeneity assumption of the model parameters within a latent group, e.g. by introducing smooth time-varying model parameters as in swj2019. Another possible extension could be to think of the model parameters not as group-specific fixed constants, but as firm-specific random coefficients with unknown distribution function. Group-specific estimates of the model parameters could then be interpreted as representative points of an unknown distribution.

\printbibliography