EconBase
← Back to paper

Identification and Estimation of Production Function with Unobserved Heterogeneity

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

101,961 characters · 9 sections · 65 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Identification and Estimation of Production Function with Unobserved Heterogeneity Hiroyuki Kasahara Vancouver School of Economics University of British Columbia [email removed] Paul Schrimpf Vancouver School of Economics University of British Columbia [email removed] Michio Suzuki Tohoku University Cabinet Office, Government of Japan [email removed]

{ \singlespacing }

center[center omitted — 13 chars of source]
abstractThis paper examines the nonparametric identifiability of production functions, considering firm heterogeneity beyond Hicks-neutral technology terms. We propose a finite mixture model to account for unobserved heterogeneity in production technology and productivity growth processes. Our analysis demonstrates that the production function for each latent type can be nonparametrically identified using four periods of panel data, relying on assumptions similar to those employed in existing literature on production function and panel data identification. By analyzing Japanese plant-level panel data, we uncover significant disparities in estimated input elasticities and productivity growth processes among latent types within narrowly defined industries. We further show that neglecting unobserved heterogeneity in input elasticities may lead to substantial and systematic bias in the estimation of productivity growth.

Introduction

Estimating a firm's production function and productivity is a critical topic in empirical economics.\footnote{Understanding how the input is related to the output is a fundamental issue in empirical industrial organization Ackerberg07 while a measure of total factor productivity is necessary to examine the effect of trade policy on productivity and to analyze the role of resource allocation on aggregate productivity Pavcnik00,KASAHARA2008,KASAHARA2013,Hsieh09. Production function estimation is also important for markup estimation hall1988relation, de2012markups,Raval2023.} Despite its importance, the standard production function estimation procedures impose an implausible assumption that production functions are common across firms except for separable Hicks-neutral productivity terms Olley1996, Levinsohn03, Wooldridge09, Ackerberg15, gandhi20. In the presence of identification issues due to the simultaneity problem Marschak44, the literature on identifying production functions that incorporate unobserved heterogeneity beyond the Hicks-neutral productivity term is scarce but a rapidly growing research area Li2017, doraszelski2018measuring, balat19, ZHANG2019, Demirer20, chen21, Raval2023.

This paper establishes the nonparametric identification of production functions from panel data when production functions and productivity growth processes are heterogeneous across firms in unobserved time-varying ways. We consider a finite mixture specification in which there are $J$ distinct time-varying production technologies, and each firm belongs to one of the $J$ latent types. Econometricians do not observe the latent type of firms. Without making any functional form assumption on production technology and productivity growth processes, we establish nonparametric identification of $J$ distinct production functions, productivity processes, and a population proportion of each type under assumptions similar to those used in the existing production function and panel data identification literature. We also address potential measurement errors in labor inputs because of unobserved working hours and labor quality.

Building on our nonparametric identification result and considering computational ease, we propose an estimation procedure for the production function with random coefficients. Under the assumption of Gaussian error terms, we develop a penalized maximum likelihood estimator for a finite mixture model of random coefficient production functions, where the form of the likelihood function is motivated by our identification argument. The EM algorithm is employed to simplify the computational complexity of maximizing the log-likelihood function of the mixture model.

As an empirical application, we investigate the extent of production technology heterogeneity across plants using panel data from Japanese manufacturing plants between 1986 and 2010. We show that the ratios of intermediate cost to sales and intermediate cost to variable cost—averaging from 1986 to 2010 at the plant level—are substantially different across plants within narrowly defined industries. For example, the 90th-10th percentile differences in the intermediate cost shares in variable costs for concrete products and electric audio equipment are large, at 0.27 and 0.67, respectively.\footnote{ZHANG2019 finds considerable heterogeneity in the labor share in sales across firms in the Chinese steel industry.} Differences in input levels cannot explain these heterogeneities; that is, a considerable cross-plant variation in the ratio of intermediate cost to sales or variable costs remains after controlling for observable inputs, presenting evidence for persistent and substantial heterogeneity in production technologies across plants.

By employing a finite mixture of random coefficients production functions, we also find substantial differences in estimated input elasticities across latent types within narrowly defined industries. To understand the consequences of neglecting unobserved heterogeneity in input elasticities on productivity growth measurement, we adopt a specification with unobserved heterogeneity as the true model and calculate the bias in productivity growth measurement when using a misspecified production function model that omits unobserved heterogeneity. Our findings indicate that ignoring unobserved heterogeneity in input elasticities can result in substantial and systematic bias in estimated productivity growth, contingent on the heterogeneous parameter estimates and the direction of productivity changes. Additionally, our analysis reveals a significantly stronger correlation between estimated productivity and investment among high capital-intensive latent type firms compared to low capital-intensive type firms, implying that unobserved disparities in input elasticities are vital in plant-level investment decisions.

As first discussed by Marschak44, ordinary least squares estimation of production functions is subject to simultaneity bias when firms make input decisions based on their productivity level Griliches98. To address the simultaneity issue, Olley1996 and Levinsohn03 develop control function approaches, which have been widely applied in empirical studies Wooldridge09, Ackerberg15. Despite their popularity, the control function approach has faced potential identification issues as highlighted in the literature. bond05 and Ackerberg15 discuss identification issues due to collinearity under two flexible inputs (i.e., material and labor) in Cobb-Douglas specification. Furthermore, gandhi20 contend that if the firm's decision follows a Markovian strategy, the moment restriction utilized in the control function approach fails to provide sufficient restriction to identify flexible input elasticities due to a lack of instrumental power.

GNR exploit the first-order condition for flexible input under profit maximization and establish the identification of production functions without making any functional form assumptions. However, their result presumes that production technology is identical across plants, except for the Hicks-neutral productivity term. This paper extends the nonparametric identification approach of GNR to accommodate settings where production technologies exhibit unobserved heterogeneity across plants.

Several papers employ the first-order condition as a restriction to identify heterogenous elasticities of flexible inputs under functional form assumptions VanBiesebroeck03, Li2017, doraszelski2018measuring, balat19, ZHANG2019.\footnote{As Solow57 first illustrates, the flexible input elasticities are identified with their input revenue share under the Cobb-Douglas production functions.} doraszelski2018measuring develop a framework to identify plant-level time-varying labor-augmenting productivity in addition to Hicks-neutral productivity under the constant elasticity of substitution (CES) production function, allowing for two-dimensional heterogeneity. ZHANG2019 proposes an estimation method based on the CES production function that accounts for heterogeneity in capital, labor, and material-augmenting efficiency across firms. Li2017 use the flexible input cost ratio to construct a control variable for latent technology to identify flexible inputs' elasticities while imposing the timing assumption suggested by Ackerberg15 to identify the labor and capital coefficients under the Cobb-Douglas specification. balat19 also consider the Cobb-Douglas production function with heterogeneity in the efficiency of using skilled and unskilled labor. Demirer20 extends the framework of doraszelski2018measuring by relaxing the parametric assumption of the CES production function but assumes that the labor-augmenting technology is the only additional source of individual-level heterogeneity other than the Hicks neutral productivity. Raval2023 demonstrates the importance of accommodating non-neutral productivity differences across firms when estimating markups using flexible inputs. dewitte2022 illustrate the significance of accounting for unobserved heterogeneity in productivity growth processes when analyzing export premia and the contributions of exporting firms to aggregate productivity.

These papers identify firm-specific input elasticities, factor-augmenting technologies, or productivity growth processes but impose parametric assumptions or limit the sources of heterogeneity. Our paper complements these studies by establishing nonparametric identification of heterogenous production functions and productivity growth processes without imposing any functional form assumptions or limiting sources of unobserved heterogeneity.

Cheng21 extend the k-means clustering approach of Bonhomme15ecma to multi-dimensional clustering in random coefficient production functions in a nonlinear GMM framework, building upon the dynamic panel approach Arellanobond91, Blundell1998, Blundell2000. Cheng21 consider an asymptotic setting when the time dimension $T\rightarrow \infty$ while our identification is based on $T$ being fixed. Our identification result with fixed $T$ is useful in empirical applications where firm-level panel datasets have limited time dimensions.

Our paper also contributes to the literature on identifying dynamic panel data models with unobserved heterogeneity by relaxing the existing identification conditions. Specifically, Proposition 3 demonstrates that the mixing probabilities and the type-specific time-varying probability distributions across latent types can be identified from panel data with four periods under the Markov assumption and other regularity conditions. This result improves upon the findings of Kasahara09, who established the identification of dynamic panel data models under the Markov assumption but imposed the stationarity and required panel data with six periods.\footnote{HigginsJochmans21 point out that the type-specific distribution is identified only up to an arbitrary ordering of the latent types that differs across different points in Kasahara09. Our argument for identifying a common order of the latent types is based on that of HigginsJochmans21. } Hu12 consider a non-stationary case and establish the identification of a continuous mixture dynamic panel data model using a panel dataset with five time periods. However, their result is limited, as they only establish the identification of type-specific distributions for the third to fifth periods, leaving the identification for the first two periods unresolved. In contrast, we identify the type-specific distributions across all four periods from panel data of length four.

A key condition for our identification analysis is that the observed variable must follow a first-order Markov process within the subpopulation specified by latent type. Proposition (ref) demonstrates that this Markov assumption is satisfied under our structural model assumptions, including the Markovian investment strategy (Assumption (ref)(b)) and the monotonicity of flexible input demands for productivity and wage shocks (Assumption (ref)(b)), both of which are standard assumptions in the production function literature Olley1996,Levinsohn03. Another identifying condition is a rank condition in Assumption (ref) which requires that the changes in the value of the observed vector, $\boldsymbol z_{t}$, must induce sufficiently different changes in the value of the type specific conditional density function of $\boldsymbol z_t$ given the past value $\boldsymbol z_{t-1}$ across latent types. As illustrated in the Cobb-Douglas example in Appendix (ref), this condition is satisfied when input elasticities are sufficiently different across latent types.

The remainder of this paper is organized as follows. Section 2 presents evidence of heterogeneity in production technologies across plants, using panel data from Japanese manufacturing plants. Section 3 introduces the setup for our production function models and discusses the assumptions. Section 4 provides the main identification results, while Section 5 develops an estimator for the production function using a finite mixture model. In Section 6, we present empirical results based on the Japanese manufacturing plants. Section 7 concludes the paper.

Evidence for production technology heterogeneity

In order to underscore the significance of accounting for unobserved heterogeneity in production functions, we first present a set of stylized facts that clearly indicate the presence of heterogeneity beyond Hicks-neutral technology components in production functions. Our analysis employs panel data derived from Japanese manufacturing facilities, spanning the years 1986 to 2010. A comprehensive discussion of the dataset can be found in Section (ref).

For illustration, consider a plant with the Cobb-Douglas production technology: \[ \log Y_{it} =\beta_0+ \beta^i_{m} \log M_{it}+ \beta^i_{\ell} \log L_{it} + \beta^i_{k} \log K_{it} +\omega_{it},\smallskip \] where $Y_{it}$, $M_{it}$, $L_{it}$, and $K_{it}$ denote the output, intermediate input, labor, and capital of plant $i$ in year $t$, respectively, while $\omega_{it}$ represents the total factor productivity (TFP) that follows a first-order Markov process. The superscript $i$ in $\beta^i_m$, $\beta_\ell^i$, and $\beta^i_k$ signifies the variation in output elasticities of inputs across different plants.

We assume that firms consider their output and input prices as given, and that the intermediate input and labor are flexibly selected after $\omega_{it}$ has been fully observed. Consequently, a plant's profit maximization implies the following relationships:

equation[equation omitted — 195 chars of source]

where $P_{Y,t}$, $P_{M,t}$, and $W_{t}$ represent the prices of output, intermediate input, and labor, respectively.

In most existing empirical studies, production functions are estimated under the assumption that the coefficients $\beta^i_m$, $\beta^i_\ell$, and $\beta^i_k$ do not vary across plants. This assumption can be tested in light of ((ref)) by examining whether the intermediate input share, $ \frac{P_{M,t}M_{it}}{P_{Y,t} Y_{it}}$, and the ratio of intermediate cost to variable cost (i.e., the sum of intermediate and labor costs), $\frac{P_{M,t}M_{it}}{P_{M,t}M_{it}+W_{t}L_{it}}$, remain constant across plants. To investigate this, we calculate the plant-level averages over the maximum 25-year period, during which a plant remained in the market between 1986 and 2010, as follows: \[ \overline{ \left(\frac{P_{M,t}M_{it}}{P_{Y,t} Y_{it}}\right)}_i=\frac{1}{25}\sum_{t=1986}^{2010} \frac{P_{M,t}M_{it}}{P_{Y,t} Y_{it}}\quad\text{and}\quad \overline{\left(\frac{P_{M,t}M_{it}}{P_{M,t}M_{it}+W_{t}L_{it}}\right)}_i=\frac{1}{25}\sum_{t=1986}^{2010}\frac{P_{M,t}M_{it}}{P_{M,t}M_{it}+W_{t}L_{it}}. \] Subsequently, we analyze the extent of variation across plants within a narrowly defined industry.

Figure (ref) and Figure (ref) display histograms illustrating plant-level averages of intermediate input shares, $\overline{ \left(\frac{P_{M,t}M_{it}}{P_{Y,t} Y_{it}}\right)}_i$, for all plants within the concrete products and electric audio equipment industries, respectively. Both figures exhibit substantial variation in intermediate shares. The disparity between the 90th and 10th percentiles reaches up to 0.28 for concrete products, an industry typically regarded as having homogeneous technology.

figure[figure omitted — 484 chars of source]
figure[figure omitted — 520 chars of source]

The variation in intermediate shares between plants might be indicative of differences in markups; however, the ratio of intermediate costs to total variable costs is less susceptible to markup discrepancies. Figures (ref) and (ref) depict histograms of plant-level averages of the ratio of intermediate costs to total variable costs, $\overline{\left(\frac{P_{M,t}M_{it}}{P_{M,t}M_{it}+W_{t}L_{it}}\right)}_i$, for the concrete products and electric audio equipment industries, respectively. The significant variation in intermediate cost shares implies that heterogeneous markups are not the primary explanation for the observed variation in intermediate input shares presented in Figures (ref) and (ref).

By comparing the degree of dispersion in input shares within the 2-digit industry classification with that within the 3-digit or 4-digit industry classification, we can examine the extent to which classifying industries at a more refined level helps control for heterogeneity in production technology.

Figures (ref), (ref), and (ref) contrast histograms of plant-level averages of intermediate input shares for ceramics and clay (2-digit), cement products (3-digit), and concrete products (4-digit) industries. These figures suggest that while dispersion decreases somewhat from the 2-digit to the 4-digit level, the degree of heterogeneity remains notably high even at the 4-digit industry classification. Likewise, Figures (ref), (ref), and (ref) reveal that the dispersion of material shares does not decline considerably when transitioning from electric parts, devices, and circuits (2-digit) to electric devices (3-digit) and subsequently to electric audio equipment (4-digit).

As illustrated in Table (ref), the 90th-10th percentile difference in intermediate input shares decreases from 0.38 (2-digit) to 0.28 (4-digit) for the ceramics and clay industry. In contrast, the 90th-10th percentile difference for the electric parts industry changes only marginally from 2-digit to 4-digit, ranging from 0.61 to 0.62. Similar patterns are observed for the plant-level averages of intermediate cost shares in variable costs. The 90th-10th percentile differences in intermediate cost shares for concrete products (4-digit) and electric audio equipment (4-digit) are substantial, measuring 0.27 and 0.67, respectively. Therefore, classifying industries at a more refined level does not substantially reduce the heterogeneity in output elasticities with respect to intermediate and labor inputs.

figure[figure omitted — 776 chars of source]
figure[figure omitted — 788 chars of source]
table[table omitted — 776 chars of source]
table[table omitted — 746 chars of source]
table[table omitted — 774 chars of source]

Analogous patterns are observed across various industries. Table (ref) displays the average differences between the 90th and 10th percentiles for all industries, classified at the 2-digit, 3-digit, and 4-digit levels, with their corresponding standard deviations presented in parentheses. The findings reveal that dispersion decreases only marginally when refining industry classification from the 2-digit to the 4-digit levels. In general, substantial dispersion persists even at the 4-digit classification level, indicating that output elasticities of variable inputs exhibit variation across plants, even when adopting a more detailed industry classification.

The implications in ((ref)) are valid exclusively under the Cobb-Douglas production function. For a more generalized production function, the elasticities of output for inputs are dependent on the levels of material, labor, and capital inputs, even without heterogeneity in production technology. Consequently, we investigate whether the intermediate input cost-to-output value ratio remains similar across firms, even after adjusting for variations in capital, labor, and intermediate inputs. To achieve this, we regress $\frac{P_{M,t}M_{it}}{P_{Y,t}Y_{it}}$ or $\frac{P_{M,t}M_{it}}{P_{M,t}M_{it}+W_{t}L_{it}}$ on second-order polynomials of the natural logarithm of materials, the number of workers, and capital, resulting in residuals denoted by $e_{it}$. Subsequently, we compute the plant-level average $\hat\xi_i := (1/25)\sum_{t=1986}^{2010} e_{it}$ to assess production technology heterogeneity, conditional on inputs.

Table (ref) presents the averages of the 90th-10th percentile differences in $\hat\xi_i$ for $\frac{P_{M,t}M_{it}}{P_{Y,t}Y_{it}}$ or $\frac{P_{M,t}M_{it}}{P_{M,t}M_{it}+W_{t}L_{it}}$ across all industries, classified at the 2-digit, 3-digit, and 4-digit levels. The findings reveal considerable variation in production technology after accounting for observable inputs, even at the 4-digit industry classification. This provides further evidence supporting the existence of heterogeneity in production technology.

The Model

We consolidate the notation as follows. Let $:=$ stand for "equal by definition". Bold letters denote vectors or matrices. For a continuous random variable $X$, a calligraphic letter $\mathcal{X}$ denotes its support, while its probability density function is represented by $g_{X}(x)$ for $x \in \mathcal{X}$.

Output, capital, intermediate inputs, labor input in effective units of labor, and total wage bills are denoted by $(Y_{it},K_{it},M_{it},L_{it},B_{it}) \in \mathcal{Y} \times \mathcal{K} \times \mathcal{M} \times \mathcal{L} \times \mathcal{B} \subset \mathbb{R}_{++}^5$, respectively, where $\mathcal{Y}$, $\mathcal{K}$, $\mathcal{M}$, $\mathcal{L}$, and $\mathcal{B}$ are the supports of the corresponding variables. We assume that $(Y_{it},K_{it},M_{it},L_{it},B_{it})$ are continuously distributed with strictly positive density on connected supports. We combine capital, intermediate, and labor inputs into a vector as ${\boldsymbol X}_{it} := (K_{it},M_{it},L_{it})' \in \mathcal{X} := \mathcal{K} \times \mathcal{M} \times \mathcal{L}$.

We allow firms' production technologies to differ beyond Hick's neutral productivity shocks. Specifically, we use a finite mixture specification to capture the unobserved heterogeneity in firms' production technologies as well as the process through which Hick's neutral productivity shocks evolve. We assume that there are $J$ unobserved types. Define the latent random variable $D_{i} \in\{1,2,...,J\}$ representing the type of firm $i$ such that $D_i=j$ if the production technology of firm $i$ is of the $j$-th type. In the following, the superscript $j$ indicates that the functions are specific to technology type $j$, while the subscript $t$ indicates that the functions are specific to period $t$.

For the $j$-th type of production technology at time $t$, the output is related to inputs as follows:

equation[equation omitted — 178 chars of source]

where $\epsilon_{it}$ is an idiosyncratic productivity shock with its density function $g_{\epsilon_t}^j(\cdot)$, and $\omega_{it}$ follows an exogenous first-order stationary Markov process given by:

align[align omitted — 148 chars of source]

where $\eta_{it}$ is an innovation to the productivity process. As indicated by the subscript $t$ in $F_{t}^j(\cdot)$, $h_t^j(\cdot)$, $g_{\epsilon_t}^j(\cdot)$, and $g_{\eta_t}^j(\cdot)$, production functions and productivity processes differ not only between latent types but also across periods. For example, this reflects type-specific aggregate shocks or type-specific biased technological changes.

We assume that labor input in effective units of labor, $L_{it}$, is not directly observable because firms differ in their labor quality and working hours, while we only observe the number of workers. Instead, $L_{it}$ is related to the number of workers, denoted by $\widetilde L_{it}$, as

equation[equation omitted — 75 chars of source]

Equation ((ref)) imposes a specific structure on how labor input in effective units of labor is related to the observed number of workers, where $\psi_t^j$ represents latent worker quality and working hours specific to type $j$. With this specification, measurement errors in observed labor input (i.e., the number of workers) are captured by the latent type-specific value of $\psi_t^j$.

The total wage bills, $B_{it}$, are related to the labor input in effective units as follows:

equation[equation omitted — 71 chars of source]

where $P_{L,t}$ represents the market wage. The random variable $v_{it}$ is a transitory wage shock known to firm $i$ at the time of choosing intermediate and labor inputs, while $\zeta_{it}$ is an idiosyncratic wage shock not included in the information set when firm $i$ selects these two inputs. Thus, $\zeta_{it} \in \boldsymbol{\mathcal{I}}_{it}$, but $v_{it} \notin \boldsymbol{\mathcal{I}}_{it}$.

We assume that firms make flexible choices regarding both $M_{it}$ and $L_{it}$ after observing their serially correlated productivity shock, $\omega_{it}$, but before observing $\epsilon_{it}$. In contrast, $K_{it}$ is predetermined at the end of the previous period, prior to the observation of the serially correlated productivity shock $\omega_{it}$. Denote the information available to a firm for making decisions on $M_{it}$ and $L_{it}$ by $\boldsymbol{\mathcal{I}}_{it}$.

We present the model assumptions. For continuous random variables $Z_{it}$ and $W_{it}$, we denote the probability density function and expectation conditional on $D_i=j$ as $g_{Z_t}^j(z_t) := g_{Z_t|D=j}(z_t|D_i=j)$ and $E^j[Z_{it}] := E[Z_{it}|D_i=j]$, respectively. Additionally, we denote the probability density function of $Z_{it}$ conditional on $D_i=j$ and $W_{it}=w_t$ as $g_{Z_t|W_t}^j(z_t|w_t)$. The unconditional probability density function of $Z_{it}$ is denoted by $g_{Z_t}(z_t)$.

assumption(a) Each firm belongs to one of the $J$ types, where the population probability of belonging to type $j$ is given by $\pi^j=\text{Pr}(D_i=j)$, and $J$ is known to econometricians. (b) A firm knows its type, i.e., $D_i\in \boldsymbol{\mathcal{I}}_{it}$.
assumption(a) $(v_{it},\omega_{it}) \in\boldsymbol{\mathcal{I}}_{it}$. (b) $(\epsilon_{it},\zeta_{it})\not\in\boldsymbol{\mathcal{I}}_{it}$. (c) For the $j$-th type, $\eta_{it}$ and $v_{it}$ are mean-zero i.i.d. continuous random variables on $\mathbb{R}$ with its probability density functions $g_{\eta_t}^j(\cdot)$ and $g_{v_t}^j(\cdot)$, respectively, while $(\epsilon_{it},\zeta_{it})$ is a mean-zero i.i.d. continuous random variable on $\mathbb{R}^2$ with the joint probability density function $g_{\epsilon_t,\zeta_t}^j(\cdot,\cdot)$. (d) The unconditional mean of $\omega_{it}$ is zero, i.e., $E_t^j[\omega_{it}]=0$ for every $t$.
assumption(a) $K_{it}\in \boldsymbol{\mathcal{I}}_{it}$ but $K_{it}\not\in \boldsymbol{\mathcal{I}}_{it-1}$. (b) the conditional density function of $K_{it}$ given $\boldsymbol{\mathcal{I}}_{t-1}$ is type specific and only depends on $K_{it-1}$ and $\omega_{it-1}$, i.e., $g_{K_t|\boldsymbol{\mathcal{I}}_{t-1},D}(K_{it}|\boldsymbol{\mathcal{I}}_{i,t-1},D_i=j) = g_{K_t|K_{it-1},\omega_{t-1}}^j(K_{it}|K_{it-1},\omega_{it-1})$.
assumption(a) $M_{it}$ and $L_{it}$ are chosen at time $t$ by maximizing expected profit conditional on $\boldsymbol{\mathcal{I}}_{it}$ as \begin{align} (M_{it},L_{it})&=(\mathbb{M}_t^j(K_{it},\omega_{it},v_{it}),\mathbb{L}_t^j(K_{it},\omega_{it},v_{it}))\nonumber\\ &: = \operatorname*{\arg\!\max}_{(M,L)\in\mathcal{M}\times\mathcal{L}} P_{Y,t}E^j[e^{\epsilon_{it}}]e^{\omega_{it}}F_t^j(K_{it},L,M)-P_{M,t} M -E^j[ e^{\zeta_{it}}]e^{v_{it}} P_{L,t} L, \end{align} where $(\mathbb{M}_t^j(K_{it},\omega_{it},v_{it}),\mathbb{L}_t^j(K_{it},\omega_{it},v_{it}))$ is a type-specific deterministic function of $(K_{it}, \omega_{it},v_{it})$. (b) For any given $K_{it}\in\mathcal{K}$, $(\mathbb{M}_t^j(K_{it},\omega_{it},v_{it}),\mathbb{L}_t^j(K_{it},\omega_{it},v_{it}))$ is invertible with respect to $(\omega_{it},v_{it})$ with probability one. (c) For any given $K_{it}\in\mathcal{K}$, the function $F_{t}^t(K_{it},L_{it},M_{it})$ is continuously differentiable and strictly concave in $(L_{it}, M_{it})$.
assumption(a) A firm is a price taker. (b) The intermediate input price $P_{M,t}$, the output price $P_{Y,t}$, and the market wage $P_{L,t}$ at time $t$ are common across firms. (c) $(P_{M,t},P_{Y,t},P_{L,t})\in \boldsymbol{\mathcal{I}}_{it} $ and $(P_{M,t},P_{Y,t})$ is known to an econometrician.
assumptionLabor input in effective unit of labour $L_{it}$ is not directly observable but $L_{it}$ is related to the observed number of workers as in ((ref)) with $\sum_{j=1}^J \pi^j e^{\psi_t^j}=1$

Under Assumptions (ref)-(ref), the information set at the time of choosing $M_{it}$ and $L_{it}$ is given by $\boldsymbol{\mathcal{I}}_{it} =\{D_i,\omega_{it},v_{it}, K_{it}, P_{M,t},P_{Y,t}, P_{L,t}, \boldsymbol V_{it-1},\boldsymbol V_{it-2},...\}$ with $\boldsymbol V_{it}=\{\zeta_{it},\epsilon_{it},\omega_{it},v_{it}, K_{it}, P_{M,t},P_{Y,t},P_{L,t}\}$.

Assumption (ref)(a) presumes that the number of types is known. Kasahara09 and Kasahara2014 discuss nonparametric identification of a lower bound for the number of types. Kasahara2019 and HaoKasahara2022 develop a likelihood-based testing procedure for the number of types in multivariate and panel data normal mixture models. Assumption (ref)(b) assumes that a firm is aware of its type.

Assumption (ref)(a)(b) asserts that $(\omega_{it},v_{it})$ is known when $L_{it}$ and $M_{it}$ are chosen, while $(\epsilon_{it},\zeta_{it})$ is not known when $L_{it}$ and $M_{it}$ are chosen. The presence of the wage shock $v_{it}$ in ((ref)) provides an additional source of variation for $L_{it}$ beyond $\omega_{it}$ and $K_{it}$; as a result, $L_{it}$ and $M_{it}$ are not collinear, avoiding the identification problem discussed by bond05 and Ackerberg15. Assumption (ref)(c) introduces notation for the probability density function of $\eta_{it}$, $v_{it}$, $\epsilon_{it}$, and $\zeta_{it}$, allowing for correlation between $\epsilon_{it}$ and $\zeta_{it}$. Assumption (ref)(d) is a normalization assumption to identify the location of $F_{t}^j$.

Assumption (ref)(a) presumes that $K_{it}$ is determined at time $t-1$, implying that $(\eta_{it},\omega_{it},v_{it})$ is not known when $K_{it}$ is chosen. Assumption (ref)(b) can be explicitly derived from a dynamic model of investment decisions with convex/non-convex adjustment costs under the first-order Markov productivity process ((ref)).

Assumption (ref)(a) introduces the demand function for $M_{it}$ and $L_{it}$, derived from the static profit maximization problem when $M_{it}$ and $L_{it}$ are flexibly chosen in each period. Assumption (ref)(b) holds when there is a one-to-one relationship between $(M_{it},L_{it})$ and $(\omega_{it},v_{it})$, except for a set of measure zero conditional on the value of $K_{it}$, and is satisfied in the case of a Cobb-Douglas function. Under Assumption (ref)(c), the first-order condition for the maximization problem ((ref)) characterizes the optimal choice for $L_{it}$ and $M_{it}$.

Assumption (ref)(a) posits that the firm has no market power. Under Assumption (ref)(b), the intermediate input price $P_{M,t}$ cannot be used for instrumenting $M_{it}$. When intermediate prices are exogenous and heterogeneous across firms, the production function could be identified using the intermediate input prices as instruments doraszelski2018measuring. In Assumption (ref)(c), an alternative approach assumes that a firm is subject to an idiosyncratic price shock $\xi_{it}$ such that, for example, $P_{Y,it}=\exp(\xi_{it})P_{Y,t}$ with $\xi_{it}\not\in \boldsymbol{\mathcal{I}}_{it}$, then $\xi_{it}$ plays a similar role to $\epsilon_{it}$. We may assume that $(P_{M,t},P_{Y,t})$ is not known to the econometrician by treating $P_{M,t}/P_{Y,t}$ as parameters to be estimated; in such a case, we can identify the production function up to scale.

Appendix (ref) discusses an alternative assumption to Assumption (ref) when a firm produces differentiated products and faces a demand function with constant price elasticity.

Assumption (ref) suggests that the quality of workers or the average working hours per worker differ across types and periods, as captured by the parameter $\psi_t^j$, which leads to the systematic difference in the average wage of workers latent types. The assumption that $\sum_{j=1}^J \pi^j e^{\psi_t^j}=1$ serves as a normalization for identification.

Nonparametric identification

Assume that we have panel data for firms $i=1,...,n$ over periods $t=1,...,T$ consisting of output, capital, intermediate inputs, the number of workers, and total wage bills, denoted by $(Y_{it},B_{it},K_{it},M_{it},\widetilde L_{it})\in \mathcal{Y}\times\mathcal{B}\times\mathcal{K}\times \mathcal{M}\times\tilde{\mathcal{L}}$, respectively. For brevity, define $\boldsymbol{\boldsymbol{\widetilde X}}:= (K_{it},M_{it},\widetilde L_{it}) \in \widetilde{ \mathcal{X}}:=\mathcal{K}\times \mathcal{M}\times\tilde{\mathcal{L}}$. Each firm's observation $\{Y_{it},B_{it},\widetilde {\boldsymbol X}_{it}\}_{t=1}^T$ is randomly sampled from a population distribution with a density function given by $g_{\{Y_{t},B_{t},\widetilde {\boldsymbol X}_{t}\}_{t=1}^T}(\{Y_{t},B_{t},\widetilde {\boldsymbol X}_{t}\}_{t=1}^T)$.

Let $g_{\epsilon_t}(\epsilon):= \int_{\mathbb{R}}g_{\epsilon_t,\zeta_t}(\epsilon,\zeta)d\zeta$ and $g_{\zeta_t}(\zeta):= \int_{\mathbb{R}} g_{\epsilon_t,\zeta_t}(\epsilon,\zeta)d\epsilon$ be the probability density functions of $\epsilon$ and $\zeta$, respectively. Under Assumptions (ref), (ref), (ref)(a), (ref)(a), and (ref), the first order conditions with respect to $M_{it}$ and $L_{it}$ for maximizing the expected static profit in Assumption (ref)(a) are given by:

equation[equation omitted — 239 chars of source]

where $F_{M,t}^j(\boldsymbol X):=\frac{\partial F_{t}^j(X) }{\partial M}$, $F_{L,t}^j(X):=\frac{\partial F_{t}^j(\boldsymbol X) }{\partial L}$, $E_t^j[e^{\epsilon}] :=\int e^{\epsilon} dg^j_{\epsilon_t}(\epsilon)d\epsilon$, and $E_t^j[e^{\zeta}] :=\int e^{\zeta} dg^j_{\zeta_t}(\zeta)d\zeta$. Rearranging equations ((ref)), ((ref)), ((ref)), and ((ref)) gives a system of equations:

equation[equation omitted — 660 chars of source]

where

align*[align* omitted — 361 chars of source]

For notational brevity, we drop the subscript $i$ in the rest of this section. Because $Y_t=\frac{P_{M,t}M_t}{S^m_t P_{Y,t}}$ and $B_t=\frac{S^\ell_t P_{M,t}M_t}{S^m_t}$, there exists a one-to-one relationship between $(Y_t,B_t)$ and $(S_t^m,S_t^\ell)$ given $\widetilde {\boldsymbol X}_t$ under Assumption (ref). Therefore, denoting $\boldsymbol S_t=(S_t^m,S_t^\ell)\in \mathcal{S}$, we consider $\{\boldsymbol S_t,\widetilde {\boldsymbol X}_t\}_{t=1}^T$ in place of $\{Y_{t},B_{t},\widetilde {\boldsymbol X}_{t}\}_{t=1}^T$ as our data.

Let $\boldsymbol Z_t := (\boldsymbol S_t,\widetilde {\boldsymbol X}_t)\in \mathcal{Z}:= \mathcal{S}\times\widetilde{\mathcal{X}}$. We assume that the population density function, denoted by $g_{\boldsymbol Z_1,...,\boldsymbol Z_T}(\{\boldsymbol z_{t}\}_{t=1}^T)$, is directly identified from the data. We are interested in identifying the model structure \[ \boldsymbol{\theta} := \left\{ \pi^j, \{g_{v_t}^j(\cdot), g_{\epsilon_t,\zeta_t}^j(\cdot), ,\Gamma_{M,t}^j(\cdot),\Gamma_{L,t}^j(\cdot),P_{L,t},\psi_t^j\}_{t=1}^T,\{F_t^j(\cdot)\}_{t=2}^T,\{h_t^j(\cdot),g_{\eta_t}^j(\cdot)\}_{t=3}^T\right\}_{j=1}^J \] from the population density function $\text{g}_{\boldsymbol Z_1,...,\boldsymbol Z_T}(\{\boldsymbol z_{t}\}_{t=1}^T)$ given a set of restrictions in ((ref)) under Assumptions (ref)-(ref).

We first establish the nonparametric identification of model structure $\boldsymbol\theta$ when $J=1$ as follows.

propositionSuppose that $J=1$ and Assumption (ref)-(ref) holds with $T\geq 3$. Then, $\boldsymbol\theta$ is uniquely determined from the population density function $\text{g}_{\boldsymbol Z_1,...,\boldsymbol Z_T}(\{\boldsymbol z_{t}\}_{t=1}^T)$.
remarkProposition (ref) extends the identification result of GNR to the setting where $L_{it}$ is contemporaneously determined rather than predetermined.

When $J\geq 2$, the probability density function of ${\{\boldsymbol Z_t\}_{t=1}^T}$ follows an $J$-term mixture distribution

equation[equation omitted — 327 chars of source]

The number of type $J$ is defined to be the smallest integer $J$ such that the density function of ${\{\boldsymbol Z_t\}_{t=1}^T}$ admits the representation ((ref)).

propositionSuppose that Assumptions (ref)-(ref) hold. Then, the probability density function of ${\{\boldsymbol Z_t\}_{t=1}^T}$ defined in ((ref)) can be written as \begin{align} g_{\boldsymbol Z_1,...,\boldsymbol Z_T}(\{\boldsymbol z_t\}_{t=1}^T) & = \sum_{j=1}^J \pi^j g_{{\boldsymbol Z}_1}^j(\boldsymbol z_1) \prod_{t=2}^T g_{{\boldsymbol Z}_t|{\boldsymbol Z}_{t-1}}^j({\boldsymbol z}_t|{\boldsymbol z}_{t-1})\\ &=\sum_{j=1}^J \pi^j \left(g_{\boldsymbol S_1|\widetilde {\boldsymbol X}_1}^j(\boldsymbol s_1|\widetilde {\boldsymbol x}_1)\prod_{t=2}^Tg_{\boldsymbol S_t|\widetilde {\boldsymbol X}_t}^j(\boldsymbol s_t|\widetilde {\boldsymbol x}_{t})\right) \times \left( g_{\widetilde {\boldsymbol X}_1}^j(\widetilde {\boldsymbol x}_1) \prod_{t=2}^Tg_{\widetilde {\boldsymbol X}_t|\widetilde {\boldsymbol X}_{t-1}}^j(\widetilde {\boldsymbol x}_t|\widetilde {\boldsymbol x}_{t-1})\right), \end{align} where $g_{\boldsymbol S_t|\widetilde {\boldsymbol X}_t}^j(\boldsymbol s_t|\widetilde {\boldsymbol x}_t)$ is the type $j$'s conditional probability density function of $\boldsymbol S_t$ given $\widetilde {\boldsymbol X}_t=\widetilde {\boldsymbol x}_t$, $g_{\widetilde {\boldsymbol X}_1}^j(\widetilde {\boldsymbol x}_1)$ is the type $j$'s marginal probability density function of $\widetilde {\boldsymbol X}_1$, $g_{{\boldsymbol Z}_1}^j(\boldsymbol z_1) =g_{\widetilde {\boldsymbol X}_1}^j(\widetilde {\boldsymbol x}_1)g_{\boldsymbol S_1|\widetilde {\boldsymbol X}_1}^j(\boldsymbol s_t|\widetilde {\boldsymbol x}_1)$, and $g_{{\boldsymbol Z}_t|{\boldsymbol Z}_{t-1}}^j({\boldsymbol z}_t|{\boldsymbol z}_{t-1})=g_{\boldsymbol S_t|\widetilde {\boldsymbol X}_t}^j(\boldsymbol s_t|\widetilde {\boldsymbol x}_{t}) g_{\widetilde {\boldsymbol X}_t|\widetilde {\boldsymbol X}_{t-1}}^j(\widetilde {\boldsymbol x}_t|\widetilde {\boldsymbol x}_{t-1})$.

Therefore, under the stated model assumption, ${\{\boldsymbol Z_t\}_{t=1}^T}$ follows a first order Markov process within subpopulation specified by type. The result of Proposition (ref) allows us to establish the nonparametric identification of $\{\pi^{j},g_{\boldsymbol Z_1}^{j}(\boldsymbol z_1), g_{\boldsymbol Z_2|\boldsymbol Z_1}^{j}(\boldsymbol z_2|\boldsymbol z_1), ... , g_{\boldsymbol Z_T|\boldsymbol Z_{T-1},...,\boldsymbol Z_1}^{j}(\boldsymbol z_T|\boldsymbol z_{T-1}) \}_{j=1}^J$ by extending the argument in Kasahara09, carolletal10jns, and Hu12.

We now establish identification when $T=4$. Define

equation[equation omitted — 962 chars of source]

where $\bar \lambda^j_2(\boldsymbol a,{\boldsymbol z}_2):=\pi^j g_{{\boldsymbol Z}_2|{\boldsymbol Z}_1}^j({\boldsymbol z}_2|\boldsymbol a)g_{{\boldsymbol Z}_1}^j(\boldsymbol a)$, $\lambda^j_3({\boldsymbol z}_3|{\boldsymbol z}_2):= g_{{\boldsymbol Z}_3|{\boldsymbol Z}_2}^j({\boldsymbol z}_3|{\boldsymbol z}_2)$, and $\lambda^j_{4}(\boldsymbol b|{\boldsymbol z}_3):=g_{{\boldsymbol Z}_4|{\boldsymbol Z}_3}^j(\boldsymbol b|{\boldsymbol z}_3)$.

assumptionThere exists a value $\boldsymbol z_3^*$ that satisfies the following condition: for every $\boldsymbol z_3\in\mathcal{Z}_3$, we can find $(\bar {\boldsymbol z}_2,\check{\boldsymbol z}_2,\bar {\boldsymbol z}_3)\in \mathcal{Z}_2\times \mathcal{Z}_2\times \mathcal{Z}_3$, $(\boldsymbol a_1,...,\boldsymbol a_{J})\in \mathcal{Z}_1^{J}$ and $(\boldsymbol b_1,...,\boldsymbol b_{J-1})\in \mathcal{Z}_4^{J-1}$ such that (a) $\boldsymbol L_{{\boldsymbol z}_3^*}$, $\boldsymbol L_{{\boldsymbol z}_3}$, $\boldsymbol L_{\bar {\boldsymbol z}_3}$, $\bar {\boldsymbol L}_{\check{\boldsymbol z}_2}$, and $\bar {\boldsymbol L}_{\bar {\boldsymbol z}_2}$ are non-singular, and (b) all the diagonal elements of ${\boldsymbol D}_{\boldsymbol z_3,\boldsymbol z_3}:= {\boldsymbol D}_{{\boldsymbol z}_3|\check{\boldsymbol z}_2} {\boldsymbol D}_{\bar {\boldsymbol z}_3|\check{\boldsymbol z}_2} ^{-1} {\boldsymbol D}_{\bar {\boldsymbol z}_3|\bar {\boldsymbol z}_2}{\boldsymbol D}_{{\boldsymbol z}_3|\bar {\boldsymbol z}_2}^{-1}$ take distinct values. Furthermore, (c) for every $(\boldsymbol z_2,\boldsymbol z_3)\in\mathcal{Z}_2\times \mathcal{Z}_3$, $g^j_{{\boldsymbol Z}_3|{\boldsymbol Z}_2}({\boldsymbol z}_3|{\boldsymbol z}_2)> 0$ for $j=1,...,J$.
propositionSuppose that Assumptions (ref)-(ref) hold and $T\geq 4$. Then, \\ $\{\pi^j, g_{{\boldsymbol Z}_1}^j({\boldsymbol z}_1),g^j_{{\boldsymbol Z}_2|{\boldsymbol Z}_1}({\boldsymbol z}_2|{\boldsymbol z}_1),..., g_{{\boldsymbol Z}_T|{\boldsymbol Z}_{T-1}}^j({\boldsymbol z}_T|{\boldsymbol z}_{T-1})\}_{j=1}^J$ is uniquely determined from $g_{{\boldsymbol Z}_1,...,{\boldsymbol Z}_T}(\{{\boldsymbol z}_{t}\}_{t=1}^T)$ up to a common permutation of the latent types.
remarkUnder the additional assumption of stationarity, i.e., the conditional density function $g_{{\boldsymbol Z}_t|{\boldsymbol Z}_{t-1}}^j({\boldsymbol z}_{t}|{\boldsymbol z}_{t-1})$ does not depend on $t$ for $t=2,...,T$, Kasahara09 establish the nonparametric identification of the model ((ref)) when $T=6$ while Hu12 show that $T=4$ suffices for identification.
remarkConsidering serially correlated continuous unobserved variables $\{X_t^*\}_{t=1}^T$, Hu12 analyze the nonparametric identification of the model \[ g_{{\boldsymbol Z}_1,...,{\boldsymbol Z}_T}(\{{\boldsymbol z}_{t}\}_{t=1}^T)= \int g_{{\boldsymbol Z}_1|X_1^*}({\boldsymbol z}_{1},x_{1}^*)\prod_{t=2}^Tg_{{\boldsymbol Z}_t,X_t^*|{\boldsymbol Z}_{t-1},X_{t-1}^*}({\boldsymbol z}_{t},x_{t}^*|{\boldsymbol z}_{t-1},x_{t-1}^*)d(\{x_t^*\}_{t=1}^T). \] Given panel data ${\{\boldsymbol {\boldsymbol Z}_t\}_{t=1}^T}$ with $T=5$, Theorem 1 and Corollary 1 of Hu12 state that, under their Assumptions 1-4, $g_{{\boldsymbol Z}_3,X_3^*}({\boldsymbol z}_{3},x_{3}^*)$, $g_{{\boldsymbol Z}_4,X_4^*|{\boldsymbol Z}_3,X_3^*}({\boldsymbol z}_{4},x_{4}^*|{\boldsymbol z}_{3},x_{3}^*)$, and\\ $g_{{\boldsymbol Z}_5,X_5^*|{\boldsymbol Z}_4,X_4^*}({\boldsymbol z}_{5},x_{5}^*|{\boldsymbol z}_{4},x_{4}^*)$ are nonparametrically identified, but the identification of $g_{{\boldsymbol Z}_1,X_1^*}({\boldsymbol z}_{1},x_{1}^*)$, $g_{{\boldsymbol Z}_2,X_2^*|{\boldsymbol Z}_1,X_1^*}({\boldsymbol z}_{2},x_{2}^*|{\boldsymbol z}_{1},x_{1}^*)$, and $g_{{\boldsymbol Z}_3,X_3^*|{\boldsymbol Z}_2,X_2^*}({\boldsymbol z}_{3},x_{3}^*|{\boldsymbol z}_{2},x_{2}^*)$ remains unresolved. Our Proposition (ref) shows a new identification result that, for a model in which unobserved heterogeneity is discrete and finite, we can nonparametrically identify the type-specific distribution of $\{{\boldsymbol Z}_{t}\}_{t=1}^T$, including the first two periods of the data, from $T=4$ periods of panel data without imposing stationarity.
remarkIn the identification argument of Kasahara09, the type-specific distribution is identified only up to an arbitrary ordering of the latent types that differs across different evaluation points in $\{{\boldsymbol z}_1,...,{\boldsymbol z}_T\}$. Our proof of Proposition (ref) on the identification of the common order of the latent types is based on HigginsJochmans21.
remarkAssumption (ref)(a) assumes the rank condition of matrices ${\boldsymbol L}_{{\boldsymbol z}^*_3}$, ${\boldsymbol L}_{{\boldsymbol z}_3}$, ${\boldsymbol L}_{\bar {\boldsymbol z}_3}$, $\bar {\boldsymbol L}_{\check{\boldsymbol z}_2}$, and $\bar {\boldsymbol L}_{\bar {\boldsymbol z}_2}$ defined in ((ref)), of which elements are constructed by evaluating $g^j_{{\boldsymbol Z}_4|{\boldsymbol Z}_3}({\boldsymbol z}_{4}|{\boldsymbol z}_3)$ and $\pi^jg^j_{{\boldsymbol Z}_2|{\boldsymbol Z}_1}({\boldsymbol z}_{2}|{\boldsymbol z}_1)g^j_{{\boldsymbol Z}_1}({\boldsymbol z}_1)$ at different points. These conditions are similar to the assumption stated in Proposition 1 of Kasahara and Shimotsu (2009), implying that all columns in these matrices must be linearly independent. For example, because each column of ${\boldsymbol L}_{{\boldsymbol z}_3}$ represents the type-specific conditional density function of $\boldsymbol z_4$ across different values of $\boldsymbol z_4$ given $\boldsymbol z_3$, the changes in the value of $\boldsymbol z_4$ must induce sufficiently different changes in the values of conditional density function across types. One needs to find only one set of values $(\check{\boldsymbol z}_2,\bar {\boldsymbol z}_2,\bar {\boldsymbol z}_3)\in \mathcal{Z}_2^2\times \mathcal{Z}_3$ and one set of $J-1$ and $J$ points of ${\boldsymbol Z}_1$ and ${\boldsymbol Z}_4$ to construct nonsingular $\boldsymbol L_{{\boldsymbol z}_3^*}$, $\boldsymbol L_{{\boldsymbol z}_3}$, $\boldsymbol L_{\bar {\boldsymbol z}_3}$, $\bar {\boldsymbol L}_{\check{\boldsymbol z}_2}$, and $\bar {\boldsymbol L}_{\bar {\boldsymbol z}_2}$ for each ${\boldsymbol z}_3\in \mathcal{Z}_3$ and these rank conditions are not stringent when ${\boldsymbol Z}_t$ has continuous support. The identification of $g_{{\boldsymbol Z}_4|{\boldsymbol Z}_3}^j({\boldsymbol z}_4|{\boldsymbol z}_3)$ and $\pi^jg_{{\boldsymbol Z}_2|{\boldsymbol Z}_1}^j({\boldsymbol z}_2|{\boldsymbol z}_1)g_{{\boldsymbol Z}_1}^j({\boldsymbol z}_1)$ at all other points of ${\boldsymbol Z}_4$, $\boldsymbol Z_2$, and ${\boldsymbol Z}_1$ follows without any further requirement on the rank condition.

Once the type-specific distribution of $\{{\boldsymbol Z}_t\}_{t=1}^T$ is identified, we can use the argument in the proof of Proposition (ref) to prove nonparametric identification for the model structure of each type.

propositionSuppose that Assumptions (ref)-(ref) hold and $T\geq 4$. Then, $\boldsymbol\theta$ is uniquely determined from $g_{{\boldsymbol Z}_1,...,{\boldsymbol Z}_T}(\{{\boldsymbol z}_{t}\}_{t=1}^T)$.

Therefore, type-specific production functions, as well as the distribution of unobserved variables, can be identified nonparametrically. In the estimation, we focus on the case where the type-specific production functions are Cobb-Douglas.

example[Random Coefficients Model] Consider a Cobb-Douglas production function with time-varying random coefficients: \begin{equation} \tilde f_t^j(\tilde {\boldsymbol X}_t) = \beta_{0,t}^j + \beta_{k,t}^j \ln K_{t} + \beta_{m,t}^j \ln M_{t}+ \beta_{\ell,t}^j( \psi_t^j + \ln \tilde L_t), \end{equation} where $\tilde f_t^j(\tilde {\boldsymbol X}_t) := \ln F_t^j(K_t,e^{\psi_t^j} \widetilde L_t,M_t)$ while the intermediate and labor cost share equations are given by \begin{align*} \ln S_t ^m& = \ln(\beta_{m,t}^j) + \ln E_t^j[e^{\epsilon}] - \epsilon_t,\quad \ln S_t ^\ell-\ln S_t ^m = \ln(\beta_{\ell,t}^j/\beta_{m,t}^j) - \ln \left(E_t^j[e^{\zeta}]\right) + \zeta_t,\\ \ln M_{it}- \ln \tilde L_{it} & = \ln (P_{L,t}/P_{M,t})+ \ln \left(\beta_{m,t}^j/\beta_{\ell,t}^j\right) +\ln \left(E_t^j[e^{\zeta}]\right) + \psi_t^j +v_{it}, \end{align*} Under Assumptions (ref)-(ref), $\boldsymbol\theta=\{\pi^j, \{g^j_{v_t}(\cdot), \beta_{m,t}^j, \beta_{\ell,t}^j,g_{\epsilon_t,\zeta_t}^j(\cdot),\psi_t^j\}_{t=1}^4,\{\beta_{0,t}^j, \beta_{k,t}^j \}_{t=2}^4, \{h_t^j(\cdot), g^j_{\eta_t}(\cdot)\}_{t=3}^4 \}$ for $j=1,...,J$ is nonparametrically identified from the panel data $\{\boldsymbol S_t,\widetilde{\boldsymbol X}_t\}_{t=1}^4$.

In Appendix (ref), we discuss the sufficient conditions under which Assumption (ref) holds when the production function is Cobb-Douglas as in Example (ref).

Estimation of production function with finite mixture random coefficients models

In this section, we present a finite mixture model of the random coefficient Cobb-Douglas production function, based on our nonparametric identification analysis and with consideration for computational efficiency. We develop a penalized maximum likelihood estimator for this model.

Let us denote the logarithmic values of $(Y_{it},K_{it},M_{it},\tilde L_{it},S_{it}^m,S_{it}^\ell,B_{it})$ using the corresponding lowercase letters, such that $(y_{it},k_{it},m_{it},\tilde \ell_{it},s_{it}^m,s_{it}^\ell,b_{it})$, where $y_{it} :=\log Y_{it}$, and so on. Define $\boldsymbol s_{it}:=(s_{it}^m,s_{it}^\ell)$ and $\tilde {\boldsymbol x}_{it}:=(k_{it},m_{it},\tilde \ell_{it})$. For estimation purposes, we assume that the data generation follows the parametric assumptions outlined below.

assumption(a) $T$ is fixed at $T\geq 4$ and $N\rightarrow\infty$. (b) Equation ((ref)) holds with \begin{equation} Y_{it}= F_t^j(K_{it},M_{it},e^{\psi_t^j} \widetilde L_{it})e^{\omega_{it}+ \epsilon_{it}}\ with\ F_t^j(K_{it},M_{it},e^{\psi_t^j} \widetilde L_{it})=\exp((\beta_{0,t}^j + \beta_{\ell}^j \psi_t^j)+ \beta_{k}^j k_{it} + \beta_m^j m_{it} + \beta_{\ell}^j \tilde \ell_{it} ). \end{equation} (c) $(\epsilon_{it},\zeta_{it})^\top |D_i=j \overset{d}{\sim} N\left(\boldsymbol 0, \boldsymbol\Sigma_{\epsilon\zeta}\right)$ with $\boldsymbol\Sigma_{\epsilon\zeta}= \begin{pmatrix} (\sigma_\epsilon^j)^2&\rho_{\epsilon\zeta}^j\sigma_\epsilon^j\sigma_\zeta^j\\ \rho_{\epsilon\zeta}^j\sigma_\epsilon^j\sigma_\zeta^j&(\sigma_\zeta^j)^2 \end{pmatrix}$, $g_{\eta}^j(\eta)=\phi(\eta/\sigma_{\eta}^j)/\sigma_{\eta}^j$, and $g_{v}^j(v)=\phi(v/\sigma_{v}^j)/\sigma_{v}^j$, where $\phi(t) = \exp(-t^2/2)/\sqrt{2}\pi$. Furthermore, we assume $h_t^j(\omega_{it})=\rho_\omega^j \omega_{it}$ in ((ref)) so that \begin{equation} \omega_{it}=\rho_\omega^j \omega_{it-1} + \eta_{it}. \end{equation} (d) Conditional on being type $j$, $k_{it}$ given $(k_{it-1},\omega_{it-1})$ is normally distributed with mean $\rho_{k0}^j+\rho_{kk}^j k_{it-1}+ \rho_{k\omega}^j\omega_{it-1}$ and variance $(\sigma_k^j)^2$ while the distribution of $(k_{i1},\omega_{i1})$ follows a bivariate normal distribution with mean $\boldsymbol{\mu}_1^j$ and variance $\boldsymbol\Sigma_1^j$.

Assumption (ref)(a) posits that the length of panel data is short, while Assumption (ref)(b) enforces the Cobb-Douglas functional form assumption. Assumptions (ref)(c) and (ref)(d) impose Gaussian distribution assumptions under Assumptions (ref) and (ref), where $\omega_{it}$ follows a first-order autoregressive (AR(1)) process.

In equation ((ref)), since $\ln L_{it} = \psi_t^j +\tilde \ell_{it}$, the intercept term encompasses both $\beta_{0,t}^j$ and $\beta_{\ell}^j\psi_t^j$. The latter term captures the variation in worker quality across types. The normality assumption in Assumptions (ref)(c) and (ref)(d) could potentially be relaxed; for instance, by employing the maximum smoothed likelihood estimator of finite mixture models proposed by lhc11, in which the type-specific distribution of $\epsilon_{it}$ and $\zeta_{it}$ is nonparametrically specified. Additionally, kasaharashimotsu15jasa develop a likelihood-based procedure to test the number of components in normal mixture regression models.

Suppose we have a random sample of $n$ independent observations $\{\{\boldsymbol S_{it},\widetilde {\boldsymbol X}_{it}\}_{t=1}^T\}_{i=1}^n$ from the $J$-component mixture model $\sum_{j=1}^J \pi^j g_{\{\boldsymbol S_t,\widetilde {\boldsymbol X}_t\}_{t=1}^T}^j(\{\boldsymbol s_{t},\widetilde {\boldsymbol x}_{t}\}_{t=1}^T)$ that satisfies Assumptions (ref)-(ref). We propose a penalized maximum likelihood estimator (PMLE) that directly maximizes the log-likelihood function of a finite mixture model of production functions. The likelihood function is a parametric version of ((ref)). To address the issue of unbounded likelihood for a normal mixture model hartigan85book, we introduce a penalty term to the log-likelihood function. The maximum likelihood estimator, which leverages distributional information, is consistent even when $T$ is small, provided $T \geq 4$. Given the nonparametric identification result established in Proposition (ref), if the parametric assumptions are invalid and the parametric model is misspecified, the parametric maximum likelihood estimator converges in probability to the pseudo-true value of the parameter that minimizes the Kullback-Leibler Information Criterion between the density of the parametric model and the true population density White1982.

Our estimation procedure is based on the two-stage identification proof from Proposition (ref). To address the computational complexity of maximizing the log-likelihood function for the finite mixture model, the EM algorithm is employed.

Under Assumptions (ref)-(ref), (ref), the first order conditions for the expected profit maximization imply that

align[align omitted — 346 chars of source]

where $\alpha_t:= \ln (P_{L,t}/P_{M,t})$ and ((ref)) follows from ((ref)), $v_{it}=b_{it}-(\psi^j +\tilde \ell_t + \ln P_{L,t}+ \zeta_{it})$, and $s_{it}^\ell-s_{it}^m= b_{it} - (\ln P_{M,t}+ m_{it})$.

Collect the model parameter $\boldsymbol\theta$ into $\boldsymbol\pi$, $\boldsymbol{\theta}_1$ and $\boldsymbol{\theta}_2$ as \[ \boldsymbol\theta:=(\boldsymbol\pi',\boldsymbol\theta_1',\boldsymbol\theta_2')'\quad\text{with}\quad\boldsymbol{\theta}_1= (\boldsymbol{\alpha}',(\boldsymbol \theta_1^1)',...,(\boldsymbol{\theta}_1^J)')'\ \text{and}\ \boldsymbol{\theta}_2 = ((\boldsymbol{\theta}_2^1)',...,(\boldsymbol{\theta}_2^J)')', \] where $\boldsymbol\theta_{1}^j =(\beta_m^j,\beta_\ell^j, \psi^j, (\sigma_{\epsilon}^j)^2,(\sigma_{\zeta}^j)^2,(\sigma_v^j)^2)'$ and $\boldsymbol\theta_2^j=(\beta_{2}^j,...,\beta_T^j, \beta_k^j,(\boldsymbol\mu_1^j)',\text{vech}(\boldsymbol\Sigma_1^j)',\rho_{k0}^j,\rho_{kk}^j,$\\ $\rho_{k\omega}^j,(\sigma_k^j)^2,\rho_{\omega}^j,(\sigma_\eta^j)^2)'$ for $j=1,...,J$, and $\boldsymbol{\alpha}=(\alpha_1,....,\alpha_T)'$.

Denote $\boldsymbol\theta^j:=((\boldsymbol\theta_1^j)',(\boldsymbol\theta_2^j)')'$. Then, under Assumption (ref), we may write the probability density function of $\{\boldsymbol s_{it},\tilde{\boldsymbol x}_{it}\}_{t=1}^T$ for type $j$ as

equation[equation omitted — 598 chars of source]

where the exact expression for $L_{1i}(\boldsymbol\theta_1^j,\boldsymbol{\alpha})$ and $L_{2i}(\boldsymbol \theta_2^j,\boldsymbol \theta_1^j)$ is derived below.

To deal with the issue of unbounded log-likelihood function of a normal mixture model hartigan85book,HaoKasahara2022, we estimate the model parameter $\boldsymbol\theta$ by a penalized likelihood method proposed by chentan09jmva. Let $(\hat\sigma_{\epsilon,0}^2,\hat\sigma_{\zeta,0}^2,\hat\sigma_{v,0}^2, \hat\sigma_{k,0}^2,\hat\sigma_{\eta,0}^2, \hat{\boldsymbol\Sigma}_{1,0})$ be the estimator of $(\sigma_{\epsilon}^2,\sigma_{\zeta}^2,\sigma_{v}^2, \hat\sigma_{k}^2,\hat\sigma_{\eta}^2, \hat{\boldsymbol\Sigma}_{1})$ for the one-component model with $J=1$. Then, we consider the following penalized maximum likelihood estimator (PMLE):

equation[equation omitted — 391 chars of source]

where \[ Q_i(\boldsymbol\theta):= \log\left(\sum_{j=1}^J \pi^j L_{i}(\boldsymbol \theta^j,\boldsymbol\alpha)\right) \quad\text{with}\quad L_{i}(\boldsymbol \theta^j,\boldsymbol\alpha):=L_{1i}({\boldsymbol \theta}_1^j,{\boldsymbol{\alpha}})L_{2i}({\boldsymbol\theta}_1^j,{\boldsymbol\theta}_2^j) \] and

align[align omitted — 435 chars of source]

The expression for $L_{1i}({\boldsymbol \theta}_1^j,{\boldsymbol{\alpha}})$ and $L_{2i}({\boldsymbol\theta}_1^j,{\boldsymbol\theta}_2^j)$ are derived below.

To reduce the computational burden of finding the penalized maximum likelihood estimator, we follow a three-stage procedure.

In the first stage, from equations ((ref))-((ref)), we can express $\epsilon_{it}$, $\zeta_{it}$, and $v_{it}$ as a function of $\boldsymbol s_{it}$, $\tilde \ell_{it}-m_{it}$, $\boldsymbol\theta_1^j$, and $\alpha_t$ as

align[align omitted — 495 chars of source]

Then, we estimate $\boldsymbol\theta_1$ by maximizing the log likelihood function as

align*[align* omitted — 442 chars of source]

where, under Assumption (ref)(b), the likelihood function $L_{1i}(\boldsymbol\theta_1^j,\boldsymbol{\alpha})$ is given by

align*[align* omitted — 631 chars of source]

with $ \epsilon^*(\boldsymbol s_{it};\boldsymbol \theta_1^j)$, $\zeta^*(\boldsymbol s_{it};\boldsymbol \theta_1^j) $, and $v^*(\tilde \ell_{it}-m_{it} ;\boldsymbol\theta_1^j,\alpha_t)$ defined in ((ref))-((ref)).

In the second stage, from ((ref)), $\epsilon_{it} = E[s_{it}^m| \tilde{\boldsymbol x}_{it}]-s_{it}^{m}$, and $y_{it}+s_{it}^m=m_{it}+\ln (P_{M,t}/P_{Y,t})$, we have

equation[equation omitted — 250 chars of source]

where $\beta_t^j:=\beta_{0,t}^j +\ln (P_{M,t}/P_{Y,t})-\ln \beta_m^j - 0.5(\sigma_{\epsilon}^j)^2$.

In view of equation ((ref)), by a change of variables, we may relate the density function of $m_{it}$ conditional on $\tilde \ell_{it}-m_{it}$ and $k_{it}$ to the density function of $\omega_{it}$, denoted by $g_{\omega,t}$, as $g_t^j(m_{it}|\ell_{it}-m_{it},k_{it})=(1-\beta^j_m-\beta^j_\ell)g_{\omega,t}^j(\omega_t^*(m_{it},\tilde \ell_{it}-m_{it},k_{it};\boldsymbol\theta^j))$. Then, from ((ref))-((ref)) and Assumptions (ref)-(ref), we have

align[align omitted — 548 chars of source]

where $g_{\omega|k,1}^j(\omega_{i1} |k_{i1})$ is the density function of $\omega_{i1}$ conditional on $k_{i1}$, $g_{k,t}^j(k_{it}|k_{it-1},\omega_{it-1})$ is the density function of $k_{it}$ given $(k_{it-1},\omega_{it-1})$, $\omega_{it}^*(\boldsymbol\theta^j):=\omega_t^*(m_{it},\ell_{it}-m_{it},k_{it};\boldsymbol\theta^j) $, and

equation[equation omitted — 155 chars of source]

Therefore, under Assumption (ref), it follows from ((ref)) and ((ref))-((ref)) that

align*[align* omitted — 485 chars of source]

where

align*[align* omitted — 737 chars of source]

Given the first stage estimate $\tilde{\boldsymbol\theta}_1$, the second stage estimates parameters $\boldsymbol{\pi}$ and $\boldsymbol\theta_2$ by maximizing the log-likelihood function as \[ (\tilde{\boldsymbol{\pi}}_2,\tilde{\boldsymbol\theta}_2) = \operatorname*{\arg\!\max}_{\boldsymbol{\pi},\boldsymbol\theta_2} \sum_{i=1}^n \log\left(\sum_{j=1}^J \pi^j L_{1i}(\tilde{\boldsymbol\theta}_1^j,\tilde{\boldsymbol{\alpha}}) L_{2i}(\tilde{\boldsymbol\theta}_1^j,\boldsymbol\theta_2^j) \right)+ \sum_{j=1}^J \left\{ \sum_{s\in\{k,\eta\}} p_n((\sigma_s^j)^2;\hat\sigma_{s,0}^2) + p_n(\boldsymbol\Sigma_1^j;\hat{\boldsymbol\Sigma}_{1,0})\right\}. \]

Finally, using $\tilde{\boldsymbol{\theta}} := (\tilde{\boldsymbol{\pi}}_2', \tilde{\boldsymbol{\theta}}_1', \tilde{\boldsymbol{\theta}}_2')'$ as an initial value, we obtain the PMLE $\hat{\boldsymbol{\theta}}$ by maximizing the full-information penalized log-likelihood function as in Equation ((ref)) using the EM algorithm.

Let the true value of the model parameter be denoted by $\boldsymbol{\theta}^*$. When the number of types $J$ is correctly specified, the Fisher Information matrix is given by \[ \boldsymbol{I}(\boldsymbol{\theta}^*) := -\mathbb{E}\left[ \nabla_{\boldsymbol{\theta}\boldsymbol{\theta}'} Q_i(\boldsymbol{\theta}^*) \right] = \mathbb{E}\left[ \nabla_{\boldsymbol{\theta}}Q_i(\boldsymbol{\theta}^*) \nabla_{\boldsymbol{\theta}'} Q_i(\boldsymbol{\theta}^*) \right], \] which is positive definite. The following proposition demonstrates that the PMLE is consistent and asymptotically normal.

propositionSuppose that Assumptions (ref)-(ref) hold. Then, $\hat{\boldsymbol\theta}\overset{p}{\rightarrow} \boldsymbol \theta^*$ and $\sqrt{n}(\hat{\boldsymbol\theta}-\boldsymbol\theta^*) \overset{d}{\rightarrow} N(\boldsymbol 0, \boldsymbol{{I}}(\boldsymbol\theta^*))$.

Empirical Application

Data

We utilize plant-level panel data from the Census of Manufacture of Japan spanning 1986-2010. This dataset encompasses production information for manufacturing plants in Japan. Our analysis focuses on plants with 30 or more employees, as detailed data are consistently available only for these establishments.\footnote{The survey employs distinct questionnaires based on plant size: 1. Plants with 30 or more employees; 2. Plants with 4-29 employees; 3. Plants with 1-3 employees. The questionnaire for plants with 30 or more employees provides more comprehensive information. For instance, beginning in 2000, the census collects fixed asset data every five years, rather than annually, for plants with fewer than 30 employees.}

At the 4-digit industry classification level accessible in the Census of Manufacture, we identify a total of 276 industries. In our empirical application, we primarily concentrate on concrete products and electric audio equipment for two reasons: 1. Both industries have a substantial number of observations; 2. The former exhibits relatively small variation in intermediate input share, while the latter displays significant variation, as demonstrated in Section (ref). Thus, examining these two industries proves useful for assessing the significance of unobserved heterogeneity.

Output ($Y$) is defined as the sum of shipments, revenue from repair and maintenance services, and revenue from performing subcontracted work. Initial capital value ($K$) is determined as the fixed asset value minus land, and subsequent capital values are constructed using the perpetual inventory method. The observed labor input ($\tilde{L}$) is represented by the number of employees. The intermediate input ($M$) is defined as the sum of material input, energy input, and subcontracting expenses for consigned production.

Flow data, such as shipments and various production costs, pertain to the calendar year. The number of employees refers to the value at the end of the year, while the stock of fixed assets corresponds to the beginning of the period. Table (ref) presents summary statistics for the variables employed in our empirical analysis.

table[table omitted — 995 chars of source]

Estimation of Production Function

This section presents estimation results for a random-coefficient Cobb-Douglas production function featuring three technology types and two unobserved labor types within each technology type.

Table (ref) and Table (ref) present the parameter estimates for the concrete products and electric audio equipment industries, considering both the unobserved heterogeneity case ($J = 3 \times 2 = 6$) and the homogeneous case ($J = 1$). The estimated coefficients in both industries demonstrate economically significant differences in output elasticities associated with labor, capital, and intermediate inputs across various firm types.

Comparing the two industries, the variation in $\hat \beta_m^j$ across types is more substantial for electric audio equipment than for concrete products, which aligns with the dispersion of intermediate input shares discussed in Section (ref). As $\hat \beta_\ell^j$ and $\hat\beta_k^j$ also exhibit variation across types, the ratio of output elasticities between capital and labor, $\hat\beta_k^j/\hat\beta_\ell^j$, varies as well. For electric audio equipment, the value of $\hat\beta_k^j/\hat\beta_\ell^j$ ranges from 0.29 (Type 1 and 2) to 1.14 (Type 3 and 4), while for concrete products, it spans from 0.56 to 0.78. As demonstrated below, the value of $\hat\beta_k^j/\hat\beta_\ell^j$ serves as a crucial determinant of the capital investment response to productivity $\hat\omega_{it}$.

Furthermore, the productivity growth processes display variations across latent types, as indicated by the estimated AR(1) coefficient $\rho_\omega^j$ and standard deviation $\sigma_\eta^j$. In both industries, the estimated AR(1) coefficient for the homogeneous case ($J = 1$) is considerably larger than those for the unobserved heterogeneity case ($J = 6$). This suggests that neglecting unobserved heterogeneity may result in an upward bias in the estimates of the AR(1) coefficient for productivity processes.

In various latent types, the returns to scale for concrete products, denoted by $(\hat{\beta}_m^j + \hat{\beta}_\ell^j + \hat{\beta}_k^j)$, is approximately 0.7, whereas for electrical audio equipment, it ranges between 0.78 and 0.91. When considering the homogeneous case, the returns to scale are comparatively lower at 0.63 and 0.65 for these respective industries. The estimates of $(\hat{\psi}^j)$ indicate the existence of considerable unobserved heterogeneity in labor quality or working hours among manufacturing plants.

table[table omitted — 2,204 chars of source]
table[table omitted — 2,194 chars of source]

Figures (ref) and (ref) display the distribution of posterior type probabilities, defined by $\hat\pi^j_i := \frac{\hat\pi^j L_{i}(\hat\theta^j)}{\sum_{k=1}^J \hat\pi^{k} L_{i}(\hat\theta^{k})}$ for $j=1,...,J$, across plants for the model with $J=6$. The posterior probabilities for each type are concentrated around 0 or 1. In the subsequent analysis, we assign one of the $J$ types to each plant based on its posterior type probability that achieves the highest value across the $J$ types.

figure[figure omitted — 244 chars of source]
figure[figure omitted — 251 chars of source]

Ignoring unobserved heterogeneity may lead to significant biases in measuring productivity growth. To examine this issue, we consider a specification with $J=6$ as the true model and compute the bias in measuring productivity growth when using a misspecified model with $J=1$. Specifically, let $\Delta\omega_{it}:= \Delta y_{it}-(\hat\beta_{t}^j+\hat\beta_{m}^j \Delta m_{it}+ \hat\beta_{\ell}^j \Delta \tilde\ell_{it} + \hat\beta_{k}^j \Delta k_{it} + \Delta \hat\epsilon_{it}^j)$ for $j=1,2,...,6$ be the estimated productivity growth when $J=6$ and let $\Delta \tilde\omega_{it}:= \Delta y_{it}-(\bar\beta_t+ \bar\beta_{m} \Delta m_{it}+ \bar\beta_{\ell}\Delta \tilde\ell_{it} + \bar\beta_{k} \Delta k_{it} + \Delta \bar\epsilon_{it})$ be the estimated productivity growth when $J=1$, where $\{\hat\beta_{t}^j,\hat\beta_{m}^j,\hat\beta_{\ell}^j, \hat\beta_{k}^j\}_{j=1}^6$ and $\{\bar\beta_{t},\bar\beta_{m}^j,\bar\beta_{\ell}^j, \bar\beta_{k}^j\}$ denote estimated coefficients when $J=6$ and $J=1$, respectively. Then, we compute the bias as \[ \Delta\tilde \omega_{it}=\Delta\omega_{it} + \underbrace{ (\bar \beta_{m}-\hat\beta_{m}^j) \Delta m_{it}+ (\bar\beta_{\ell} - \hat\beta_{\ell}^j) \Delta \ell_{it} + (\bar\beta_{k}^j- \hat\beta_{k}^j) \Delta k_{it} +(\Delta \bar\epsilon_{it}- \Delta \hat\epsilon_{it}^j)}_{:=\text{Bias}_{it}}. \]

The first row of Table (ref), labeled as $\frac{\text{Mean of } |\text{Bias}_{it}|}{\text{Mean of }|\Delta\tilde\omega_{it}|}$, presents the ratio of the average absolute value of bias to the average productivity growth within each of three subsamples, which are classified by technology types. The magnitude of the bias is around 0.10 for concrete products, while it ranges from 0.23 to 0.35 for electric audio equipment.

The second row of Table (ref), denoted by $\frac{\text{Mean of } \text{Bias}_{it}}{\text{Mean of } |\Delta\tilde\omega_{it}|}\Big|_{\Delta\omega_{it}>0}$, reports the ratio of the average value of bias to the average productivity growth conditional on positive productivity growth measured by the model with $J=6$. Note that $\text{Bias}_{it}\approx (\hat\beta_{m}^j-\bar \beta_{m}) \Delta m_{it}$ and $Corr(\Delta \omega_{it},\Delta m_{it})>0$. Consequently, the average bias conditional on $\Delta\omega_{it}>0$ tends to be positive when $\hat\beta_m^j > \bar\beta_m^j$. The empirical results confirm this pattern: for both concrete products and electric audio equipment, Types 5-6 have $\hat\beta_m^j$ higher than $\bar\beta_m$, and thus the estimated bias $\frac{\text{Mean of } \text{Bias}_{it}}{\text{Mean of } |\Delta\tilde\omega_{it}|}\Big|_{\Delta\omega_{it}>0}$ is positive for these types, while the estimated bias when $\Delta\omega_{it}>0$ is negative for Types 1-2 that have lower values of $\hat\beta_m^j$.

These findings imply that neglecting unobserved heterogeneity could lead to significant bias in estimating productivity growth, and the bias is likely to exhibit a systematic pattern depending on the values of $\hat\beta_m^j$.

table[table omitted — 857 chars of source]

As an application of using the estimated productivity growth in empirical analysis, we now investigate whether unobserved heterogeneity, as captured by type-specific production function parameters, is significant for investment decisions. Specifically, for each subsample classified by type, we estimate the following linear investment model: \[ \frac{I_{it}}{K_{it}} = \alpha_0 + {\alpha^j_\omega} \hat\omega_{it} + \text{quadratic of $k_{it}$} + \zeta_{it}, \] where ${I_{it}}/{K_{it}}$ represents the ratio of investment to capital stock.

Table (ref) displays the OLS estimates of $\alpha^j_\omega$ in the first row as well as the quantile regression estimates of $\alpha^j_\omega$ at the 10th, 25th, 50th, 75th, and 90th percentiles across different types for $J=1$ and $6$ for concrete products. Table (ref) presents the same estimates for the electric audio equipment industry. When $J=1$, the OLS coefficient of $\omega_{it}$ is estimated significantly at $0.06$ for concrete products and at $0.5$ for electric audio equipment.

For the model with $J=6$, the estimated coefficients of $\omega_{it}$ differ considerably across different types of plants, indicating that the investment response to a productivity shock varies across plants. Both OLS and quantile regression results show that the estimated coefficients tend to be higher for the types with higher $\hat\beta_m^j$ and $\hat\beta_k^j/\hat\beta_\ell^j$, suggesting that firms invest more given a positive productivity shock if their production technology features high material shares and high capital-labor ratios. In the case of quantile regressions, this pattern is particularly pronounced for firms with high investment ratios.

Overall, these results highlight the importance of accounting for unobserved heterogeneity in the production function when estimating plant-level productivity and its impact on investment.

table[table omitted — 2,380 chars of source]
table[table omitted — 2,385 chars of source]

Conclusion

This paper establishes the nonparametric identifiability of production functions when unobserved heterogeneity exists across firms in the form of latent technology groups. Building on our nonparametric identification analysis and considering computational simplicity, we propose an estimation procedure for the production function with random coefficients employing a finite mixture specification.

Our analysis of Japanese plant-level panel data reveals a substantial degree of variation in estimated input elasticities and productivity growth processes across latent types within narrowly defined industries. We demonstrate that neglecting unobserved heterogeneity in input elasticities can lead to significant and systematic bias in estimated productivity growth. Moreover, we highlight the critical role played by unobserved disparities in input elasticities in plant-level investment decisions. Specifically, we find that the correlation between estimated productivity and investment is notably stronger among high capital-intensive latent type firms than among low capital-intensive type firms.

As future research topics, we may extend our framework in several directions. First, the framework may be extended to explicitly account for plant- or firm-specific biased technological change, which may be continuously distributed. This can be achieved by integrating the structural assumption discussed in doraszelski2018measuring, ZHANG2019, Demirer20, and Raval2023 into our framework. The identification approach of this paper---based on panel data with a Markov structure---is different from, but complementary to, the structural approaches of these existing papers. Adopting both identification approaches simultaneously can provide an empirical framework for estimating production functions that incorporates a broader range of unobserved heterogeneity.

Second, although we establish nonparametric identification, we employ a parametric finite mixture model for estimation due to computational complexity. An important research direction would be to relax the parametric assumption and develop a nonparametric estimator for a finite mixture model of the production function. This could be achieved, for instance, by extending the maximum smooth likelihood estimator proposed by lhc11 or the estimator presented by Bonhomme16as.

Finally, the assumption of perfect competition or monopolistic competition with constant price elasticity may not be realistic. If data on quantities and prices are available separately, our framework can be employed to estimate the production function using the quantity of output, instead of relying on sales as a proxy for output. However, it is frequently the case that firm- or plant-level datasets do not contain output quantities and prices separately. Therefore, developing a framework for concurrently identifying the production function and demand structure from revenue data, as demonstrated in the work by kasahara_sugita2020, while incorporating unobserved heterogeneity, is an important future research topic.