Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
101,961 characters · 9 sections · 65 citation commands
Identification and Estimation of Production Function with Unobserved Heterogeneity Hiroyuki Kasahara Vancouver School of Economics University of British Columbia [email removed] Paul Schrimpf Vancouver School of Economics University of British Columbia [email removed] Michio Suzuki Tohoku University Cabinet Office, Government of Japan [email removed]
{ \singlespacing }
Estimating a firm's production function and productivity is a critical topic in empirical economics.\footnote{Understanding how the input is related to the output is a fundamental issue in empirical industrial organization Ackerberg07 while a measure of total factor productivity is necessary to examine the effect of trade policy on productivity and to analyze the role of resource allocation on aggregate productivity Pavcnik00,KASAHARA2008,KASAHARA2013,Hsieh09. Production function estimation is also important for markup estimation hall1988relation, de2012markups,Raval2023.} Despite its importance, the standard production function estimation procedures impose an implausible assumption that production functions are common across firms except for separable Hicks-neutral productivity terms Olley1996, Levinsohn03, Wooldridge09, Ackerberg15, gandhi20. In the presence of identification issues due to the simultaneity problem Marschak44, the literature on identifying production functions that incorporate unobserved heterogeneity beyond the Hicks-neutral productivity term is scarce but a rapidly growing research area Li2017, doraszelski2018measuring, balat19, ZHANG2019, Demirer20, chen21, Raval2023.
This paper establishes the nonparametric identification of production functions from panel data when production functions and productivity growth processes are heterogeneous across firms in unobserved time-varying ways. We consider a finite mixture specification in which there are $J$ distinct time-varying production technologies, and each firm belongs to one of the $J$ latent types. Econometricians do not observe the latent type of firms. Without making any functional form assumption on production technology and productivity growth processes, we establish nonparametric identification of $J$ distinct production functions, productivity processes, and a population proportion of each type under assumptions similar to those used in the existing production function and panel data identification literature. We also address potential measurement errors in labor inputs because of unobserved working hours and labor quality.
Building on our nonparametric identification result and considering computational ease, we propose an estimation procedure for the production function with random coefficients. Under the assumption of Gaussian error terms, we develop a penalized maximum likelihood estimator for a finite mixture model of random coefficient production functions, where the form of the likelihood function is motivated by our identification argument. The EM algorithm is employed to simplify the computational complexity of maximizing the log-likelihood function of the mixture model.
As an empirical application, we investigate the extent of production technology heterogeneity across plants using panel data from Japanese manufacturing plants between 1986 and 2010. We show that the ratios of intermediate cost to sales and intermediate cost to variable cost—averaging from 1986 to 2010 at the plant level—are substantially different across plants within narrowly defined industries. For example, the 90th-10th percentile differences in the intermediate cost shares in variable costs for concrete products and electric audio equipment are large, at 0.27 and 0.67, respectively.\footnote{ZHANG2019 finds considerable heterogeneity in the labor share in sales across firms in the Chinese steel industry.} Differences in input levels cannot explain these heterogeneities; that is, a considerable cross-plant variation in the ratio of intermediate cost to sales or variable costs remains after controlling for observable inputs, presenting evidence for persistent and substantial heterogeneity in production technologies across plants.
By employing a finite mixture of random coefficients production functions, we also find substantial differences in estimated input elasticities across latent types within narrowly defined industries. To understand the consequences of neglecting unobserved heterogeneity in input elasticities on productivity growth measurement, we adopt a specification with unobserved heterogeneity as the true model and calculate the bias in productivity growth measurement when using a misspecified production function model that omits unobserved heterogeneity. Our findings indicate that ignoring unobserved heterogeneity in input elasticities can result in substantial and systematic bias in estimated productivity growth, contingent on the heterogeneous parameter estimates and the direction of productivity changes. Additionally, our analysis reveals a significantly stronger correlation between estimated productivity and investment among high capital-intensive latent type firms compared to low capital-intensive type firms, implying that unobserved disparities in input elasticities are vital in plant-level investment decisions.
As first discussed by Marschak44, ordinary least squares estimation of production functions is subject to simultaneity bias when firms make input decisions based on their productivity level Griliches98. To address the simultaneity issue, Olley1996 and Levinsohn03 develop control function approaches, which have been widely applied in empirical studies Wooldridge09, Ackerberg15. Despite their popularity, the control function approach has faced potential identification issues as highlighted in the literature. bond05 and Ackerberg15 discuss identification issues due to collinearity under two flexible inputs (i.e., material and labor) in Cobb-Douglas specification. Furthermore, gandhi20 contend that if the firm's decision follows a Markovian strategy, the moment restriction utilized in the control function approach fails to provide sufficient restriction to identify flexible input elasticities due to a lack of instrumental power.
GNR exploit the first-order condition for flexible input under profit maximization and establish the identification of production functions without making any functional form assumptions. However, their result presumes that production technology is identical across plants, except for the Hicks-neutral productivity term. This paper extends the nonparametric identification approach of GNR to accommodate settings where production technologies exhibit unobserved heterogeneity across plants.
Several papers employ the first-order condition as a restriction to identify heterogenous elasticities of flexible inputs under functional form assumptions VanBiesebroeck03, Li2017, doraszelski2018measuring, balat19, ZHANG2019.\footnote{As Solow57 first illustrates, the flexible input elasticities are identified with their input revenue share under the Cobb-Douglas production functions.} doraszelski2018measuring develop a framework to identify plant-level time-varying labor-augmenting productivity in addition to Hicks-neutral productivity under the constant elasticity of substitution (CES) production function, allowing for two-dimensional heterogeneity. ZHANG2019 proposes an estimation method based on the CES production function that accounts for heterogeneity in capital, labor, and material-augmenting efficiency across firms. Li2017 use the flexible input cost ratio to construct a control variable for latent technology to identify flexible inputs' elasticities while imposing the timing assumption suggested by Ackerberg15 to identify the labor and capital coefficients under the Cobb-Douglas specification. balat19 also consider the Cobb-Douglas production function with heterogeneity in the efficiency of using skilled and unskilled labor. Demirer20 extends the framework of doraszelski2018measuring by relaxing the parametric assumption of the CES production function but assumes that the labor-augmenting technology is the only additional source of individual-level heterogeneity other than the Hicks neutral productivity. Raval2023 demonstrates the importance of accommodating non-neutral productivity differences across firms when estimating markups using flexible inputs. dewitte2022 illustrate the significance of accounting for unobserved heterogeneity in productivity growth processes when analyzing export premia and the contributions of exporting firms to aggregate productivity.
These papers identify firm-specific input elasticities, factor-augmenting technologies, or productivity growth processes but impose parametric assumptions or limit the sources of heterogeneity. Our paper complements these studies by establishing nonparametric identification of heterogenous production functions and productivity growth processes without imposing any functional form assumptions or limiting sources of unobserved heterogeneity.
Cheng21 extend the k-means clustering approach of Bonhomme15ecma to multi-dimensional clustering in random coefficient production functions in a nonlinear GMM framework, building upon the dynamic panel approach Arellanobond91, Blundell1998, Blundell2000. Cheng21 consider an asymptotic setting when the time dimension $T\rightarrow \infty$ while our identification is based on $T$ being fixed. Our identification result with fixed $T$ is useful in empirical applications where firm-level panel datasets have limited time dimensions.
Our paper also contributes to the literature on identifying dynamic panel data models with unobserved heterogeneity by relaxing the existing identification conditions. Specifically, Proposition 3 demonstrates that the mixing probabilities and the type-specific time-varying probability distributions across latent types can be identified from panel data with four periods under the Markov assumption and other regularity conditions. This result improves upon the findings of Kasahara09, who established the identification of dynamic panel data models under the Markov assumption but imposed the stationarity and required panel data with six periods.\footnote{HigginsJochmans21 point out that the type-specific distribution is identified only up to an arbitrary ordering of the latent types that differs across different points in Kasahara09. Our argument for identifying a common order of the latent types is based on that of HigginsJochmans21. } Hu12 consider a non-stationary case and establish the identification of a continuous mixture dynamic panel data model using a panel dataset with five time periods. However, their result is limited, as they only establish the identification of type-specific distributions for the third to fifth periods, leaving the identification for the first two periods unresolved. In contrast, we identify the type-specific distributions across all four periods from panel data of length four.
A key condition for our identification analysis is that the observed variable must follow a first-order Markov process within the subpopulation specified by latent type. Proposition (ref) demonstrates that this Markov assumption is satisfied under our structural model assumptions, including the Markovian investment strategy (Assumption (ref)(b)) and the monotonicity of flexible input demands for productivity and wage shocks (Assumption (ref)(b)), both of which are standard assumptions in the production function literature Olley1996,Levinsohn03. Another identifying condition is a rank condition in Assumption (ref) which requires that the changes in the value of the observed vector, $\boldsymbol z_{t}$, must induce sufficiently different changes in the value of the type specific conditional density function of $\boldsymbol z_t$ given the past value $\boldsymbol z_{t-1}$ across latent types. As illustrated in the Cobb-Douglas example in Appendix (ref), this condition is satisfied when input elasticities are sufficiently different across latent types.
The remainder of this paper is organized as follows. Section 2 presents evidence of heterogeneity in production technologies across plants, using panel data from Japanese manufacturing plants. Section 3 introduces the setup for our production function models and discusses the assumptions. Section 4 provides the main identification results, while Section 5 develops an estimator for the production function using a finite mixture model. In Section 6, we present empirical results based on the Japanese manufacturing plants. Section 7 concludes the paper.
In order to underscore the significance of accounting for unobserved heterogeneity in production functions, we first present a set of stylized facts that clearly indicate the presence of heterogeneity beyond Hicks-neutral technology components in production functions. Our analysis employs panel data derived from Japanese manufacturing facilities, spanning the years 1986 to 2010. A comprehensive discussion of the dataset can be found in Section (ref).
For illustration, consider a plant with the Cobb-Douglas production technology: \[ \log Y_{it} =\beta_0+ \beta^i_{m} \log M_{it}+ \beta^i_{\ell} \log L_{it} + \beta^i_{k} \log K_{it} +\omega_{it},\smallskip \] where $Y_{it}$, $M_{it}$, $L_{it}$, and $K_{it}$ denote the output, intermediate input, labor, and capital of plant $i$ in year $t$, respectively, while $\omega_{it}$ represents the total factor productivity (TFP) that follows a first-order Markov process. The superscript $i$ in $\beta^i_m$, $\beta_\ell^i$, and $\beta^i_k$ signifies the variation in output elasticities of inputs across different plants.
We assume that firms consider their output and input prices as given, and that the intermediate input and labor are flexibly selected after $\omega_{it}$ has been fully observed. Consequently, a plant's profit maximization implies the following relationships:
where $P_{Y,t}$, $P_{M,t}$, and $W_{t}$ represent the prices of output, intermediate input, and labor, respectively.
In most existing empirical studies, production functions are estimated under the assumption that the coefficients $\beta^i_m$, $\beta^i_\ell$, and $\beta^i_k$ do not vary across plants. This assumption can be tested in light of ((ref)) by examining whether the intermediate input share, $ \frac{P_{M,t}M_{it}}{P_{Y,t} Y_{it}}$, and the ratio of intermediate cost to variable cost (i.e., the sum of intermediate and labor costs), $\frac{P_{M,t}M_{it}}{P_{M,t}M_{it}+W_{t}L_{it}}$, remain constant across plants. To investigate this, we calculate the plant-level averages over the maximum 25-year period, during which a plant remained in the market between 1986 and 2010, as follows: \[ \overline{ \left(\frac{P_{M,t}M_{it}}{P_{Y,t} Y_{it}}\right)}_i=\frac{1}{25}\sum_{t=1986}^{2010} \frac{P_{M,t}M_{it}}{P_{Y,t} Y_{it}}\quad\text{and}\quad \overline{\left(\frac{P_{M,t}M_{it}}{P_{M,t}M_{it}+W_{t}L_{it}}\right)}_i=\frac{1}{25}\sum_{t=1986}^{2010}\frac{P_{M,t}M_{it}}{P_{M,t}M_{it}+W_{t}L_{it}}. \] Subsequently, we analyze the extent of variation across plants within a narrowly defined industry.
Figure (ref) and Figure (ref) display histograms illustrating plant-level averages of intermediate input shares, $\overline{ \left(\frac{P_{M,t}M_{it}}{P_{Y,t} Y_{it}}\right)}_i$, for all plants within the concrete products and electric audio equipment industries, respectively. Both figures exhibit substantial variation in intermediate shares. The disparity between the 90th and 10th percentiles reaches up to 0.28 for concrete products, an industry typically regarded as having homogeneous technology.
The variation in intermediate shares between plants might be indicative of differences in markups; however, the ratio of intermediate costs to total variable costs is less susceptible to markup discrepancies. Figures (ref) and (ref) depict histograms of plant-level averages of the ratio of intermediate costs to total variable costs, $\overline{\left(\frac{P_{M,t}M_{it}}{P_{M,t}M_{it}+W_{t}L_{it}}\right)}_i$, for the concrete products and electric audio equipment industries, respectively. The significant variation in intermediate cost shares implies that heterogeneous markups are not the primary explanation for the observed variation in intermediate input shares presented in Figures (ref) and (ref).
By comparing the degree of dispersion in input shares within the 2-digit industry classification with that within the 3-digit or 4-digit industry classification, we can examine the extent to which classifying industries at a more refined level helps control for heterogeneity in production technology.
Figures (ref), (ref), and (ref) contrast histograms of plant-level averages of intermediate input shares for ceramics and clay (2-digit), cement products (3-digit), and concrete products (4-digit) industries. These figures suggest that while dispersion decreases somewhat from the 2-digit to the 4-digit level, the degree of heterogeneity remains notably high even at the 4-digit industry classification. Likewise, Figures (ref), (ref), and (ref) reveal that the dispersion of material shares does not decline considerably when transitioning from electric parts, devices, and circuits (2-digit) to electric devices (3-digit) and subsequently to electric audio equipment (4-digit).
As illustrated in Table (ref), the 90th-10th percentile difference in intermediate input shares decreases from 0.38 (2-digit) to 0.28 (4-digit) for the ceramics and clay industry. In contrast, the 90th-10th percentile difference for the electric parts industry changes only marginally from 2-digit to 4-digit, ranging from 0.61 to 0.62. Similar patterns are observed for the plant-level averages of intermediate cost shares in variable costs. The 90th-10th percentile differences in intermediate cost shares for concrete products (4-digit) and electric audio equipment (4-digit) are substantial, measuring 0.27 and 0.67, respectively. Therefore, classifying industries at a more refined level does not substantially reduce the heterogeneity in output elasticities with respect to intermediate and labor inputs.
Analogous patterns are observed across various industries. Table (ref) displays the average differences between the 90th and 10th percentiles for all industries, classified at the 2-digit, 3-digit, and 4-digit levels, with their corresponding standard deviations presented in parentheses. The findings reveal that dispersion decreases only marginally when refining industry classification from the 2-digit to the 4-digit levels. In general, substantial dispersion persists even at the 4-digit classification level, indicating that output elasticities of variable inputs exhibit variation across plants, even when adopting a more detailed industry classification.
The implications in ((ref)) are valid exclusively under the Cobb-Douglas production function. For a more generalized production function, the elasticities of output for inputs are dependent on the levels of material, labor, and capital inputs, even without heterogeneity in production technology. Consequently, we investigate whether the intermediate input cost-to-output value ratio remains similar across firms, even after adjusting for variations in capital, labor, and intermediate inputs. To achieve this, we regress $\frac{P_{M,t}M_{it}}{P_{Y,t}Y_{it}}$ or $\frac{P_{M,t}M_{it}}{P_{M,t}M_{it}+W_{t}L_{it}}$ on second-order polynomials of the natural logarithm of materials, the number of workers, and capital, resulting in residuals denoted by $e_{it}$. Subsequently, we compute the plant-level average $\hat\xi_i := (1/25)\sum_{t=1986}^{2010} e_{it}$ to assess production technology heterogeneity, conditional on inputs.
Table (ref) presents the averages of the 90th-10th percentile differences in $\hat\xi_i$ for $\frac{P_{M,t}M_{it}}{P_{Y,t}Y_{it}}$ or $\frac{P_{M,t}M_{it}}{P_{M,t}M_{it}+W_{t}L_{it}}$ across all industries, classified at the 2-digit, 3-digit, and 4-digit levels. The findings reveal considerable variation in production technology after accounting for observable inputs, even at the 4-digit industry classification. This provides further evidence supporting the existence of heterogeneity in production technology.
We consolidate the notation as follows. Let $:=$ stand for "equal by definition". Bold letters denote vectors or matrices. For a continuous random variable $X$, a calligraphic letter $\mathcal{X}$ denotes its support, while its probability density function is represented by $g_{X}(x)$ for $x \in \mathcal{X}$.
Output, capital, intermediate inputs, labor input in effective units of labor, and total wage bills are denoted by $(Y_{it},K_{it},M_{it},L_{it},B_{it}) \in \mathcal{Y} \times \mathcal{K} \times \mathcal{M} \times \mathcal{L} \times \mathcal{B} \subset \mathbb{R}_{++}^5$, respectively, where $\mathcal{Y}$, $\mathcal{K}$, $\mathcal{M}$, $\mathcal{L}$, and $\mathcal{B}$ are the supports of the corresponding variables. We assume that $(Y_{it},K_{it},M_{it},L_{it},B_{it})$ are continuously distributed with strictly positive density on connected supports. We combine capital, intermediate, and labor inputs into a vector as ${\boldsymbol X}_{it} := (K_{it},M_{it},L_{it})' \in \mathcal{X} := \mathcal{K} \times \mathcal{M} \times \mathcal{L}$.
We allow firms' production technologies to differ beyond Hick's neutral productivity shocks. Specifically, we use a finite mixture specification to capture the unobserved heterogeneity in firms' production technologies as well as the process through which Hick's neutral productivity shocks evolve. We assume that there are $J$ unobserved types. Define the latent random variable $D_{i} \in\{1,2,...,J\}$ representing the type of firm $i$ such that $D_i=j$ if the production technology of firm $i$ is of the $j$-th type. In the following, the superscript $j$ indicates that the functions are specific to technology type $j$, while the subscript $t$ indicates that the functions are specific to period $t$.
For the $j$-th type of production technology at time $t$, the output is related to inputs as follows:
where $\epsilon_{it}$ is an idiosyncratic productivity shock with its density function $g_{\epsilon_t}^j(\cdot)$, and $\omega_{it}$ follows an exogenous first-order stationary Markov process given by:
where $\eta_{it}$ is an innovation to the productivity process. As indicated by the subscript $t$ in $F_{t}^j(\cdot)$, $h_t^j(\cdot)$, $g_{\epsilon_t}^j(\cdot)$, and $g_{\eta_t}^j(\cdot)$, production functions and productivity processes differ not only between latent types but also across periods. For example, this reflects type-specific aggregate shocks or type-specific biased technological changes.
We assume that labor input in effective units of labor, $L_{it}$, is not directly observable because firms differ in their labor quality and working hours, while we only observe the number of workers. Instead, $L_{it}$ is related to the number of workers, denoted by $\widetilde L_{it}$, as
Equation ((ref)) imposes a specific structure on how labor input in effective units of labor is related to the observed number of workers, where $\psi_t^j$ represents latent worker quality and working hours specific to type $j$. With this specification, measurement errors in observed labor input (i.e., the number of workers) are captured by the latent type-specific value of $\psi_t^j$.
The total wage bills, $B_{it}$, are related to the labor input in effective units as follows:
where $P_{L,t}$ represents the market wage. The random variable $v_{it}$ is a transitory wage shock known to firm $i$ at the time of choosing intermediate and labor inputs, while $\zeta_{it}$ is an idiosyncratic wage shock not included in the information set when firm $i$ selects these two inputs. Thus, $\zeta_{it} \in \boldsymbol{\mathcal{I}}_{it}$, but $v_{it} \notin \boldsymbol{\mathcal{I}}_{it}$.
We assume that firms make flexible choices regarding both $M_{it}$ and $L_{it}$ after observing their serially correlated productivity shock, $\omega_{it}$, but before observing $\epsilon_{it}$. In contrast, $K_{it}$ is predetermined at the end of the previous period, prior to the observation of the serially correlated productivity shock $\omega_{it}$. Denote the information available to a firm for making decisions on $M_{it}$ and $L_{it}$ by $\boldsymbol{\mathcal{I}}_{it}$.
We present the model assumptions. For continuous random variables $Z_{it}$ and $W_{it}$, we denote the probability density function and expectation conditional on $D_i=j$ as $g_{Z_t}^j(z_t) := g_{Z_t|D=j}(z_t|D_i=j)$ and $E^j[Z_{it}] := E[Z_{it}|D_i=j]$, respectively. Additionally, we denote the probability density function of $Z_{it}$ conditional on $D_i=j$ and $W_{it}=w_t$ as $g_{Z_t|W_t}^j(z_t|w_t)$. The unconditional probability density function of $Z_{it}$ is denoted by $g_{Z_t}(z_t)$.
Under Assumptions (ref)-(ref), the information set at the time of choosing $M_{it}$ and $L_{it}$ is given by $\boldsymbol{\mathcal{I}}_{it} =\{D_i,\omega_{it},v_{it}, K_{it}, P_{M,t},P_{Y,t}, P_{L,t}, \boldsymbol V_{it-1},\boldsymbol V_{it-2},...\}$ with $\boldsymbol V_{it}=\{\zeta_{it},\epsilon_{it},\omega_{it},v_{it}, K_{it}, P_{M,t},P_{Y,t},P_{L,t}\}$.
Assumption (ref)(a) presumes that the number of types is known. Kasahara09 and Kasahara2014 discuss nonparametric identification of a lower bound for the number of types. Kasahara2019 and HaoKasahara2022 develop a likelihood-based testing procedure for the number of types in multivariate and panel data normal mixture models. Assumption (ref)(b) assumes that a firm is aware of its type.
Assumption (ref)(a)(b) asserts that $(\omega_{it},v_{it})$ is known when $L_{it}$ and $M_{it}$ are chosen, while $(\epsilon_{it},\zeta_{it})$ is not known when $L_{it}$ and $M_{it}$ are chosen. The presence of the wage shock $v_{it}$ in ((ref)) provides an additional source of variation for $L_{it}$ beyond $\omega_{it}$ and $K_{it}$; as a result, $L_{it}$ and $M_{it}$ are not collinear, avoiding the identification problem discussed by bond05 and Ackerberg15. Assumption (ref)(c) introduces notation for the probability density function of $\eta_{it}$, $v_{it}$, $\epsilon_{it}$, and $\zeta_{it}$, allowing for correlation between $\epsilon_{it}$ and $\zeta_{it}$. Assumption (ref)(d) is a normalization assumption to identify the location of $F_{t}^j$.
Assumption (ref)(a) presumes that $K_{it}$ is determined at time $t-1$, implying that $(\eta_{it},\omega_{it},v_{it})$ is not known when $K_{it}$ is chosen. Assumption (ref)(b) can be explicitly derived from a dynamic model of investment decisions with convex/non-convex adjustment costs under the first-order Markov productivity process ((ref)).
Assumption (ref)(a) introduces the demand function for $M_{it}$ and $L_{it}$, derived from the static profit maximization problem when $M_{it}$ and $L_{it}$ are flexibly chosen in each period. Assumption (ref)(b) holds when there is a one-to-one relationship between $(M_{it},L_{it})$ and $(\omega_{it},v_{it})$, except for a set of measure zero conditional on the value of $K_{it}$, and is satisfied in the case of a Cobb-Douglas function. Under Assumption (ref)(c), the first-order condition for the maximization problem ((ref)) characterizes the optimal choice for $L_{it}$ and $M_{it}$.
Assumption (ref)(a) posits that the firm has no market power. Under Assumption (ref)(b), the intermediate input price $P_{M,t}$ cannot be used for instrumenting $M_{it}$. When intermediate prices are exogenous and heterogeneous across firms, the production function could be identified using the intermediate input prices as instruments doraszelski2018measuring. In Assumption (ref)(c), an alternative approach assumes that a firm is subject to an idiosyncratic price shock $\xi_{it}$ such that, for example, $P_{Y,it}=\exp(\xi_{it})P_{Y,t}$ with $\xi_{it}\not\in \boldsymbol{\mathcal{I}}_{it}$, then $\xi_{it}$ plays a similar role to $\epsilon_{it}$. We may assume that $(P_{M,t},P_{Y,t})$ is not known to the econometrician by treating $P_{M,t}/P_{Y,t}$ as parameters to be estimated; in such a case, we can identify the production function up to scale.
Appendix (ref) discusses an alternative assumption to Assumption (ref) when a firm produces differentiated products and faces a demand function with constant price elasticity.
Assumption (ref) suggests that the quality of workers or the average working hours per worker differ across types and periods, as captured by the parameter $\psi_t^j$, which leads to the systematic difference in the average wage of workers latent types. The assumption that $\sum_{j=1}^J \pi^j e^{\psi_t^j}=1$ serves as a normalization for identification.
Assume that we have panel data for firms $i=1,...,n$ over periods $t=1,...,T$ consisting of output, capital, intermediate inputs, the number of workers, and total wage bills, denoted by $(Y_{it},B_{it},K_{it},M_{it},\widetilde L_{it})\in \mathcal{Y}\times\mathcal{B}\times\mathcal{K}\times \mathcal{M}\times\tilde{\mathcal{L}}$, respectively. For brevity, define $\boldsymbol{\boldsymbol{\widetilde X}}:= (K_{it},M_{it},\widetilde L_{it}) \in \widetilde{ \mathcal{X}}:=\mathcal{K}\times \mathcal{M}\times\tilde{\mathcal{L}}$. Each firm's observation $\{Y_{it},B_{it},\widetilde {\boldsymbol X}_{it}\}_{t=1}^T$ is randomly sampled from a population distribution with a density function given by $g_{\{Y_{t},B_{t},\widetilde {\boldsymbol X}_{t}\}_{t=1}^T}(\{Y_{t},B_{t},\widetilde {\boldsymbol X}_{t}\}_{t=1}^T)$.
Let $g_{\epsilon_t}(\epsilon):= \int_{\mathbb{R}}g_{\epsilon_t,\zeta_t}(\epsilon,\zeta)d\zeta$ and $g_{\zeta_t}(\zeta):= \int_{\mathbb{R}} g_{\epsilon_t,\zeta_t}(\epsilon,\zeta)d\epsilon$ be the probability density functions of $\epsilon$ and $\zeta$, respectively. Under Assumptions (ref), (ref), (ref)(a), (ref)(a), and (ref), the first order conditions with respect to $M_{it}$ and $L_{it}$ for maximizing the expected static profit in Assumption (ref)(a) are given by:
where $F_{M,t}^j(\boldsymbol X):=\frac{\partial F_{t}^j(X) }{\partial M}$, $F_{L,t}^j(X):=\frac{\partial F_{t}^j(\boldsymbol X) }{\partial L}$, $E_t^j[e^{\epsilon}] :=\int e^{\epsilon} dg^j_{\epsilon_t}(\epsilon)d\epsilon$, and $E_t^j[e^{\zeta}] :=\int e^{\zeta} dg^j_{\zeta_t}(\zeta)d\zeta$. Rearranging equations ((ref)), ((ref)), ((ref)), and ((ref)) gives a system of equations:
where
For notational brevity, we drop the subscript $i$ in the rest of this section. Because $Y_t=\frac{P_{M,t}M_t}{S^m_t P_{Y,t}}$ and $B_t=\frac{S^\ell_t P_{M,t}M_t}{S^m_t}$, there exists a one-to-one relationship between $(Y_t,B_t)$ and $(S_t^m,S_t^\ell)$ given $\widetilde {\boldsymbol X}_t$ under Assumption (ref). Therefore, denoting $\boldsymbol S_t=(S_t^m,S_t^\ell)\in \mathcal{S}$, we consider $\{\boldsymbol S_t,\widetilde {\boldsymbol X}_t\}_{t=1}^T$ in place of $\{Y_{t},B_{t},\widetilde {\boldsymbol X}_{t}\}_{t=1}^T$ as our data.
Let $\boldsymbol Z_t := (\boldsymbol S_t,\widetilde {\boldsymbol X}_t)\in \mathcal{Z}:= \mathcal{S}\times\widetilde{\mathcal{X}}$. We assume that the population density function, denoted by $g_{\boldsymbol Z_1,...,\boldsymbol Z_T}(\{\boldsymbol z_{t}\}_{t=1}^T)$, is directly identified from the data. We are interested in identifying the model structure \[ \boldsymbol{\theta} := \left\{ \pi^j, \{g_{v_t}^j(\cdot), g_{\epsilon_t,\zeta_t}^j(\cdot), ,\Gamma_{M,t}^j(\cdot),\Gamma_{L,t}^j(\cdot),P_{L,t},\psi_t^j\}_{t=1}^T,\{F_t^j(\cdot)\}_{t=2}^T,\{h_t^j(\cdot),g_{\eta_t}^j(\cdot)\}_{t=3}^T\right\}_{j=1}^J \] from the population density function $\text{g}_{\boldsymbol Z_1,...,\boldsymbol Z_T}(\{\boldsymbol z_{t}\}_{t=1}^T)$ given a set of restrictions in ((ref)) under Assumptions (ref)-(ref).
We first establish the nonparametric identification of model structure $\boldsymbol\theta$ when $J=1$ as follows.
When $J\geq 2$, the probability density function of ${\{\boldsymbol Z_t\}_{t=1}^T}$ follows an $J$-term mixture distribution
The number of type $J$ is defined to be the smallest integer $J$ such that the density function of ${\{\boldsymbol Z_t\}_{t=1}^T}$ admits the representation ((ref)).
Therefore, under the stated model assumption, ${\{\boldsymbol Z_t\}_{t=1}^T}$ follows a first order Markov process within subpopulation specified by type. The result of Proposition (ref) allows us to establish the nonparametric identification of $\{\pi^{j},g_{\boldsymbol Z_1}^{j}(\boldsymbol z_1), g_{\boldsymbol Z_2|\boldsymbol Z_1}^{j}(\boldsymbol z_2|\boldsymbol z_1), ... , g_{\boldsymbol Z_T|\boldsymbol Z_{T-1},...,\boldsymbol Z_1}^{j}(\boldsymbol z_T|\boldsymbol z_{T-1}) \}_{j=1}^J$ by extending the argument in Kasahara09, carolletal10jns, and Hu12.
We now establish identification when $T=4$. Define
where $\bar \lambda^j_2(\boldsymbol a,{\boldsymbol z}_2):=\pi^j g_{{\boldsymbol Z}_2|{\boldsymbol Z}_1}^j({\boldsymbol z}_2|\boldsymbol a)g_{{\boldsymbol Z}_1}^j(\boldsymbol a)$, $\lambda^j_3({\boldsymbol z}_3|{\boldsymbol z}_2):= g_{{\boldsymbol Z}_3|{\boldsymbol Z}_2}^j({\boldsymbol z}_3|{\boldsymbol z}_2)$, and $\lambda^j_{4}(\boldsymbol b|{\boldsymbol z}_3):=g_{{\boldsymbol Z}_4|{\boldsymbol Z}_3}^j(\boldsymbol b|{\boldsymbol z}_3)$.
Once the type-specific distribution of $\{{\boldsymbol Z}_t\}_{t=1}^T$ is identified, we can use the argument in the proof of Proposition (ref) to prove nonparametric identification for the model structure of each type.
Therefore, type-specific production functions, as well as the distribution of unobserved variables, can be identified nonparametrically. In the estimation, we focus on the case where the type-specific production functions are Cobb-Douglas.
In Appendix (ref), we discuss the sufficient conditions under which Assumption (ref) holds when the production function is Cobb-Douglas as in Example (ref).
In this section, we present a finite mixture model of the random coefficient Cobb-Douglas production function, based on our nonparametric identification analysis and with consideration for computational efficiency. We develop a penalized maximum likelihood estimator for this model.
Let us denote the logarithmic values of $(Y_{it},K_{it},M_{it},\tilde L_{it},S_{it}^m,S_{it}^\ell,B_{it})$ using the corresponding lowercase letters, such that $(y_{it},k_{it},m_{it},\tilde \ell_{it},s_{it}^m,s_{it}^\ell,b_{it})$, where $y_{it} :=\log Y_{it}$, and so on. Define $\boldsymbol s_{it}:=(s_{it}^m,s_{it}^\ell)$ and $\tilde {\boldsymbol x}_{it}:=(k_{it},m_{it},\tilde \ell_{it})$. For estimation purposes, we assume that the data generation follows the parametric assumptions outlined below.
Assumption (ref)(a) posits that the length of panel data is short, while Assumption (ref)(b) enforces the Cobb-Douglas functional form assumption. Assumptions (ref)(c) and (ref)(d) impose Gaussian distribution assumptions under Assumptions (ref) and (ref), where $\omega_{it}$ follows a first-order autoregressive (AR(1)) process.
In equation ((ref)), since $\ln L_{it} = \psi_t^j +\tilde \ell_{it}$, the intercept term encompasses both $\beta_{0,t}^j$ and $\beta_{\ell}^j\psi_t^j$. The latter term captures the variation in worker quality across types. The normality assumption in Assumptions (ref)(c) and (ref)(d) could potentially be relaxed; for instance, by employing the maximum smoothed likelihood estimator of finite mixture models proposed by lhc11, in which the type-specific distribution of $\epsilon_{it}$ and $\zeta_{it}$ is nonparametrically specified. Additionally, kasaharashimotsu15jasa develop a likelihood-based procedure to test the number of components in normal mixture regression models.
Suppose we have a random sample of $n$ independent observations $\{\{\boldsymbol S_{it},\widetilde {\boldsymbol X}_{it}\}_{t=1}^T\}_{i=1}^n$ from the $J$-component mixture model $\sum_{j=1}^J \pi^j g_{\{\boldsymbol S_t,\widetilde {\boldsymbol X}_t\}_{t=1}^T}^j(\{\boldsymbol s_{t},\widetilde {\boldsymbol x}_{t}\}_{t=1}^T)$ that satisfies Assumptions (ref)-(ref). We propose a penalized maximum likelihood estimator (PMLE) that directly maximizes the log-likelihood function of a finite mixture model of production functions. The likelihood function is a parametric version of ((ref)). To address the issue of unbounded likelihood for a normal mixture model hartigan85book, we introduce a penalty term to the log-likelihood function. The maximum likelihood estimator, which leverages distributional information, is consistent even when $T$ is small, provided $T \geq 4$. Given the nonparametric identification result established in Proposition (ref), if the parametric assumptions are invalid and the parametric model is misspecified, the parametric maximum likelihood estimator converges in probability to the pseudo-true value of the parameter that minimizes the Kullback-Leibler Information Criterion between the density of the parametric model and the true population density White1982.
Our estimation procedure is based on the two-stage identification proof from Proposition (ref). To address the computational complexity of maximizing the log-likelihood function for the finite mixture model, the EM algorithm is employed.
Under Assumptions (ref)-(ref), (ref), the first order conditions for the expected profit maximization imply that
where $\alpha_t:= \ln (P_{L,t}/P_{M,t})$ and ((ref)) follows from ((ref)), $v_{it}=b_{it}-(\psi^j +\tilde \ell_t + \ln P_{L,t}+ \zeta_{it})$, and $s_{it}^\ell-s_{it}^m= b_{it} - (\ln P_{M,t}+ m_{it})$.
Collect the model parameter $\boldsymbol\theta$ into $\boldsymbol\pi$, $\boldsymbol{\theta}_1$ and $\boldsymbol{\theta}_2$ as \[ \boldsymbol\theta:=(\boldsymbol\pi',\boldsymbol\theta_1',\boldsymbol\theta_2')'\quad\text{with}\quad\boldsymbol{\theta}_1= (\boldsymbol{\alpha}',(\boldsymbol \theta_1^1)',...,(\boldsymbol{\theta}_1^J)')'\ \text{and}\ \boldsymbol{\theta}_2 = ((\boldsymbol{\theta}_2^1)',...,(\boldsymbol{\theta}_2^J)')', \] where $\boldsymbol\theta_{1}^j =(\beta_m^j,\beta_\ell^j, \psi^j, (\sigma_{\epsilon}^j)^2,(\sigma_{\zeta}^j)^2,(\sigma_v^j)^2)'$ and $\boldsymbol\theta_2^j=(\beta_{2}^j,...,\beta_T^j, \beta_k^j,(\boldsymbol\mu_1^j)',\text{vech}(\boldsymbol\Sigma_1^j)',\rho_{k0}^j,\rho_{kk}^j,$\\ $\rho_{k\omega}^j,(\sigma_k^j)^2,\rho_{\omega}^j,(\sigma_\eta^j)^2)'$ for $j=1,...,J$, and $\boldsymbol{\alpha}=(\alpha_1,....,\alpha_T)'$.
Denote $\boldsymbol\theta^j:=((\boldsymbol\theta_1^j)',(\boldsymbol\theta_2^j)')'$. Then, under Assumption (ref), we may write the probability density function of $\{\boldsymbol s_{it},\tilde{\boldsymbol x}_{it}\}_{t=1}^T$ for type $j$ as
where the exact expression for $L_{1i}(\boldsymbol\theta_1^j,\boldsymbol{\alpha})$ and $L_{2i}(\boldsymbol \theta_2^j,\boldsymbol \theta_1^j)$ is derived below.
To deal with the issue of unbounded log-likelihood function of a normal mixture model hartigan85book,HaoKasahara2022, we estimate the model parameter $\boldsymbol\theta$ by a penalized likelihood method proposed by chentan09jmva. Let $(\hat\sigma_{\epsilon,0}^2,\hat\sigma_{\zeta,0}^2,\hat\sigma_{v,0}^2, \hat\sigma_{k,0}^2,\hat\sigma_{\eta,0}^2, \hat{\boldsymbol\Sigma}_{1,0})$ be the estimator of $(\sigma_{\epsilon}^2,\sigma_{\zeta}^2,\sigma_{v}^2, \hat\sigma_{k}^2,\hat\sigma_{\eta}^2, \hat{\boldsymbol\Sigma}_{1})$ for the one-component model with $J=1$. Then, we consider the following penalized maximum likelihood estimator (PMLE):
where \[ Q_i(\boldsymbol\theta):= \log\left(\sum_{j=1}^J \pi^j L_{i}(\boldsymbol \theta^j,\boldsymbol\alpha)\right) \quad\text{with}\quad L_{i}(\boldsymbol \theta^j,\boldsymbol\alpha):=L_{1i}({\boldsymbol \theta}_1^j,{\boldsymbol{\alpha}})L_{2i}({\boldsymbol\theta}_1^j,{\boldsymbol\theta}_2^j) \] and
The expression for $L_{1i}({\boldsymbol \theta}_1^j,{\boldsymbol{\alpha}})$ and $L_{2i}({\boldsymbol\theta}_1^j,{\boldsymbol\theta}_2^j)$ are derived below.
To reduce the computational burden of finding the penalized maximum likelihood estimator, we follow a three-stage procedure.
In the first stage, from equations ((ref))-((ref)), we can express $\epsilon_{it}$, $\zeta_{it}$, and $v_{it}$ as a function of $\boldsymbol s_{it}$, $\tilde \ell_{it}-m_{it}$, $\boldsymbol\theta_1^j$, and $\alpha_t$ as
Then, we estimate $\boldsymbol\theta_1$ by maximizing the log likelihood function as
where, under Assumption (ref)(b), the likelihood function $L_{1i}(\boldsymbol\theta_1^j,\boldsymbol{\alpha})$ is given by
with $ \epsilon^*(\boldsymbol s_{it};\boldsymbol \theta_1^j)$, $\zeta^*(\boldsymbol s_{it};\boldsymbol \theta_1^j) $, and $v^*(\tilde \ell_{it}-m_{it} ;\boldsymbol\theta_1^j,\alpha_t)$ defined in ((ref))-((ref)).
In the second stage, from ((ref)), $\epsilon_{it} = E[s_{it}^m| \tilde{\boldsymbol x}_{it}]-s_{it}^{m}$, and $y_{it}+s_{it}^m=m_{it}+\ln (P_{M,t}/P_{Y,t})$, we have
where $\beta_t^j:=\beta_{0,t}^j +\ln (P_{M,t}/P_{Y,t})-\ln \beta_m^j - 0.5(\sigma_{\epsilon}^j)^2$.
In view of equation ((ref)), by a change of variables, we may relate the density function of $m_{it}$ conditional on $\tilde \ell_{it}-m_{it}$ and $k_{it}$ to the density function of $\omega_{it}$, denoted by $g_{\omega,t}$, as $g_t^j(m_{it}|\ell_{it}-m_{it},k_{it})=(1-\beta^j_m-\beta^j_\ell)g_{\omega,t}^j(\omega_t^*(m_{it},\tilde \ell_{it}-m_{it},k_{it};\boldsymbol\theta^j))$. Then, from ((ref))-((ref)) and Assumptions (ref)-(ref), we have
where $g_{\omega|k,1}^j(\omega_{i1} |k_{i1})$ is the density function of $\omega_{i1}$ conditional on $k_{i1}$, $g_{k,t}^j(k_{it}|k_{it-1},\omega_{it-1})$ is the density function of $k_{it}$ given $(k_{it-1},\omega_{it-1})$, $\omega_{it}^*(\boldsymbol\theta^j):=\omega_t^*(m_{it},\ell_{it}-m_{it},k_{it};\boldsymbol\theta^j) $, and
Therefore, under Assumption (ref), it follows from ((ref)) and ((ref))-((ref)) that
where
Given the first stage estimate $\tilde{\boldsymbol\theta}_1$, the second stage estimates parameters $\boldsymbol{\pi}$ and $\boldsymbol\theta_2$ by maximizing the log-likelihood function as \[ (\tilde{\boldsymbol{\pi}}_2,\tilde{\boldsymbol\theta}_2) = \operatorname*{\arg\!\max}_{\boldsymbol{\pi},\boldsymbol\theta_2} \sum_{i=1}^n \log\left(\sum_{j=1}^J \pi^j L_{1i}(\tilde{\boldsymbol\theta}_1^j,\tilde{\boldsymbol{\alpha}}) L_{2i}(\tilde{\boldsymbol\theta}_1^j,\boldsymbol\theta_2^j) \right)+ \sum_{j=1}^J \left\{ \sum_{s\in\{k,\eta\}} p_n((\sigma_s^j)^2;\hat\sigma_{s,0}^2) + p_n(\boldsymbol\Sigma_1^j;\hat{\boldsymbol\Sigma}_{1,0})\right\}. \]
Finally, using $\tilde{\boldsymbol{\theta}} := (\tilde{\boldsymbol{\pi}}_2', \tilde{\boldsymbol{\theta}}_1', \tilde{\boldsymbol{\theta}}_2')'$ as an initial value, we obtain the PMLE $\hat{\boldsymbol{\theta}}$ by maximizing the full-information penalized log-likelihood function as in Equation ((ref)) using the EM algorithm.
Let the true value of the model parameter be denoted by $\boldsymbol{\theta}^*$. When the number of types $J$ is correctly specified, the Fisher Information matrix is given by \[ \boldsymbol{I}(\boldsymbol{\theta}^*) := -\mathbb{E}\left[ \nabla_{\boldsymbol{\theta}\boldsymbol{\theta}'} Q_i(\boldsymbol{\theta}^*) \right] = \mathbb{E}\left[ \nabla_{\boldsymbol{\theta}}Q_i(\boldsymbol{\theta}^*) \nabla_{\boldsymbol{\theta}'} Q_i(\boldsymbol{\theta}^*) \right], \] which is positive definite. The following proposition demonstrates that the PMLE is consistent and asymptotically normal.
We utilize plant-level panel data from the Census of Manufacture of Japan spanning 1986-2010. This dataset encompasses production information for manufacturing plants in Japan. Our analysis focuses on plants with 30 or more employees, as detailed data are consistently available only for these establishments.\footnote{The survey employs distinct questionnaires based on plant size: 1. Plants with 30 or more employees; 2. Plants with 4-29 employees; 3. Plants with 1-3 employees. The questionnaire for plants with 30 or more employees provides more comprehensive information. For instance, beginning in 2000, the census collects fixed asset data every five years, rather than annually, for plants with fewer than 30 employees.}
At the 4-digit industry classification level accessible in the Census of Manufacture, we identify a total of 276 industries. In our empirical application, we primarily concentrate on concrete products and electric audio equipment for two reasons: 1. Both industries have a substantial number of observations; 2. The former exhibits relatively small variation in intermediate input share, while the latter displays significant variation, as demonstrated in Section (ref). Thus, examining these two industries proves useful for assessing the significance of unobserved heterogeneity.
Output ($Y$) is defined as the sum of shipments, revenue from repair and maintenance services, and revenue from performing subcontracted work. Initial capital value ($K$) is determined as the fixed asset value minus land, and subsequent capital values are constructed using the perpetual inventory method. The observed labor input ($\tilde{L}$) is represented by the number of employees. The intermediate input ($M$) is defined as the sum of material input, energy input, and subcontracting expenses for consigned production.
Flow data, such as shipments and various production costs, pertain to the calendar year. The number of employees refers to the value at the end of the year, while the stock of fixed assets corresponds to the beginning of the period. Table (ref) presents summary statistics for the variables employed in our empirical analysis.
This section presents estimation results for a random-coefficient Cobb-Douglas production function featuring three technology types and two unobserved labor types within each technology type.
Table (ref) and Table (ref) present the parameter estimates for the concrete products and electric audio equipment industries, considering both the unobserved heterogeneity case ($J = 3 \times 2 = 6$) and the homogeneous case ($J = 1$). The estimated coefficients in both industries demonstrate economically significant differences in output elasticities associated with labor, capital, and intermediate inputs across various firm types.
Comparing the two industries, the variation in $\hat \beta_m^j$ across types is more substantial for electric audio equipment than for concrete products, which aligns with the dispersion of intermediate input shares discussed in Section (ref). As $\hat \beta_\ell^j$ and $\hat\beta_k^j$ also exhibit variation across types, the ratio of output elasticities between capital and labor, $\hat\beta_k^j/\hat\beta_\ell^j$, varies as well. For electric audio equipment, the value of $\hat\beta_k^j/\hat\beta_\ell^j$ ranges from 0.29 (Type 1 and 2) to 1.14 (Type 3 and 4), while for concrete products, it spans from 0.56 to 0.78. As demonstrated below, the value of $\hat\beta_k^j/\hat\beta_\ell^j$ serves as a crucial determinant of the capital investment response to productivity $\hat\omega_{it}$.
Furthermore, the productivity growth processes display variations across latent types, as indicated by the estimated AR(1) coefficient $\rho_\omega^j$ and standard deviation $\sigma_\eta^j$. In both industries, the estimated AR(1) coefficient for the homogeneous case ($J = 1$) is considerably larger than those for the unobserved heterogeneity case ($J = 6$). This suggests that neglecting unobserved heterogeneity may result in an upward bias in the estimates of the AR(1) coefficient for productivity processes.
In various latent types, the returns to scale for concrete products, denoted by $(\hat{\beta}_m^j + \hat{\beta}_\ell^j + \hat{\beta}_k^j)$, is approximately 0.7, whereas for electrical audio equipment, it ranges between 0.78 and 0.91. When considering the homogeneous case, the returns to scale are comparatively lower at 0.63 and 0.65 for these respective industries. The estimates of $(\hat{\psi}^j)$ indicate the existence of considerable unobserved heterogeneity in labor quality or working hours among manufacturing plants.
Figures (ref) and (ref) display the distribution of posterior type probabilities, defined by $\hat\pi^j_i := \frac{\hat\pi^j L_{i}(\hat\theta^j)}{\sum_{k=1}^J \hat\pi^{k} L_{i}(\hat\theta^{k})}$ for $j=1,...,J$, across plants for the model with $J=6$. The posterior probabilities for each type are concentrated around 0 or 1. In the subsequent analysis, we assign one of the $J$ types to each plant based on its posterior type probability that achieves the highest value across the $J$ types.
Ignoring unobserved heterogeneity may lead to significant biases in measuring productivity growth. To examine this issue, we consider a specification with $J=6$ as the true model and compute the bias in measuring productivity growth when using a misspecified model with $J=1$. Specifically, let $\Delta\omega_{it}:= \Delta y_{it}-(\hat\beta_{t}^j+\hat\beta_{m}^j \Delta m_{it}+ \hat\beta_{\ell}^j \Delta \tilde\ell_{it} + \hat\beta_{k}^j \Delta k_{it} + \Delta \hat\epsilon_{it}^j)$ for $j=1,2,...,6$ be the estimated productivity growth when $J=6$ and let $\Delta \tilde\omega_{it}:= \Delta y_{it}-(\bar\beta_t+ \bar\beta_{m} \Delta m_{it}+ \bar\beta_{\ell}\Delta \tilde\ell_{it} + \bar\beta_{k} \Delta k_{it} + \Delta \bar\epsilon_{it})$ be the estimated productivity growth when $J=1$, where $\{\hat\beta_{t}^j,\hat\beta_{m}^j,\hat\beta_{\ell}^j, \hat\beta_{k}^j\}_{j=1}^6$ and $\{\bar\beta_{t},\bar\beta_{m}^j,\bar\beta_{\ell}^j, \bar\beta_{k}^j\}$ denote estimated coefficients when $J=6$ and $J=1$, respectively. Then, we compute the bias as \[ \Delta\tilde \omega_{it}=\Delta\omega_{it} + \underbrace{ (\bar \beta_{m}-\hat\beta_{m}^j) \Delta m_{it}+ (\bar\beta_{\ell} - \hat\beta_{\ell}^j) \Delta \ell_{it} + (\bar\beta_{k}^j- \hat\beta_{k}^j) \Delta k_{it} +(\Delta \bar\epsilon_{it}- \Delta \hat\epsilon_{it}^j)}_{:=\text{Bias}_{it}}. \]
The first row of Table (ref), labeled as $\frac{\text{Mean of } |\text{Bias}_{it}|}{\text{Mean of }|\Delta\tilde\omega_{it}|}$, presents the ratio of the average absolute value of bias to the average productivity growth within each of three subsamples, which are classified by technology types. The magnitude of the bias is around 0.10 for concrete products, while it ranges from 0.23 to 0.35 for electric audio equipment.
The second row of Table (ref), denoted by $\frac{\text{Mean of } \text{Bias}_{it}}{\text{Mean of } |\Delta\tilde\omega_{it}|}\Big|_{\Delta\omega_{it}>0}$, reports the ratio of the average value of bias to the average productivity growth conditional on positive productivity growth measured by the model with $J=6$. Note that $\text{Bias}_{it}\approx (\hat\beta_{m}^j-\bar \beta_{m}) \Delta m_{it}$ and $Corr(\Delta \omega_{it},\Delta m_{it})>0$. Consequently, the average bias conditional on $\Delta\omega_{it}>0$ tends to be positive when $\hat\beta_m^j > \bar\beta_m^j$. The empirical results confirm this pattern: for both concrete products and electric audio equipment, Types 5-6 have $\hat\beta_m^j$ higher than $\bar\beta_m$, and thus the estimated bias $\frac{\text{Mean of } \text{Bias}_{it}}{\text{Mean of } |\Delta\tilde\omega_{it}|}\Big|_{\Delta\omega_{it}>0}$ is positive for these types, while the estimated bias when $\Delta\omega_{it}>0$ is negative for Types 1-2 that have lower values of $\hat\beta_m^j$.
These findings imply that neglecting unobserved heterogeneity could lead to significant bias in estimating productivity growth, and the bias is likely to exhibit a systematic pattern depending on the values of $\hat\beta_m^j$.
As an application of using the estimated productivity growth in empirical analysis, we now investigate whether unobserved heterogeneity, as captured by type-specific production function parameters, is significant for investment decisions. Specifically, for each subsample classified by type, we estimate the following linear investment model: \[ \frac{I_{it}}{K_{it}} = \alpha_0 + {\alpha^j_\omega} \hat\omega_{it} + \text{quadratic of $k_{it}$} + \zeta_{it}, \] where ${I_{it}}/{K_{it}}$ represents the ratio of investment to capital stock.
Table (ref) displays the OLS estimates of $\alpha^j_\omega$ in the first row as well as the quantile regression estimates of $\alpha^j_\omega$ at the 10th, 25th, 50th, 75th, and 90th percentiles across different types for $J=1$ and $6$ for concrete products. Table (ref) presents the same estimates for the electric audio equipment industry. When $J=1$, the OLS coefficient of $\omega_{it}$ is estimated significantly at $0.06$ for concrete products and at $0.5$ for electric audio equipment.
For the model with $J=6$, the estimated coefficients of $\omega_{it}$ differ considerably across different types of plants, indicating that the investment response to a productivity shock varies across plants. Both OLS and quantile regression results show that the estimated coefficients tend to be higher for the types with higher $\hat\beta_m^j$ and $\hat\beta_k^j/\hat\beta_\ell^j$, suggesting that firms invest more given a positive productivity shock if their production technology features high material shares and high capital-labor ratios. In the case of quantile regressions, this pattern is particularly pronounced for firms with high investment ratios.
Overall, these results highlight the importance of accounting for unobserved heterogeneity in the production function when estimating plant-level productivity and its impact on investment.
This paper establishes the nonparametric identifiability of production functions when unobserved heterogeneity exists across firms in the form of latent technology groups. Building on our nonparametric identification analysis and considering computational simplicity, we propose an estimation procedure for the production function with random coefficients employing a finite mixture specification.
Our analysis of Japanese plant-level panel data reveals a substantial degree of variation in estimated input elasticities and productivity growth processes across latent types within narrowly defined industries. We demonstrate that neglecting unobserved heterogeneity in input elasticities can lead to significant and systematic bias in estimated productivity growth. Moreover, we highlight the critical role played by unobserved disparities in input elasticities in plant-level investment decisions. Specifically, we find that the correlation between estimated productivity and investment is notably stronger among high capital-intensive latent type firms than among low capital-intensive type firms.
As future research topics, we may extend our framework in several directions. First, the framework may be extended to explicitly account for plant- or firm-specific biased technological change, which may be continuously distributed. This can be achieved by integrating the structural assumption discussed in doraszelski2018measuring, ZHANG2019, Demirer20, and Raval2023 into our framework. The identification approach of this paper---based on panel data with a Markov structure---is different from, but complementary to, the structural approaches of these existing papers. Adopting both identification approaches simultaneously can provide an empirical framework for estimating production functions that incorporates a broader range of unobserved heterogeneity.
Second, although we establish nonparametric identification, we employ a parametric finite mixture model for estimation due to computational complexity. An important research direction would be to relax the parametric assumption and develop a nonparametric estimator for a finite mixture model of the production function. This could be achieved, for instance, by extending the maximum smooth likelihood estimator proposed by lhc11 or the estimator presented by Bonhomme16as.
Finally, the assumption of perfect competition or monopolistic competition with constant price elasticity may not be realistic. If data on quantities and prices are available separately, our framework can be employed to estimate the production function using the quantity of output, instead of relying on sales as a proxy for output. However, it is frequently the case that firm- or plant-level datasets do not contain output quantities and prices separately. Therefore, developing a framework for concurrently identifying the production function and demand structure from revenue data, as demonstrated in the work by kasahara_sugita2020, while incorporating unobserved heterogeneity, is an important future research topic.