EconBase
← All papers

Estimating the Number of Components in Panel Data Finite Mixture Regression Models with an Application to Production Function Heterogeneity

Yu Hao, Hiroyuki Kasahara

arXiv 11 Jun 2025 · Econometrics

arXiv:2506.09666 · PDF · DOI · OpenAlex · Extracted main text

Abstract

This paper develops statistical methods for determining the number of components in panel data finite mixture regression models with regression errors independently distributed as normal or more flexible normal mixtures. We analyze the asymptotic properties of the likelihood ratio test (LRT) and information criteria (AIC and BIC) for model selection in both conditionally independent and dynamic panel settings. Unlike cross-sectional normal mixture models, we show that panel data structures eliminate higher-order degeneracy problems while retaining issues of unbounded likelihood and infinite Fisher information. Addressing these challenges, we derive the asymptotic null distribution of the LRT statistic as the maximum of random variables and develop a sequential testing procedure for consistent selection of the number of components. Our theoretical analysis also establishes the consistency of BIC and the inconsistency of AIC. Empirical application to Chilean manufacturing data reveals significant heterogeneity in production technology, with substantial variation in output elasticities of material inputs and factor-augmented technological processes within narrowly defined industries, indicating plant-specific variation in production functions beyond Hicks-neutral technological differences. These findings contrast sharply with the standard practice of assuming a homogeneous production function and highlight the necessity of accounting for unobserved plant heterogeneity in empirical production analysis.

Citation extraction

67
references
145
in-text mentions
67
distinct cited
8
self-citations
20,303
main-text words

appendix boundary found by appendix_command · 43% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Kasahara, H. and Shimotsu, K (2014) Non-parametric identification and estimation of the number of components in multivariate mixtures self1.00063100%
2Liu, X. and Shao, Y (2003) Asymptotics for likelihood ratio tests under loss of identifiability0.9285380%
3Kasahara, H., Schrimpf, P., and Suzuki, M (2022) Identification and estimation of production function with unobserved heterogeneity self0.87462100%
4Kleibergen, F. and Paap, R (2006) Generalized reduced rank tests using the singular value decomposition0.87462100%
5Kasahara, H. and Shimotsu, K (2019) Testing the Order of Multivariate Normal Mixture Models self0.87452100%
6Doraszelski, U. and Jaumandreu, J (2018) Measuring the bias of technological change0.84333100%
7Keribin, C (2000) Consistent estimation of the order of mixture models0.81142100%
8Kasahara, H. and Shimotsu, K (2015) Testing the number of components in normal mixture regression models self0.7946350%
9Kasahara, H. and Shimotsu, K (2009) Nonparametric Identification of Finite Mixture Models of Dynamic Discrete Choices self0.73732100%
10Dacunha-Castelle, D. and Gassiat, E (1999) Testing the order of a model using locally conic parametrization: Population mixtures and stationary ARMA processes0.73732100%

Showing the top 10 of 67 scored citations.