arXiv 28 Jan 2026 · Statistics — Methodology
arXiv:2601.20197 · PDF · DOI · OpenAlex · Extracted main text
Finite mixture models are widely used in econometric analyses to capture unobserved heterogeneity. This paper shows that maximum likelihood estimation of finite mixtures of parametric densities can suffer from substantial finite-sample bias in all parameters under mild regularity conditions. The bias arises from the influence of outliers in component densities with unbounded or large support and increases with the degree of overlap among mixture components. I show that maximizing the classification-mixture likelihood function, equipped with a consistent classifier, yields parameter estimates that are less biased than those obtained by standard maximum likelihood estimation (MLE). I then derive the asymptotic distribution of the resulting estimator and provide conditions under which oracle efficiency is achieved. Monte Carlo simulations show that conventional mixture MLE exhibits pronounced finite-sample bias, which diminishes as the sample size or the statistical distance between component densities tends to infinity. The simulations further show that the proposed estimation strategy generally outperforms standard MLE in finite samples in terms of both bias and mean squared errors under relatively weak assumptions. An empirical application to latent group panel structures using health administrative data shows that the proposed approach reduces out-of-sample prediction error by approximately 17.6% relative to the best results obtained from standard MLE procedures.
appendix boundary found by appendix_command · 72% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Su, Liangjun and Shi, Zhentao and Phillips, Peter C. B (2016) Identifying Latent Structures in Panel Data | 1.000 | 7 | 3 | 100% |
| 2 | Chen, Jiahua (2023) Statistical Inference Under Mixture Models | 1.000 | 6 | 3 | 100% |
| 3 | Redner, Richard A. and Walker, Homer F (1984) Mixture Densities, Maximum Likelihood and the Em Algorithm | 1.000 | 6 | 3 | 100% |
| 4 | Frühwirth-Schnatter, Sylvia (2006) Finite mixture and Markov switching models | 0.928 | 4 | 3 | 100% |
| 5 | Celeux, Gilles and Govaert, Gérard (1992) A classification EM algorithm for clustering and two stochastic versions | 0.874 | 5 | 2 | 100% |
| 6 | Dempster, A. P. and Laird, N. M. and Rubin, D. B (1977) Maximum Likelihood from Incomplete Data via the EM Algorithm | 0.843 | 3 | 3 | 100% |
| 7 | McLachlan, Geoffrey J. and Lee, Sharon X. and Rathnayake, Suren I (2019) Finite Mixture Models | 0.843 | 3 | 3 | 100% |
| 8 | Balakrishnan, Sivaraman and Wainwright, Martin J. and Yu, Bin (2017) Statistical guarantees for the EM algorithm: From population to sample-based analysis | 0.811 | 4 | 2 | 100% |
| 9 | Tanaka, Kentaro (2009) Strong Consistency of the Maximum Likelihood Estimator for Finite Mixtures of Location-Scale Distributions When Penalty is Impos… | 0.811 | 4 | 2 | 100% |
| 10 | Bonhomme, Stéphane and Manresa, Elena (2015) Grouped Patterns of Heterogeneity in Panel Data | 0.737 | 3 | 2 | 100% |
Showing the top 10 of 88 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Policy Learning with Observational Data : The Case of Hepatitis C Treatment for HIV/HCV Co-Infected Patients | 0.737 | 5 | 2 |