EconBase
← All papers

Compound Estimation for Binomials

Yan Chen, Lihua Lei

arXiv 31 Dec 2025 · Econometrics

arXiv:2512.25042 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Many applications involve estimating the mean of multiple binomial outcomes as a common problem -- assessing intergenerational mobility of census tracts, estimating prevalence of infectious diseases across countries, and measuring click-through rates for different demographic groups. The most standard approach is to report the plain average of each outcome. Despite simplicity, the estimates are noisy when the sample sizes or mean parameters are small. In contrast, the Empirical Bayes (EB) methods are able to boost the average accuracy by borrowing information across tasks. Nevertheless, the EB methods require a Bayesian model where the parameters are sampled from a prior distribution which, unlike the commonly-studied Gaussian case, is unidentified due to discreteness of binomial measurements. Even if the prior distribution is known, the computation is difficult when the sample sizes are heterogeneous as there is no simple joint conjugate prior for the sample size and mean parameter. In this paper, we consider the compound decision framework which treats the sample size and mean parameters as fixed quantities. We develop an approximate Stein's Unbiased Risk Estimator (SURE) for the average mean squared error given any class of estimators. For a class of machine learning-assisted linear shrinkage estimators, we establish asymptotic optimality, regret bounds, and valid inference. Unlike existing work, we work with the binomials directly without resorting to Gaussian approximations. This allows us to work with small sample sizes and/or mean parameters in both one-sample and two-sample settings. We demonstrate our approach using three datasets on firm discrimination, education outcomes, and innovation rates.

Citation extraction

47
references
83
in-text mentions
47
distinct cited
1
self-citations
11,511
main-text words

appendix boundary found by appendix_command · 28% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Chen, Jiafeng (2025) Empirical Bayes when estimation precision predicts parameters1.00083100%
2Xie, Xianchao and Kou, SC and Brown, Lawrence D (2012) SURE estimates for a heteroscedastic hierarchical model0.87492100%
3Bell, Alex and Chetty, Raj and Jaravel, Xavier and Petkova, Neviana… (2019) Who becomes an inventor in America? The importance of exposure to innovation0.81142100%
4Kline, Patrick and Rose, Evan K. and Walters, Christopher R (2024) A Discrimination Report Card0.81142100%
5Chen, Jiafeng and Lei, Lihua and Sudijono, Timothy and Sun, Liyang a… (2025) Compound Selection Decisions: An Almost SURE Approach self0.73732100%
6Kline, Patrick and Walters, Christopher (2021) Reasonable Doubt: Experimental Detection of Job-Level Employment Discrimination0.73732100%
7Li, Jessie (2024) Inference for constrained extremum estimators0.64441100%
8Gang, Bowen and Fu, Luella and James, Gareth and Sun, Wenguang (2023) Ranking and Selection in Large-Scale Inference of Heteroscedastic Units0.58531100%
9Leiner, James and Duan, Boyan and Wasserman, Larry and Ramdas, Aaditya (2025) Data fission: splitting a single data point0.58531100%
10Neufeld, Anna and Dharamshi, Ameer and Gao, Lucy L and Witten, Daniela (2024) Data thinning for convolution-closed distributions0.58531100%

Showing the top 10 of 47 scored citations.