arXiv 18 Mar 2024 · Statistics — Methodology
arXiv:2403.11954 · PDF · DOI · OpenAlex · Extracted main text
While there is a rich literature on robust methodologies for contamination in continuously distributed data, contamination in categorical data is largely overlooked. This is regrettable because many datasets are categorical and oftentimes suffer from contamination. Examples include inattentive responding and bot responses in questionnaires or zero-inflated count data. We propose a novel class of contamination-robust estimators of models for categorical data, coined $C$-estimators (“$C$” for categorical). We show that the countable and possibly finite sample space of categorical data results in non-standard theoretical properties. Notably, in contrast to classic robustness theory, $C$-estimators can be simultaneously robust and fully efficient at the postulated model. In addition, a certain particularly robust specification fails to be asymptotically Gaussian at the postulated model, but is asymptotically Gaussian in the presence of contamination. We furthermore propose a diagnostic test to identify categorical outliers and demonstrate the enhanced robustness of $C$-estimators in a simulation study.
appendix boundary found by appendix_command · 41% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Lindsay, B. G (1994) Efficiency versus robustness: The case for minimum Hellinger distance and related methods | 1.000 | 16 | 5 | 100% |
| 2 | Victoria-Feser, M.-P. & Ronchetti, E. M (1997) Robust estimation for grouped data | 1.000 | 5 | 3 | 100% |
| 3 | Ruckstuhl, A. F. & Welsh, A. H (2001) Robust fitting of the binomial model | 0.974 | 13 | 7 | 92% |
| 4 | Hampel, F. R., Ronchetti, E. M., Rousseeuw, P. J., & Stahel, W. A (1986) Robust statistics: The approach based on influence functions | 0.928 | 4 | 3 | 100% |
| 5 | Markatou, M., Basu, A., & Lindsay, B (1997) Weighted likelihood estimating equations: The discrete case with applications to logistic regression | 0.928 | 4 | 3 | 100% |
| 6 | Huber, P. J (1964) Robust estimation of a location parameter | 0.843 | 3 | 3 | 100% |
| 7 | Arias, V. B., Garrido, L., Jenaro, C., Martinez-Molina, A., & Arias, B (2020) A little garbage in, lots of garbage out: Assessing the impact of careless responding in personality survey data | 0.737 | 3 | 2 | 100% |
| 8 | Csiszár, I (1963) Eine informationstheoretische Ungleichung und ihre Anwendung auf den Beweis der Ergodizität von Markoffschen Ketten | 0.737 | 3 | 2 | 100% |
| 9 | Huber, P. J. & Ronchetti, E. M (2009) Robust Statistics | 0.737 | 3 | 2 | 100% |
| 10 | Ilagan, M. J. & Falk, C. F (2023) Supervised classes, unsupervised mixing proportions: Detection of bots in a Likert-type questionnaire | 0.737 | 3 | 2 | 100% |
Showing the top 10 of 27 scored citations.