arXiv 9 Jul 2022 · Statistics — Methodology
arXiv:2207.04299 · PDF · DOI · OpenAlex · Extracted main text
Model diagnostics is an indispensable component of regression analysis, yet it is not well addressed in standard textbooks on generalized linear models. The lack of exposition is attributed to the fact that when outcome data are discrete, classical methods (e.g., Pearson/deviance residual analysis and goodness-of-fit tests) have limited utility in model diagnostics and treatment. This paper establishes a novel framework for model diagnostics of discrete data regression. Unlike the literature defining a single-valued quantity as the residual, we propose to use a function as a vehicle to retain the residual information. In the presence of discreteness, we show that such a functional residual is appropriate for summarizing the residual randomness that cannot be captured by the structural part of the model. We establish its theoretical properties, which leads to the innovation of new diagnostic tools including the functional-residual-vs covariate plot and Function-to-Function (Fn-Fn) plot. Our numerical studies demonstrate that the use of these tools can reveal a variety of model misspecifications, such as not properly including a higher-order term, an explanatory variable, an interaction effect, a dispersion parameter, or a zero-inflation component. The functional residual yields, as a byproduct, Liu-Zhang's surrogate residual mainly developed for cumulative link models for ordinal data (Liu and Zhang, 2018, JASA). As a general notion, it considerably broadens the diagnostic scope as it applies to virtually all parametric models for binary, ordinal and count data, all in a unified diagnostic scheme.
appendix boundary found by appendix_titled_section at “Appendix A. Additional examples for section 3.2” · 76% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Liu, D. and Zhang, H (2018) Residuals and diagnostics for ordinal regression models: a surrogate approach self | 1.000 | 7 | 4 | 100% |
| 2 | Liu, D., Li, S., Yu, Y., and Moustaki, I (2021) Assessing partial association between ordinal variables: quantification, visualization, and hypothesis testing self | 1.000 | 5 | 3 | 100% |
| 3 | Li, C. and Shepherd, B (2012) A new residual for ordinal outcomes | 0.928 | 4 | 3 | 100% |
| 4 | Franses, P. H. and Paap, R (2001) Quantitative Models in Marketing Research | 0.644 | 2 | 2 | 100% |
| 5 | Li, C. and Shepherd, B (2010) Test of association between two ordinal variables while adjusting for covariates | 0.644 | 2 | 2 | 100% |
| 6 | Fanaee-T, H. and Gama, J (2014) Event labeling combining ensemble detectors and background knowledge | 0.511 | 2 | 1 | 100% |
| 7 | Wasserstein, R. L. and Lazar, N. A (2016) The ASA Statement on p-Values: Context, Process, and Purpose | 0.511 | 2 | 1 | 100% |
| 8 | Cortez, P., Cerdeira, A., Almeida, F., Matos, T., and Reis, J (2009) Modeling wine preferences by data mining from physicochemical properties | 0.511 | 2 | 1 | 100% |
| 9 | Archer, K. J., Lemeshow, S., and Hosmer, D. W (2007) Goodness-of-fit tests for logistic regression models when data are collected using a complex sampling design | 0.405 | 1 | 1 | 100% |
| 10 | Archer, K. J. and Lemeshow, S (2006) Goodness-of-fit test for a logistic regression model fitted using survey sample data | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 30 scored citations.