EconBase
← All papers

Model diagnostics of discrete data regression: a unifying framework using functional residuals

Zewei Lin, Dungang Liu

arXiv 9 Jul 2022 · Statistics — Methodology

arXiv:2207.04299 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Model diagnostics is an indispensable component of regression analysis, yet it is not well addressed in standard textbooks on generalized linear models. The lack of exposition is attributed to the fact that when outcome data are discrete, classical methods (e.g., Pearson/deviance residual analysis and goodness-of-fit tests) have limited utility in model diagnostics and treatment. This paper establishes a novel framework for model diagnostics of discrete data regression. Unlike the literature defining a single-valued quantity as the residual, we propose to use a function as a vehicle to retain the residual information. In the presence of discreteness, we show that such a functional residual is appropriate for summarizing the residual randomness that cannot be captured by the structural part of the model. We establish its theoretical properties, which leads to the innovation of new diagnostic tools including the functional-residual-vs covariate plot and Function-to-Function (Fn-Fn) plot. Our numerical studies demonstrate that the use of these tools can reveal a variety of model misspecifications, such as not properly including a higher-order term, an explanatory variable, an interaction effect, a dispersion parameter, or a zero-inflation component. The functional residual yields, as a byproduct, Liu-Zhang's surrogate residual mainly developed for cumulative link models for ordinal data (Liu and Zhang, 2018, JASA). As a general notion, it considerably broadens the diagnostic scope as it applies to virtually all parametric models for binary, ordinal and count data, all in a unified diagnostic scheme.

Citation extraction

30
references
48
in-text mentions
30
distinct cited
2
self-citations
8,551
main-text words

appendix boundary found by appendix_titled_section at “Appendix A. Additional examples for section 3.2” · 76% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Liu, D. and Zhang, H (2018) Residuals and diagnostics for ordinal regression models: a surrogate approach self1.00074100%
2Liu, D., Li, S., Yu, Y., and Moustaki, I (2021) Assessing partial association between ordinal variables: quantification, visualization, and hypothesis testing self1.00053100%
3Li, C. and Shepherd, B (2012) A new residual for ordinal outcomes0.92843100%
4Franses, P. H. and Paap, R (2001) Quantitative Models in Marketing Research0.64422100%
5Li, C. and Shepherd, B (2010) Test of association between two ordinal variables while adjusting for covariates0.64422100%
6Fanaee-T, H. and Gama, J (2014) Event labeling combining ensemble detectors and background knowledge0.51121100%
7Wasserstein, R. L. and Lazar, N. A (2016) The ASA Statement on p-Values: Context, Process, and Purpose0.51121100%
8Cortez, P., Cerdeira, A., Almeida, F., Matos, T., and Reis, J (2009) Modeling wine preferences by data mining from physicochemical properties0.51121100%
9Archer, K. J., Lemeshow, S., and Hosmer, D. W (2007) Goodness-of-fit tests for logistic regression models when data are collected using a complex sampling design0.40511100%
10Archer, K. J. and Lemeshow, S (2006) Goodness-of-fit test for a logistic regression model fitted using survey sample data0.40511100%

Showing the top 10 of 30 scored citations.