EconBase
← All papers

Missing at Random or Not: A Semiparametric Testing Approach

Rui Duan, C. Jason Liang, Pamela Shaw, Cheng Yong Tang, Yong Chen

arXiv 25 Mar 2020 · Statistics — Methodology · 1 citations (OpenAlex)

arXiv:2003.11181 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Practical problems with missing data are common, and statistical methods have been developed concerning the validity and/or efficiency of statistical procedures. On a central focus, there have been longstanding interests on the mechanism governing data missingness, and correctly deciding the appropriate mechanism is crucially relevant for conducting proper practical investigations. The conventional notions include the three common potential classes -- missing completely at random, missing at random, and missing not at random. In this paper, we present a new hypothesis testing approach for deciding between missing at random and missing not at random. Since the potential alternatives of missing at random are broad, we focus our investigation on a general class of models with instrumental variables for data missing not at random. Our setting is broadly applicable, thanks to that the model concerning the missing data is nonparametric, requiring no explicit model specification for the data missingness. The foundational idea is to develop appropriate discrepancy measures between estimators whose properties significantly differ only when missing at random does not hold. We show that our new hypothesis testing approach achieves an objective data oriented choice between missing at random or not. We demonstrate the feasibility, validity, and efficacy of the new test by theoretical analysis, simulation studies, and a real data analysis.

Citation extraction

46
references
67
in-text mentions
46
distinct cited
3
self-citations
8,003
main-text words

appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Zhao, J. and J. Shao (2015) Semiparametric pseudo-likelihoods in generalized linear models with non-ignorable missing data1.00083100%
2Little, R. J. A. and D. B. Rubin (2019) Statistical Analysis with Missing Data\/ (3rd ed.)0.73732100%
3Molenberghs, G. and M. G. Kenward (2007) Missing Data in Clinical Studies0.73732100%
4Kim, J. K. and J. Shao (2013) Statistical methods for handling incomplete data0.64422100%
5Tsiatis, A. A (2006) Semiparametric Theory and Missing Data0.64422100%
6Wang, L., J. Shao, and F. Fang (2019) Propensity model selection with nonignorable nonresponse and instrument variable0.64422100%
7White, I. R. and J. B. Carlin (2010, sep) (2010) Bias and efficiency of multiple imputation compared with complete-case analysis for missing covariate values0.64422100%
8Hausman, J. A. (1978, nov) (1978) Specification tests in econometrics0.64422100%
9Kim, J. K. and C. L. Yu (2011) A semi-parametric estimation of mean functionals with non-ignorable missing data0.51121100%
10Molenberghs, G., G. Fitzmaurice, M. G. Kenward, A. Tsiatis, and G. V… (2014) Handbook of Missing Data Methodology0.51121100%

Showing the top 10 of 46 scored citations.