EconBase
← All papers

Partial Identification under Missing Data Using Weak Shadow Variables from Pretrained Models

Hongyu Chen, David Simchi-Levi, Ruoxuan Xiong

arXiv 17 Feb 2026 · Statistics — Machine Learning

arXiv:2602.16061 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Estimating population quantities such as mean outcomes from user feedback is fundamental to platform evaluation and social science, yet feedback is often missing not at random (MNAR): users with stronger opinions are more likely to respond, so standard estimators are biased and the estimand is not identified without additional assumptions. Existing approaches typically rely on strong parametric assumptions or bespoke auxiliary variables that may be unavailable in practice. In this paper, we develop a partial identification framework in which sharp bounds on the estimand are obtained by solving a pair of linear programs whose constraints encode the observed data structure. This formulation naturally incorporates outcome predictions from pretrained models, including large language models (LLMs), as additional linear constraints that tighten the feasible set. We call these predictions weak shadow variables: they satisfy a conditional independence assumption with respect to missingness but need not meet the completeness conditions required by classical shadow-variable methods. When predictions are sufficiently informative, the bounds collapse to a point, recovering standard identification as a special case. In finite samples, to provide valid coverage of the identified set, we propose a set-expansion estimator that achieves slower-than-$\sqrt{n}$ convergence rate in the set-identified regime and the standard $\sqrt{n}$ rate under point identification. In simulations and semi-synthetic experiments on customer-service dialogues, we find that LLM predictions are often ill-conditioned for classical shadow-variable methods yet remain highly effective in our framework. They shrink identification intervals by 75--83% while maintaining valid coverage under realistic MNAR mechanisms.

Citation extraction

36
references
60
in-text mentions
36
distinct cited
0
self-citations
17,101
main-text words

appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Heckman JJ (1979) Sample selection bias as a specification error0.92843100%
2Little RJ (1994) A class of pattern-mixture models for normal incomplete data0.92843100%
3Miao W, Liu L, Li Y, Tchetgen Tchetgen EJ, Geng Z (2024) Identification and semiparametric efficiency theory of nonignorable missing data with a shadow variable0.92843100%
4Miao W, Tchetgen Tchetgen EJ (2016) On varieties of doubly robust estimators under missingness not at random with a shadow variable0.87462100%
5Hu N, Pavlou PA, Zhang J (2017) On self-selection biases in online product reviews0.84333100%
6Chernozhukov V, Hong H, Tamer E (2007) Estimation and confidence regions for parameter sets in econometric models0.81142100%
7Rubin DB (1987) The calculation of posterior distributions by data augmentation: Comment: A noniterative sampling/importance resampling alternat…0.73732100%
8Angelopoulos AN, Bates S, Fannjiang C, Jordan MI, Zrnic T (2023) a) Prediction-powered inference0.64422100%
9d’Haultfoeuille X (2010) A new instrumental method for dealing with endogenous selection0.51121100%
10Manski CF (2003) Partial identification of probability distributions0.51121100%

Showing the top 10 of 36 scored citations.