EconBase
← All papers

On Local Overidentification and Efficiency Gains in Modern Causal Inference and Data Combination

Xiaohong Chen, Haitian Xie

arXiv 19 Oct 2025 · Econometrics

arXiv:2510.16683 · PDF · DOI · OpenAlex · Extracted main text

Abstract

This paper studies nonparametric local (over-)identification and the semiparametric efficiency in modern causal frameworks. We develop a unified approach that begins by translating structural models with latent variables into their induced statistical models of observables and then analyzes local overidentification through conditional moment restrictions. We apply this approach to three popular classes of causal models: (1) the general treatment model under unconfoundedness; (2) the negative control model, and (3) the long-term causal inference model under unobserved confounding. The first model yields a locally just-identified statistical model, implying that all regular asymptotically linear estimators of the treatment effect have the same asymptotic variance, which equals the (trivial) semiparametric efficient variance bound. In contrast, the latter two models involve nonparametric endogeneity and are naturally locally overidentified; consequently, some doubly robust orthogonal moment estimators of the average treatment effect are inefficient. Whereas existing work typically imposes strong conditions to restore local just-identification to justify the efficiency of their doubly robust orthogonal moment estimators, we characterize the semiparametric efficient variance bounds, along with efficient estimators, for the (locally) overidentified models (2) and (3). A small real data application, along with a simulation study, illustrates the semiparametric efficiency gains in model (3).

Citation extraction

34
references
83
in-text mentions
34
distinct cited
11
self-citations
8,614
main-text words

appendix boundary found by appendix_command · 63% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Imbens, Guido and Kallus, Nathan and Mao, Xiaojie and Wang, Yuhao (2024) Long-term causal inference under persistent confounding via data combination0.94118483%
2Chen, Xiaohong and Santos, Andres (2018) Overidentification in regular models self0.9098475%
3Chunrong Ai and Xiaohong Chen (2012) The semiparametric efficiency bound for models of sequential moment restrictions containing unknown functions self0.84310460%
4Chen, Xiaohong and Hong, Han and Tarozzi, Alessandro (2004) Semiparametric Efficiency in GMM Models of Nonclassical Measurement Error, Missing Data and Treatment Effects self0.84333100%
5Ai, Chunrong and Linton, Oliver and Motegi, Kaiji and Zhang, Zheng (2021) A unified framework for efficient estimation of general treatment models0.81142100%
6Cui, Yifan and Pu, Hongming and Shi, Xu and Miao, Wang and Tchetgen… (2024) Semiparametric proximal causal inference0.7374275%
7Hirano, Keisuke and Imbens, Guido W and Ridder, Geert (2003) Efficient estimation of average treatment effects using the estimated propensity score0.73732100%
8Ai, Chunrong and Chen, Xiaohong (2003) Efficient Estimation of Models with Conditional Moment Restrictions Containing Unknown Functions self0.64422100%
9Chunrong Ai and Xiaohong Chen (2007) Estimation of possibly misspecified semiparametric conditional moment restriction models with different conditioning variables self0.64422100%
10Xiaohong Chen and Han Hong and Alessandro Tarozzi (2008) Semiparametric efficiency in GMM models with auxiliary data self0.64422100%

Showing the top 10 of 34 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Higher-Order Debiased Estimators for General Treatment Models0.51122
2Semiparametric Efficiency in Policy Learning with General Treatments0.40511