EconBase
← All papers

Moment-Based Inference for Regression with Latent Dirichlet Covariates

Ziyu Jiang

arXiv 29 May 2026 · Econometrics

arXiv:2605.30718 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Topic models are often used as dimension-reduction tools before regression, with estimated document-level topic shares treated as observed covariates. This plug-in workflow creates two inferential difficulties: valid inference requires a regular first-stage-to-second-stage expansion that propagates topic-estimation uncertainty, and, at fixed document length, a document's topic mixture cannot be consistently recovered from its own words even when the population topic matrix is known. Corrected spectral moment methods for latent Dirichlet allocation (LDA) offer a starting point: when the total Dirichlet concentration is known, low-order word moments can be corrected to yield operators diagonal in the latent topic basis. We extend this to downstream regression. Under a finite LDA model with response residuals orthogonal to the low-order token moments used for identification, response-weighted word moments admit the same correction, and the resulting supervised operator identifies the regression coefficient $β$ directly, without estimating document-level topic shares. The main obstacle is that the correction depends on the unknown total concentration $α_0$. We show that, for $k\ge3$ topics and under a generic finite-probe condition, $α_0$ is identified by commutativity: at the true value a family of corrected word-moment operators commute, whereas away from it they generically do not. This yields a feasible estimator and lets uncertainty in $\hatα_0$ propagate into inference for $β$. The estimator is asymptotically linear as the number of documents grows with fixed document length, with sandwich standard errors from document-level moment contributions. Simulations show near-nominal coverage where plug-in topic-share regressions can undercover, and an application to top economics journals illustrates contrast inference for latent topic effects.

Citation extraction

16
references
18
in-text mentions
16
distinct cited
0
self-citations
13,088
main-text words

appendix boundary found by appendix_command · 50% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Ren, Yong and Wang, Yining and Zhu, Jun (2018) Spectral Learning for Supervised Topic Models0.51121100%
2Wang, Yining and Zhu, Jun (2014) Spectral Methods for Supervised Topic Models0.51121100%
3Anandkumar, Anima and Foster, Dean P. and Hsu, Daniel J. and Kakade,… (2012) A Spectral Algorithm for Latent Dirichlet Allocation0.40511100%
4Anauati, María Victoria and Galiani, Sebastian and Gálvez, Ramiro H (2016) Quantifying the Life Cycle of Scholarly Articles across Fields of Economic Research0.40511100%
5Bai, Jushan (2003) Inferential Theory for Factor Models of Large Dimensions0.40511100%
6Bai, Jushan and Ng, Serena (2006) Confidence Intervals for Diffusion Index Forecasts and Inference for Factor-Augmented Regressions0.40511100%
7Battaglia, Laura and Christensen, Timothy and Hansen, Stephen and Sa… (2024) Inference for Regression with Variables Generated by AI or Machine Learning0.40511100%
8Bybee, Leland and Kelly, Bryan and Manela, Asaf and Xiu, Dacheng (2024) Business News and Business Cycles0.40511100%
9Card, David and DellaVigna, Stefano (2013) Nine Facts about Top Journals in Economics0.40511100%
10Hansen, Stephen and McMahon, Michael and Prat, Andrea (2018) Transparency and Deliberation within the FOMC: A Computational Linguistics Approach0.40511100%

Showing the top 10 of 16 scored citations.