arXiv 29 May 2026 · Econometrics
arXiv:2605.30718 · PDF · DOI · OpenAlex · Extracted main text
Topic models are often used as dimension-reduction tools before regression, with estimated document-level topic shares treated as observed covariates. This plug-in workflow creates two inferential difficulties: valid inference requires a regular first-stage-to-second-stage expansion that propagates topic-estimation uncertainty, and, at fixed document length, a document's topic mixture cannot be consistently recovered from its own words even when the population topic matrix is known. Corrected spectral moment methods for latent Dirichlet allocation (LDA) offer a starting point: when the total Dirichlet concentration is known, low-order word moments can be corrected to yield operators diagonal in the latent topic basis. We extend this to downstream regression. Under a finite LDA model with response residuals orthogonal to the low-order token moments used for identification, response-weighted word moments admit the same correction, and the resulting supervised operator identifies the regression coefficient $β$ directly, without estimating document-level topic shares. The main obstacle is that the correction depends on the unknown total concentration $α_0$. We show that, for $k\ge3$ topics and under a generic finite-probe condition, $α_0$ is identified by commutativity: at the true value a family of corrected word-moment operators commute, whereas away from it they generically do not. This yields a feasible estimator and lets uncertainty in $\hatα_0$ propagate into inference for $β$. The estimator is asymptotically linear as the number of documents grows with fixed document length, with sandwich standard errors from document-level moment contributions. Simulations show near-nominal coverage where plug-in topic-share regressions can undercover, and an application to top economics journals illustrates contrast inference for latent topic effects.
appendix boundary found by appendix_command · 50% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Ren, Yong and Wang, Yining and Zhu, Jun (2018) Spectral Learning for Supervised Topic Models | 0.511 | 2 | 1 | 100% |
| 2 | Wang, Yining and Zhu, Jun (2014) Spectral Methods for Supervised Topic Models | 0.511 | 2 | 1 | 100% |
| 3 | Anandkumar, Anima and Foster, Dean P. and Hsu, Daniel J. and Kakade,… (2012) A Spectral Algorithm for Latent Dirichlet Allocation | 0.405 | 1 | 1 | 100% |
| 4 | Anauati, María Victoria and Galiani, Sebastian and Gálvez, Ramiro H (2016) Quantifying the Life Cycle of Scholarly Articles across Fields of Economic Research | 0.405 | 1 | 1 | 100% |
| 5 | Bai, Jushan (2003) Inferential Theory for Factor Models of Large Dimensions | 0.405 | 1 | 1 | 100% |
| 6 | Bai, Jushan and Ng, Serena (2006) Confidence Intervals for Diffusion Index Forecasts and Inference for Factor-Augmented Regressions | 0.405 | 1 | 1 | 100% |
| 7 | Battaglia, Laura and Christensen, Timothy and Hansen, Stephen and Sa… (2024) Inference for Regression with Variables Generated by AI or Machine Learning | 0.405 | 1 | 1 | 100% |
| 8 | Bybee, Leland and Kelly, Bryan and Manela, Asaf and Xiu, Dacheng (2024) Business News and Business Cycles | 0.405 | 1 | 1 | 100% |
| 9 | Card, David and DellaVigna, Stefano (2013) Nine Facts about Top Journals in Economics | 0.405 | 1 | 1 | 100% |
| 10 | Hansen, Stephen and McMahon, Michael and Prat, Andrea (2018) Transparency and Deliberation within the FOMC: A Computational Linguistics Approach | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 16 scored citations.