Mehrzad Khosravi, Hema Yoganarasimhan
arXiv 24 Sep 2026 · Econometrics
arXiv:2609.30481 · PDF · Extracted main text
Researchers increasingly rely on third-party platforms such as Similarweb and Semrush to measure web traffic when first-party analytics are unavailable. Yet these platforms report model-generated estimates rather than raw data, raising questions about whether their measures preserve the temporal and cross-source variation required for causal inference. As a motivating diagnostic, we examine reported referral traffic around two documented search-engine outages; the absence of visible discontinuities illustrates why preservation of identifying variation cannot be taken for granted. We then characterize three mechanisms, within-source smoothing, cross-source leakage, and treatment-induced calibration error, through which platform processing can generate nonclassical outcome measurement error. Analytical results and a stylized difference-in-differences simulation show that this error can attenuate, amplify, or reverse estimated treatment effects. Our findings caution against using third-party traffic measures based on black-box proprietary models for causal inference.
appendix boundary found by appendix_command · 59% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Semrush (2026) How semrush turns traffic data into traffic intelligence | 0.644 | 3 | 2 | 67% |
| 2 | Similarweb (2026) Similarweb data methodology | 0.644 | 3 | 2 | 67% |
| 3 | H. Zhao and R. Berman (2025) Strategic response of news publishers to generative AI | 0.644 | 2 | 2 | 100% |
| 4 | R. Cao, R. Koning, and R. Nanda (2021) Sampling bias in entrepreneurial experiments | 0.511 | 3 | 2 | 33% |
| 5 | K. M. Miller, J. Schmitt, and B. Skiera (2024) The impact of the general data protection regulation (GDPR) on online usage behavior | 0.511 | 2 | 2 | 50% |
| 6 | N. L. Wright (2023) When local learning scales: Entrepreneurs' initial users and market expansion | 0.511 | 2 | 2 | 50% |
| 7 | P. Zhou, D. Proserpio, and A. Goli (2026) LLMs as gatekeepers: Source concentration, factual quality, and political slant in information search | 0.511 | 2 | 2 | 50% |
| 8 | C. S. Armstrong, Y. Konchitchki, and B. Zhang (2023) Digital traffic, financial performance, and stock valuation | 0.405 | 1 | 1 | 100% |
| 9 | G. Burtch, D. Lee, and Z. Chen (2024) The consequences of generative ai for online knowledge communities | 0.405 | 1 | 1 | 100% |
| 10 | J. Calzada and R. Gil (2020) What do news aggregators do? evidence from google news in spain and germany | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 34 scored citations.