EconBase
← All papers

Do Third-Party Web Traffic Estimates Preserve Causal Variation?

Mehrzad Khosravi, Hema Yoganarasimhan

arXiv 24 Sep 2026 · Econometrics

arXiv:2609.30481 · PDF · Extracted main text

Abstract

Researchers increasingly rely on third-party platforms such as Similarweb and Semrush to measure web traffic when first-party analytics are unavailable. Yet these platforms report model-generated estimates rather than raw data, raising questions about whether their measures preserve the temporal and cross-source variation required for causal inference. As a motivating diagnostic, we examine reported referral traffic around two documented search-engine outages; the absence of visible discontinuities illustrates why preservation of identifying variation cannot be taken for granted. We then characterize three mechanisms, within-source smoothing, cross-source leakage, and treatment-induced calibration error, through which platform processing can generate nonclassical outcome measurement error. Analytical results and a stylized difference-in-differences simulation show that this error can attenuate, amplify, or reverse estimated treatment effects. Our findings caution against using third-party traffic measures based on black-box proprietary models for causal inference.

Citation extraction

34
references
44
in-text mentions
34
distinct cited
0
self-citations
7,186
main-text words

appendix boundary found by appendix_command · 59% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Semrush (2026) How semrush turns traffic data into traffic intelligence0.6443267%
2Similarweb (2026) Similarweb data methodology0.6443267%
3H. Zhao and R. Berman (2025) Strategic response of news publishers to generative AI0.64422100%
4R. Cao, R. Koning, and R. Nanda (2021) Sampling bias in entrepreneurial experiments0.5113233%
5K. M. Miller, J. Schmitt, and B. Skiera (2024) The impact of the general data protection regulation (GDPR) on online usage behavior0.5112250%
6N. L. Wright (2023) When local learning scales: Entrepreneurs' initial users and market expansion0.5112250%
7P. Zhou, D. Proserpio, and A. Goli (2026) LLMs as gatekeepers: Source concentration, factual quality, and political slant in information search0.5112250%
8C. S. Armstrong, Y. Konchitchki, and B. Zhang (2023) Digital traffic, financial performance, and stock valuation0.40511100%
9G. Burtch, D. Lee, and Z. Chen (2024) The consequences of generative ai for online knowledge communities0.40511100%
10J. Calzada and R. Gil (2020) What do news aggregators do? evidence from google news in spain and germany0.40511100%

Showing the top 10 of 34 scored citations.