EconBase
← All papers

Measuring Trade Direction in a Prediction Market: Settlement Ground Truth and Trading-Cost Measurement on Polymarket

Philipp D. Dubach

arXiv 5 Oct 2026 · Finance — Trading

arXiv:2610.06412 · PDF · Extracted main text

Abstract

Trade-sign errors can change measured trading costs even when classification accuracy is high. We validate the side of Polymarket's public trade prints against the taker leg of each print's on-chain settlement. On twelve selected days between April and August 2026, spanning both exchange generations, 92.1% to 100.0% of prints match a settled taker leg, with exact side and token agreement on all 24.5 million matched pairs. Mint-and-merge settlement makes that taker leg essential: pooling maker and taker legs changes the measured buy share. Signing the full cached tape before settlement selection yields 16.6 million prints signed by every rule; equal-day balanced accuracy is 0.942 for Lee-Ready, 0.768 for the tick test and 0.670 for retrospective bulk volume classification. On 16.4 million identical eligible fills, Lee-Ready raises effective spread by 0.429 cents per share under equal-fill weights; its realised-spread difference is -0.209 cents. Under share-volume weights, taker five-minute midpoint impact is 0.689 cents, while tick and bulk classifications give -0.374 and -0.119 cents. That aggregate sign reversal disappears when tied print rows are excluded. Distortion depends on error-weighted signed outcomes, sample selection and weighting. Receipt-time ordering and unobserved future-quote age limit these delivered-quote accounting quantities; they do not identify causal impact or private information. Supporting venue and collector analyses provide descriptive diagnostics.

Citation extraction

43
references
59
in-text mentions
43
distinct cited
2
self-citations
9,353
main-text words

appendix boundary found by appendix_command · 44% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1David Easley, Marcos M. López de Prado, and Maureen O'Hara (2016) Discerning information from trade data0.7373367%
2Katrina Ellis, Roni Michaely, and Maureen O'Hara (2000) The accuracy of trade classification rules: Evidence from Nasdaq0.7373367%
3Bidisha Chakrabarty, Roberto Pascual, and Andriy Shkilko (2015) Evaluating trade classification algorithms: Bulk volume classification versus the tick rule and the Lee–Ready algorithm0.64422100%
4Philipp D. Dubach (2026) The anatomy of a decentralized prediction market: Microstructure evidence from the polymarket order book, 2026a self0.64422100%
5Simon Jurkatis (2021) Inferring trade directions in fast markets0.64422100%
6Elizabeth R. Odders-White (2000) On the occurrence and consequences of inaccurate trade classification0.64422100%
7Boka Qin and Rui Yang (2026) Polymarket-v1 database, 20260.64422100%
8Yiming Shen, Yuhan Jin, Shuohan Wu, Yanlin Wang, and Jiachi Chen (2026) The ghosts of Polymarket: When off-chain matches meet on-chain reverts, 20260.64422100%
9Charles M. C. Lee and Mark J. Ready Inferring trade direction from intraday data0.5112250%
10PMXT (2026) Polymarket orderbook archive (v2): data overview0.5112250%

Showing the top 10 of 43 scored citations.