The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
63,842 characters
Measuring Trade Direction in a Prediction Market: Settlement Ground Truth and Trading-Cost Measurement on Polymarket
\maketitle
\begin{abstract}
\noindent
Trade-sign errors can change measured trading costs even when classification accuracy is high.
We validate the side of Polymarket's public trade prints against the taker leg of each print's
on-chain settlement. On twelve selected days between April and August 2026, spanning both
exchange generations, \SI{92.1}{\percent} to \SI{100.0}{\percent} of prints match a settled
taker leg, with exact side and token agreement on all 24.5~million matched pairs. Mint-and-merge
settlement makes that taker leg essential: pooling maker and taker legs changes the measured
buy share. Signing the full cached tape before settlement selection yields 16.6 million prints
signed by every rule; equal-day balanced accuracy is 0.942 for Lee--Ready, 0.768 for the tick
test and 0.670 for retrospective bulk volume classification. On 16.4 million identical eligible
fills, Lee--Ready raises effective spread by 0.429 cents per share under equal-fill weights;
its realised-spread difference is -0.209 cents. Under share-volume weights, taker five-minute
midpoint impact is 0.689 cents, while tick and bulk classifications give -0.374 and -0.119
cents. That aggregate sign reversal disappears when tied print rows are excluded. Distortion
depends on error-weighted signed outcomes, sample selection and weighting. Receipt-time ordering
and unobserved future-quote age limit these delivered-quote accounting quantities; they do not
identify causal impact or private information. Supporting venue and collector analyses provide descriptive diagnostics.
\end{abstract}
\noindent\textbf{Keywords:} prediction markets, market microstructure, trade direction,
trade classification, trading costs, blockchain data
\noindent\textbf{JEL:} G14, G23, D47
\section{Introduction}
Signed spread-based measures of trading costs require knowing which side initiated a trade.
When the side is inferred
from prices and quotes, a small classification error can have a large economic consequence if
errors occur where the trade price lies far from the midpoint. We study this problem on
Polymarket, a prediction market whose orders are matched off-chain and whose trades settle
on-chain. Each public trade print carries a transaction hash. This permits a comparison of its
reported side with the taker order of its own settlement, and permits signed trading-cost
measures to be recomputed against that observed taker side.
Trade-classification studies have long assessed both accuracy and the consequences of errors
\citep{ellis2000accuracy,odderswhite2000occurrence,chakrabarty-2015-bvc,jurkatis-2022-fia}.
\citet{qin-yang-2026-database} also study chain-based direction and downstream proxy bias on
Polymarket V1. A buyer of one outcome can settle against a buyer of its complement through a
mint. Maker and taker legs therefore do not provide interchangeable truth labels, and a pooled
leg benchmark can misstate both the buy share and agreement with the feed. Receipt-time quotes
add a second problem: an accurately reported taker side need not agree with the side inferred
from the last quote delivered to the collector.
These two measurement problems motivate public-feed validation linked to delivered quotes
across both exchange generations, with an explicit complete-set direction convention and
paired accounting of signed-cost distortion.
Our companion study \citep{dubach-2026-anatomy} documented the venue's microstructure across
52 days and 385{,}198 markets using an archive without trade prints. Reconstructing direction
from resting-size decrements recovered the on-chain sign on only about \SI{59}{\percent} of
matched five-second buckets. It left open whether the market channel's
\texttt{last\_trade\_price} event identifies the aggressor. The second archive format studied
here contains that event, and the sample spans the 28 April 2026 migration from CLOB~V1 to
CLOB~V2. The transaction-hash join answers the question without pairing feed and settlement
observations solely by time.
The empirical argument has three parts. First, we validate the unit and coverage of the
reported sign. On twelve selected days between April and August, including two V1-exchange days,
\SI{92.1}{\percent} to \SI{100.0}{\percent} of prints match a settled taker leg under their
transaction hash. Side and token agree on every one of the 24.5~million matched pairs. We also
measure prints without settlements and settlements without prints, so exact agreement on the
matched observations is not confused with completeness of the tape.
Second, we construct signs on every cached print before selecting those with settlement truth,
then compare rules on the same observations. On the 16.6~million prints signed by all four
rules, equal-day balanced accuracy is 0.942 for Lee--Ready, 0.768 for the tick test and 0.670
for the retrospective BVC bar-majority diagnostic. Lee--Ready leads on every sampled date,
but its score ranges from 0.901 to 0.959. Restoring unmatched prints to the history changes
the scores little; rule-specific eligibility changes the comparison more. This distinction
prevents a scoring-population effect from being mistaken for a construction improvement.
A quote strictly before the print is used because message order within a receipt millisecond
is unresolved.
Third, we connect sign errors to economic measurement on identical eligible fills, in cents
of collateral per share. On 16.4 million common observations, Lee--Ready raises effective
spread by 0.429 cents under equal-fill weights, while reducing realised spread by 0.209 cents.
Under share-volume weights, tick and BVC reverse the sign of the average five-minute midpoint
impact relative to the taker; that aggregate reversal disappears when tied scored prints are
removed. The distortion is not determined by accuracy alone: it depends on the error-weighted
signed outcome, and differs with the accounting quantity, weights and eligible population.
The comparison therefore reports baseline levels, paired differences and sensitivity rather
than treating a classification score as proof of an economically reliable measure.
The broader archive supplies context and diagnostics rather than additional main
contributions. Appendix~\ref{app:venue-analysis} retains the migration, tick-change and
quoted-liquidity analyses. The migration moved every market at once and has no contemporaneous
control group; tick changes generally follow prices near the extremes; and the World Cup
comparison contains few independent event groups. Narrower spreads are consistent with a
binding grid, while game progress and the grid's contribution remain unresolved. A
listings-based health rule (Appendix~\ref{app:health}) and the two-way feed--settlement
coverage audit address collector gaps that can otherwise resemble venue events.
The main results concern measurement on twelve observed days in one venue. They do not establish
causal market-quality effects or population accuracy beyond the sampled days.
\section{Related literature}
\paragraph{Prediction markets and calibration.}
The price-level literature establishes that prediction-market prices are informative and
reasonably calibrated when expiry is near, with a longshot bias for events far in the future
\citep{wolfers-zitzewitz-2004,page-clemen-2013}, though a price need not equal the traders' mean
belief \citep{manski-2006}. These price-level questions leave a separate measurement problem:
the cost of taking an informational position depends on the trade side and the quotes used to
measure that cost.
\paragraph{Decentralised prediction-market microstructure.}
\citet{rahman-al-chami-clark-2025} survey the domain and frame its open questions without
bringing matching tick-level data. \citet{tsang-yang-2026-anatomy} study the 2024 US presidential
election markets (chiefly the Trump YES contract) over many time slices and document
episode-level patterns, with $\lambda$ estimated on aggregated rather than tick-level data.
\citet{sirolly-2025-wash} find that a quarter of the venue's lifetime volume looks like wash
trading, a practice that \citet{cong-li-tang-yang-2023-wash} quantify on unregulated
cryptocurrency exchanges. \citet{ng-2025-discovery} compare price discovery across venues, finding that
Polymarket leads around the 2024 presidential election and in politics and entertainment, while
Kalshi leads in sports and when its relative trading activity is higher; within Polymarket,
\citet{gomezcram-2026-skilled} attribute most price discovery to a few percent of accounts.
\citet{xi-moallemi-pai-wang-2026-volatility} model forward-looking volatility in
prediction markets as a function of probability, time to resolution, spreads and activity, and
\citet{dai-jia-yu-2026-settlement} show that a short-duration market settling on an asset price
that its holders can move transfers wealth from liquidity traders to manipulators. On Polymarket's
BTC 15-minute binaries, \citet{young-2026-openmarket} finds that a walk-forward logistic model over
43 microstructure features with Binance order flow does not beat, and slightly underperforms, the
probability already implied by Polymarket's own book, and that simulated trading on it loses a
little. This paper measures the trade sign and makes no forecast with it.
The election-market evidence of \citet{tsang-yang-2026-anatomy} is complemented by the
broader cross-section in our companion study \citep{dubach-2026-anatomy}:
within its 52-day, 385{,}198-market sample, 600 markets were observed simultaneously at tick level,
with microstructure computed on the full event tape joined to the on-chain record.
\citet{qin-yang-2026-database} release the first-generation exchange's on-chain trade record
with aggressor labels and study classification bias in VPIN, trade-based order-flow measures
and forecasting relationships. They deliberately omit off-chain quote flows. Their tick-rule
and BVC raw accuracies, \SI{49.83}{\percent} and \SI{50.51}{\percent}, apply to a filtered
Standard Binary sample spanning the V1 lifecycle, with routing records and specified market
types excluded. Their log-index execution records and YES-axis normalization differ from our
transaction-hash-matched, asset-local taker target. In our settlement-filtered own-rule benchmark on the two
V1 days, tick raw accuracy is 0.704 and 0.710; period, unit, sign convention and sample selection
differ, so these data do not reconcile that gap. Our distinct comparison validates public-feed
prints and pairs their delivered quotes with settlement direction across both exchange
generations to measure signed-cost distortion. We do not claim the first chain-based target
or the first economic-bias analysis on this venue.
\citet{shen-2026-ghost} document settlements that revert on-chain after the venue reported the
trade; we meet the same transactions from the feed's side.
\paragraph{Trade classification.}
In equity data, \citet{ellis2000accuracy} and \citet{odderswhite2000occurrence} find that
the quote-and-tick algorithm of \citet{lee-ready-1991} misclassifies a material share of
trades, concentrated inside the spread and in fast markets.
\citet{chakrabarty2012short} show that it classifies short sales about as accurately as long sales
and that its trade-level accuracy depends on how quotes are timed against trades, and
\citet{chakrabarty-2015-bvc} compare it with bulk volume classification and the tick rule.
\citet{easley2016discerning} argue that bulk classification can be preferable when individual signs
are noisy, and \citet{jurkatis-2022-fia} halves Lee--Ready's error rate at one-second timestamps.
On a centralised cryptocurrency exchange, \citet{ma-zhai-2021-tick} score the tick rule against the
exchange's own side flag at \SI{77}{\percent}. The
order-flow imbalance of \citet{cont-kukanov-stoikov-2014}, which has a strong relationship with
contemporaneous price changes, and the queue imbalance that \citet{gould-bonart-2016} study as a
one-tick-ahead predictor are built from the best quotes and the sizes queued at them and need no
trade sign; the trade imbalance that Cont, Kukanov and Stoikov compare with order-flow imbalance,
Kyle's $\lambda$ and the signed spreads this paper reports do. We use the recorded side of each matched settlement's taker order as the target, and validate
the feed against it by transaction hash. The public chain supplies an auditable target for this
venue, while the mint-and-merge mechanism requires a distinction between takers and the other
legs of the same transaction. Our economic comparison asks how the errors alter the signed
price distances on the same fills, rather than treating a classification score as sufficient
evidence about the reliability of a trading-cost measure. When one side dominates, raw accuracy is a
misleading score and the Matthews correlation a more reliable one \citep{chicco2020advantages}; we
compare the rules on balanced accuracy and report the Matthews correlation beside it. Against the
taker side, the feed's print carries the sign exactly, and Lee--Ready on the quote strictly before
the print recovers it with balanced accuracy of 0.942 on the common scored sample.
\section{Institutional background}
\subsection{Polymarket and the first-generation venue}
Polymarket runs a central limit order book (CLOB) for binary markets: an operator matches orders
off-chain, and trades settle on Polygon in a stablecoin against an on-chain conditional-token
contract. Each market has two tokens, one per outcome, which this paper calls assets. A complete
set, one token of each outcome, redeems for one unit of collateral; it can be minted from
collateral and merged back into it. Until April 2026 the venue operated its first-generation
exchange (``CLOB~V1''): a standard exchange contract and a neg-risk exchange contract, through
which the venue's neg-risk markets settle. Orders were signed under an EIP-712 domain (the Ethereum
standard for signing typed structured data) whose version was \texttt{1} and which embedded a
\texttt{feeRateBps} field in the signed order. The public market channel broadcasts book
snapshots, price-level changes and trade prints. The archive this paper uses kept no prints in its
first format, so in that format the aggressor is not identified, which is the origin of the
measurement problem this paper closes.
\subsection{The CLOB V2 migration}
\label{sec:migration}
On 28 April 2026 the venue shipped a coordinated upgrade: new exchange contracts, a rewritten
matching backend, a new collateral token and a new fee model. Its changelog of 17 April announced
the upgrade for about 11:00 UTC on 28 April, with about an hour of downtime, and told traders to
plan for all open orders to be wiped at the cutover; its entry of 28 April records that resting
orders from before the cutover did not migrate, and the books reopened without them. Two of the changes matter for reading the chain and
the feed.
\paragraph{New exchange contracts.} The V2 exchange addresses differ from V1, and the on-chain
event ABI was rewritten. The fill event gained fields (most importantly an explicit \texttt{uint8}
side), and the \texttt{OrdersMatched} and \texttt{FeeCharged} events, which V1 also emitted,
changed signature. A replication that filters the V2 contracts' logs on the V1 fill event's topic,
the hash of its signature, returns zero legs; Appendix~\ref{app:decoding} gives both signatures.
\paragraph{A new fee model.} V1 embedded the fee in the signed order. V2 collects fees on-chain at
match time, with category-specific taker rates and a fee-free Geopolitics category, and carries
the fee on the fill event (Appendix~\ref{app:decoding}). The feed's \texttt{fee\_rate\_bps} field
is a V1 artefact and is not the V2 fee. This paper does not analyse fee payments; fee rates enter only as
descriptors of the venue's categories in Section~\ref{sec:cutover}.
\section{Data}
\label{sec:data}
\subsection{The feed archive}
The order-book data are a tick-level archive of Polymarket's public WebSocket market channel,
captured by a third party, PMXT \citep{pmxt-archive-v2}, in two formats: the provider's first format from 21 February to 15
April 2026 and its second format from 15 April to 10 August, 171 UTC days in all with the seam day
present in both. The archive's seam precedes the venue's migration, so the last days of the V1
exchange are second-format days. Each event carries a receipt time and a native timestamp, market
and asset identifiers, an event type and, for price-changing messages, the best bid and best ask on
every row. All feed timestamps are the archive's receipt time, not the venue's native time, which
the feed re-sends on some rows of the second format and which is not usable as a partition clock.
The second format adds the trade prints with their transaction hashes, the tick-size change
messages and the book snapshots the feed sends on subscription.
Because best quotes travel on every \texttt{price\_change} row, we recover a per-asset quoted
spread series without replaying the order book. A quote's spread is the ask minus the bid. A
one-sided book, quoted on one side only, arrives with the missing side at 0 or 1, so its spread is
one minus the lone bid or the lone ask itself. A quote counts as empty when its spread is 0.9 or
more, which covers a book with no side and the venue's placeholder (a midpoint of 0.5 with a spread
of 0.9 or more), and as live otherwise; an asset-day is live when its median spread is above zero
and below 0.9.
From both formats we build one panel per (market, asset, UTC day): update counts; the median and
quartiles of the spread over the day's updates, exact to 0.001 (each is the lower edge of the 0.001
bin that holds that quantile's update); spread sums and squares; the last midpoint; the hours in
which the asset was present and the number of updates in which its book was empty; displayed size
at the best bid and ask from the \texttt{price\_change} rows at the best level and from the
snapshots; the first quote, first live quote and first trade of the day; and, for the second-format
days, trade counts and volume. The panel holds 17{,}009{,}350 asset-days on 3{,}081{,}897 assets of
1{,}540{,}951 markets; 93{,}588 assets appear in both formats. Each asset's market comes from the
feed's own market field and its outcome from the venue's metadata, the outcome labels the first
format carries and a registry, in that order, and an outcome is known for a share of 0.9993 of asset-days.
\subsection{On-chain ground truth}
\label{sec:onchain}
Ground truth comes from the events of the standard and neg-risk exchange contracts of each
generation on Polygon, read for twelve whole UTC days: 26 and 27 April on the V1 exchange, 29 and
30 April, the first two full days of the V2 exchange, and eight days stratified over May to August
(13 and 31 May, 5 and 9 June, 1, 7 and 17 July, 5 August), chosen to span the sample, to bracket
the 2 July and 10 July venue changes and to avoid the archive's June and July outages
(Section~\ref{sec:health}). Each day's block range is fixed by block timestamps and extended 300
blocks (seven and a half to ten minutes) past midnight so that the settlements of the day's last
prints are captured. We keep every \texttt{OrderFilled} leg of these contracts, 4.6 to 7.2~million
a day, and every \texttt{OrdersMatched} and \texttt{FeeCharged} event, each with its block
timestamp. The cutover day itself, 28 April, is a boundary and is excluded. Only the V2 fill event
carries a side field; on V1 the side follows from which asset the order's maker gave, so both
exchanges yield a signed taker leg per settlement (Section~\ref{sec:method-match}).
\subsection{Summary statistics}
\begin{table}[htbp]
\centering\footnotesize\setlength{\tabcolsep}{3.5pt}
\begin{tabular}{lrrrrr}
\toprule
& mean & p25 & median & p75 & p95 \\
\midrule
\multicolumn{6}{l}{\textit{Panel A: the panel, 171 days over both archive formats, per asset-day}} \\
Binned median quoted spread, live asset-days & 0.194 & 0.010 & 0.038 & 0.260 & 0.850 \\
Asset-days, live / all & \multicolumn{5}{c}{13{,}451{,}616 / 17{,}009{,}350} \\
Asset-days by format, first / second & \multicolumn{5}{c}{2{,}855{,}452 / 14{,}153{,}898} \\
Updates per asset-day, median & \multicolumn{5}{c}{2{,}163} \\
Assets observed per day, minimum and maximum over clean days & \multicolumn{5}{c}{26{,}928--302{,}256} \\
Markets listed by the venue per day, median February--July / August & \multicolumn{5}{c}{8{,}759 / 23{,}547} \\
\midrule
\multicolumn{6}{l}{\textit{Panel B: the on-chain sample, 12 days}} \\
Feed prints & \multicolumn{5}{c}{25{,}035{,}601} \\
Settled taker legs & \multicolumn{5}{c}{27{,}144{,}192} \\
Prints matched to their taker leg by transaction hash & \multicolumn{5}{c}{24{,}537{,}912} \\
\bottomrule
\end{tabular}
\caption{Summary statistics. Panel A is the harmonised panel of both archive formats; `live' marks live days, asset-days whose binned median spread is above zero and below 0.9; clean days exclude the degraded days of Appendix~\ref{app:health}. Panel B is the on-chain sample of Section~\ref{sec:identity}.}
\label{tab:summary}
\end{table}
Table~\ref{tab:summary} summarises the panel and the on-chain sample; Section~\ref{sec:composition}
and Figure~\ref{fig:series} describe how the cross-section moves over the sample.
\subsection{Archive health}
\label{sec:health}
The feed archive is a third party's capture, and its gaps look like market events unless they are
flagged. We benchmark it against the venue's own listings, from a sweep of the venue's metadata API
by listing day, and a rule flags a day as degraded when an hourly file is missing or near-empty,
when the archive first sees almost no new assets while the venue keeps listing markets, or when it
carried fewer than half of the markets listed that day. Appendix~\ref{app:health} gives the rule,
the benchmark and the list. Over the whole sample the archive carried \SI{81.3}{\percent} of the
markets the venue listed (1{,}511{,}091 of 1{,}858{,}106) on the listing day or later; on the
median clean day it carried \SI{99.98}{\percent}, and on the tenth-percentile clean day
\SI{93.6}{\percent}. Thirty-six of the 171 days are degraded, and the rule's first-seen,
carried-share and window thresholds move that count by at most three days
(Appendix~\ref{app:health}).
\section{Methodology}
\label{sec:method}
\subsection{Matching prints to settlements}
\label{sec:method-match}
Every \texttt{last\_trade\_price} event the archive's second format records carries the hash of the
Polygon transaction meant to settle the trade it reports. Each settled transaction on the
V2 exchange contracts emits one \texttt{OrderFilled} leg per order it fills, and exactly one of those
legs has the exchange contract itself in the \texttt{taker} field: the taker order. We join each
print to that leg by transaction hash. On the first-generation exchange the fill event carries no
side, but the side follows from which asset the order's maker gave: collateral for a buy, the
outcome token for a sell. The join is therefore the same on both exchanges. For each matched pair
we compare the side, the token, the price and the size, where the on-chain price and size follow
from the two filled amounts and the side.
The residual runs in both directions. On the feed's side, a print with no settled counterpart is
matched in a second stage to the unmatched taker leg of the same token and side, with the same
price to four decimals and the same size to six, that lies nearest in time within two minutes. An
hour-stratified sample of the prints that remain is classified from the transaction receipt: no
transaction under the hash, reverted, mined outside the scraped block range, or settled by a
contract we did not read; the class shares are weighted by each hour's population, and a sample of
matched hashes is run through the same lookup as a control. On the chain's side, a taker leg with
no print is a fill the archive does not contain; the scrape's tail past midnight is excluded from
that count.
\paragraph{What a settlement looks like.} The exchange settles a taker order in one of three ways,
and the legs show which. A \emph{complementary} match fills the taker against makers on the same
token with the opposite side. A \emph{mint} fills a taker buying one outcome against makers buying
the other, and creates the pair of tokens from collateral; a \emph{merge} is its mirror for sells.
We classify each transaction from its legs; one whose maker legs include both kinds, some on the
taker's token and some on its complement, is \emph{mixed}.
\subsection{Direction rules and their scoring}
\label{sec:method-rules}
The primary benchmark constructs signs on every print in each cached UTC daily tape before
joining to settlement truth. It therefore preserves unmatched prints in the tick history and
five-second bars. Only hash-matched prints can be scored. This separates an implementable
feed-based history from the settlement selection needed for validation; it does not recover
prints missing from the archive. For comparison, a settlement-filtered construction signs only the matched
prints. Both constructions reset at the start of the day.
Lee--Ready signs a print above the prevailing midpoint as a buy, below it as a sell, and at
the midpoint by the tick test. The quote rule signs a print at or above the ask as a buy and
at or below the bid as a sell, using the tick test inside the spread. Missing or crossed quotes
invoke the tick fallback. The tick test compares the price with the asset's previous print and
carries the last non-zero sign through zero ticks; observations before its first non-zero
change are unsigned. Bulk volume classification (BVC) assigns each five-second bar a buy share
from its last-price change scaled by a twenty-bar rolling standard deviation
\citep{easley2016discerning}, then assigns that majority sign to the prints in the bar. An
unchanged or first bar is unsigned. BVC estimates a volume fraction rather than individual
initiators. Its closing price and whole-day fallback volatility are unavailable at each print;
this is a retrospective bar-majority diagnostic, not a streaming classifier. Exact definitions
and the paired history comparison are in Appendix~\ref{app:rules}.
The prevailing quote is the asset's last \texttt{price\_change} strictly before the print's
receipt millisecond. Quotes stamped at that same millisecond are excluded because their order
relative to the print is unresolved. Print histories use the saved cache row order to break
asset--timestamp ties deterministically; that order is not established venue order. A
sensitivity removes tied print rows from scoring while retaining their history. It tests their
influence on the reported scores, but cannot eliminate their possible influence on subsequent
signs. Cached quotes at lags from 1~ms to 1~s permit a further timing comparison.
For each rule we report raw accuracy, balanced accuracy (the mean buy and sell recall), the
Matthews correlation and coverage on its signed subset. The primary ranking additionally uses
the intersection signed by all four rules, so every rule faces identical truth observations
and class composition. The history comparison pairs full- and selected-history predictions on
identical cache row ids, and compares their scores only where both assign a sign. It isolates
construction changes from abstention changes. Across days we report equal-day means, ranges
and observed paired differences. Twelve dates chosen to span venue changes are not random or
necessarily independent; these contrasts describe the sample without a population-significance
claim.
\paragraph{Cell-based aggregation.} Within each asset's five-second cells, we compare feed
direction with all on-chain legs pooled together and with taker legs separately. This
comparison assesses the consequences of pooling settlement legs. Feed cells use receipt times
and settlement cells use block times, so cell disagreement also reflects the clocks.
\subsection{Economic consequences on identical eligible fills}
\label{sec:method-signed}
Let $p_i$ denote the print price, $m_i$ the preceding delivered midpoint, $m'_i$ the cached
midpoint five minutes later and $s_i$ the taker sign. For any sign $a_i$, the accounting quantities
are effective spread $E_i(a)=2a_i(p_i-m_i)$, realised spread $R_i(a)=2a_i(p_i-m'_i)$ and
five-minute midpoint impact $I_i(a)=2a_i(m'_i-m_i)$. Hence $E_i(a)=R_i(a)+I_i(a)$ on every
observation. Results are reported in cents of collateral per share, with both equal-fill and share-volume
weights. They are full-spread quantities; fees, rebates and inventory costs are excluded.
All four rules and the taker are compared on identical observations, using full-tape signs.
Eligibility requires all four signs, a finite print price in probability support, an uncrossed
strictly interior current quote ($0<b_i\leq u_i<1$, with $u_i$ the ask), a finite forward midpoint in probability
support and a positive finite print size. Prints whose five-minute horizon reaches or passes
the exclusive UTC day boundary are excluded. A sequential funnel records each exclusion. The future quote's
timestamp, bid and ask were not saved, so a finite future midpoint does not establish its age,
its two-sided state or an uncrossed future book. The day-end exclusion addresses a known
boundary, not internal gaps or stale quotes. We also restrict current quote age to at most one
second and remove tied print observations as sensitivity checks.
For rule $r$, let $e_i=\mathbf{1}\{r_i\ne s_i\}$. For any of the three measures $M$,
\[
M_i(r)-M_i(s)=-2e_i M_i(s),\qquad
\overline M_w(r)-\overline M_w(s)=-2\sum_i w_i e_iM_i(s),
\]
where the non-negative weights sum to one on the common sample. Thus error frequency alone
cannot determine distortion: it is the error-weighted signed outcome that matters. For valid
uncrossed current quotes, off-midpoint Lee--Ready errors are wrong-side fills and increase
effective spread mechanically. No analogous direction of distortion follows for realised
spread or midpoint impact. The code verifies both identities for every sign and weighting.
\begin{quote}\small
\noindent\textbf{Observed signing error.} On 17 July 2026 at 15:36:43.018 UTC, a cached
print reports a buy of 115.07 shares at 0.22. Its own settlement confirms the same buy, token,
price and size. The preceding delivered bid and ask, received 160~ms earlier, are 0.22 and
0.24, giving a midpoint of 0.23. Lee--Ready therefore signs the print as a sell. Effective
spread is $-2$ cents per share under the taker sign and $+2$ cents under Lee--Ready: the
sign error adds 4 cents on this fill. This negative taker-signed distance is relative to the
delivered quote; it does not establish the execution-time venue book or a profitable trade.
Appendix~\ref{app:decoding} identifies the transaction and its settlement legs.
\end{quote}
We report signed differences in price units rather than ratios near a zero denominator, and whether the
sign of a day's mean changes. Exploratory per-asset comparisons use the same eligible assets
with at least fifty fills to describe sign changes and rank correlations; they do not treat
assets sharing a market or event as independent draws. Means with absolute value at most
$10^{-12}$ price units count as numerical zero for sign diagnostics; asset means are rounded to
twelve decimal places only for ranking. These precision conventions leave accounting means
unchanged and are not economic materiality thresholds.
These are accounting comparisons against delivered quotes. Receipt-time ordering does not
identify the execution-time venue book, and a five-minute signed midpoint change is not an
identified causal price impact or measure of private information. Realised spread here is
not a fee-adjusted profitability estimate.
\subsection{Inference}
The primary classifier and economic comparisons use the day as the summary unit
(Section~\ref{sec:method-rules}); no uncertainty interval is pooled over millions of fills.
The twelve days were selected to span the sample and venue changes rather than sampled at
random. Their mean scores, ranges and paired contrasts describe these observed days; they do
not provide population inference over all trading days. Inference for the
supporting venue analyses is stated in Appendix~\ref{app:venue-analysis}.
\section{Results}
\label{sec:results}
\subsection{The print is the taker leg}
\label{sec:identity}
\begin{table}[htbp]
\centering\scriptsize\setlength{\tabcolsep}{3pt}
\begin{tabular}{lrrrrrrrr}
\toprule
\multicolumn{9}{l}{\emph{Panel A. Coverage and identity}} \\
day & prints & with leg (\%) & legs in day & with print (\%) & side $\neq$ & token $\neq$ & price (\%) & size (\%) \\
\midrule
26 Apr (V1) & 2{,}383{,}232 & 95.9 & 2{,}399{,}429 & 95.3 & 0 & 0 & 92.3 & 99.96 \\
27 Apr (V1) & 2{,}522{,}118 & 96.8 & 2{,}534{,}187 & 96.4 & 0 & 0 & 92.4 & 99.95 \\
29 Apr & 2{,}169{,}975 & 93.9 & 2{,}150{,}324 & 94.8 & 0 & 0 & 92.5 & 99.93 \\
30 Apr & 2{,}237{,}594 & 92.1 & 2{,}178{,}955 & 94.6 & 0 & 0 & 92.1 & 99.95 \\
13 May & 2{,}251{,}602 & 99.9 & 2{,}271{,}729 & 99.0 & 0 & 0 & 92.0 & 99.96 \\
31 May & 1{,}905{,}025 & 100.0 & 1{,}969{,}244 & 96.6 & 0 & 0 & 91.8 & 99.96 \\
5 Jun & 1{,}705{,}254 & 99.9 & 2{,}349{,}031 & 72.3 & 0 & 0 & 93.0 & 99.97 \\
9 Jun & 2{,}254{,}797 & 99.8 & 2{,}385{,}448 & 94.0 & 0 & 0 & 93.7 & 99.97 \\
1 Jul & 1{,}833{,}608 & 100.0 & 2{,}508{,}976 & 72.6 & 0 & 0 & 95.1 & 99.91 \\
7 Jul & 2{,}242{,}392 & 100.0 & 2{,}547{,}188 & 88.0 & 0 & 0 & 94.4 & 99.90 \\
17 Jul & 1{,}876{,}816 & 100.0 & 1{,}977{,}096 & 94.9 & 0 & 0 & 93.4 & 99.96 \\
5 Aug & 1{,}653{,}188 & 100.0 & 1{,}704{,}195 & 97.0 & 0 & 0 & 92.7 & 99.96 \\
\bottomrule
\end{tabular}
\par\medskip
\begin{tabular}{lrrrrrrr}
\toprule
\multicolumn{8}{l}{\emph{Panel B. Settlement types and buy share}} \\
day & mint (\%) & compl. (\%) & merge (\%) & mixed (\%) & unknown (\%) & buy share, taker & buy share, all legs \\
\midrule
26 Apr (V1) & 67.8 & 21.9 & 3.5 & 6.8 & 0.0 & 0.812 & 0.848 \\
27 Apr (V1) & 67.6 & 21.5 & 3.4 & 7.4 & 0.0 & 0.806 & 0.849 \\
29 Apr & 67.6 & 22.3 & 2.8 & 7.3 & 0.0 & 0.813 & 0.850 \\
30 Apr & 66.9 & 21.8 & 3.3 & 8.1 & 0.0 & 0.807 & 0.847 \\
13 May & 68.8 & 20.8 & 3.2 & 7.2 & 0.0 & 0.819 & 0.852 \\
31 May & 68.6 & 20.3 & 3.7 & 7.4 & 0.0 & 0.825 & 0.854 \\
5 Jun & 72.2 & 19.0 & 3.0 & 5.8 & 0.1 & 0.838 & 0.866 \\
9 Jun & 72.9 & 18.3 & 2.7 & 6.0 & 0.0 & 0.848 & 0.870 \\
1 Jul & 53.5 & 20.5 & 2.6 & 4.2 & 19.2 & 0.832 & 0.861 \\
7 Jul & 62.8 & 21.4 & 2.6 & 5.4 & 7.9 & 0.824 & 0.858 \\
17 Jul & 73.8 & 17.9 & 2.1 & 5.2 & 0.9 & 0.856 & 0.884 \\
5 Aug & 64.9 & 27.1 & 1.8 & 6.0 & 0.1 & 0.811 & 0.847 \\
\bottomrule
\end{tabular}
\par\medskip
\begin{tabular}{lrrrrrrr}
\toprule
\multicolumn{8}{l}{\emph{Panel C. The residual}} \\
day & dup. prints & unmatched prints & 2nd-stage matched & reverted (\%) & no tx (\%) & other (\%) & receipts \\
\midrule
26 Apr (V1) & 1 & 96{,}602 & 63{,}068 & 20.1 & 79.9 & 0.0 & 960 \\
27 Apr (V1) & 2 & 80{,}016 & 45{,}288 & 32.0 & 68.0 & 0.0 & 960 \\
29 Apr & 52 & 132{,}425 & 50{,}502 & 41.1 & 58.9 & 0.0 & 960 \\
30 Apr & 2 & 177{,}137 & 75{,}201 & 21.2 & 78.8 & 0.0 & 960 \\
13 May & 39 & 2{,}070 & 7 & 100.0 & 0.0 & 0.0 & 888 \\
31 May & 1{,}285 & 813 & 158 & 94.2 & 5.8 & 0.0 & 625 \\
5 Jun & 4{,}880 & 1{,}511 & 613 & 87.8 & 12.2 & 0.0 & 674 \\
9 Jun & 9{,}326 & 4{,}247 & 682 & 17.5 & 6.1 & 76.4 & 588 \\
1 Jul & 11{,}641 & 610 & 21 & 99.2 & 0.0 & 0.8 & 491 \\
7 Jul & 0 & 776 & 33 & 99.9 & 0.1 & 0.0 & 577 \\
17 Jul & 0 & 810 & 22 & 100.0 & 0.0 & 0.0 & 669 \\
5 Aug & 0 & 672 & 305 & 73.2 & 26.8 & 0.0 & 329 \\
\bottomrule
\end{tabular}
\caption{Feed prints against on-chain taker legs, joined by transaction hash. Panel~A: the share of prints with a settled taker leg; the taker legs settled inside the day (the scrape's 300-block tail past midnight excluded) and the share of them with a print; the matched pairs on which the side or the token differ; the shares of matched pairs within $5\times10^{-4}$ in price and $10^{-6}$ in size. Panel~B: settlement types as shares of settled transactions, read from the legs' structure; \emph{unknown} is a transaction with a maker leg on a token absent from the market metadata; the buy share over taker legs and over all legs. Panel~C: prints that repeat a hash; the prints the hash join leaves without a settled leg; those matched in a second stage on token, side, price and size within two minutes; and the receipt classes (\emph{reverted}; \emph{no tx}, no transaction under the hash; \emph{other}, on these days always a transaction mined outside the scraped block range) of an hour-stratified sample of the rest, weighted by each hour's population. \emph{receipts} is the number of hashes in the sample.}
\label{tab:fillmatch}
\end{table}
On the eight days from May to August, \SI{99.8}{\percent} to \SI{100.0}{\percent} of prints match
a settled taker leg under their own transaction hash; on the four April days, which straddle the
migration, \SI{92.1}{\percent} to \SI{96.8}{\percent} do (Table~\ref{tab:fillmatch}). On the
24.5~million matched pairs the side and the token never differ. The size agrees to $10^{-6}$ on
\SI{99.9}{\percent} of pairs. The price agrees to $5\times10^{-4}$ on \SI{91.8}{\percent} to
\SI{95.1}{\percent} and to half a cent on every pair: the maximum absolute difference is 0.005 on
each of the twelve days, which is what rounding to the cent produces.
Each print is compared with the taker leg of its own transaction; the maker legs are not in the
comparison. The test has power because a print copied from a maker leg would carry the other token
in a mint (\SI{54}{\percent} to \SI{74}{\percent} of settlements) and the opposite side in a
complementary match, and the size would not agree. The agreement is complete within each settlement
type separately. The sample is the matched prints, so on the April days \SI{92}{\percent} to
\SI{97}{\percent} of prints are inside the test. On the two days before the migration the join runs
against the first-generation exchange, whose fill event has no side field, and with the side derived
from the legs the agreement is the same. The feed therefore gives the aggressor's side; it carries
no identities.
\paragraph{Settlements without a print.} Among settlements inside the day, the share with a print
ranges from \SI{72.3}{\percent} to \SI{99.0}{\percent}, and is at least \SI{88}{\percent} on ten of
the twelve days. Part of the gap is the hour after midnight: on 13 May, \SI{22}{\percent} of the
unmatched settlements are in the first hour, before the collector had subscribed to the day's new
markets. The two days at \SI{72}{\percent} to \SI{73}{\percent}, 5 June and 1 July, are days on
which the archive lacked whole markets or delivered late prints. On 1 July, 485{,}746 settlements
involve a maker token the archive never carried, \SI{19.2}{\percent} of the day's transactions; on
5 June the missing settlements are on known markets, but the prints arrive a median 22~seconds after
their block, where on other days they precede it, and 4{,}880 prints repeat a hash. Both look like
collector faults rather than the venue's, although delayed prints and repeated hashes could arise on
either side; Section~\ref{sec:health} and Appendix~\ref{app:health} give the archive's health by
day.
\paragraph{Prints without a settlement.} In April, 80{,}016 to 177{,}137 prints a day have no
settled leg under their hash. Of those, 45{,}288 to 75{,}201 match an unmatched settlement of the
same token, side, price and size within two minutes, and for the remainder a receipt lookup returns
no transaction at all on \SI{59}{\percent} to \SI{80}{\percent} of an hour-weighted sample, a
reverted transaction on the rest. From May on the residual is 610 to 4{,}247 prints a day, reverted
on \SI{73}{\percent} to \SI{100}{\percent} of the sample on every day but 9 June, where
\SI{76}{\percent} were mined outside the day's block range. A control of 100 matched hashes per day
returns a settled transaction on the exchange inside the range for all 100 on each of the ten days
on the second-generation exchange; on the two first-generation exchange days the same control
returns all 100 matched hashes as sent to a contract other than the exchanges, so there the lookup's
other-contract class carries no information. The feed prints a trade when the exchange accepts it
and the chain records whether it settled; the reverted class is the ghost fill of
\citet{shen-2026-ghost}. Whether the April prints without a receipt were resubmitted under another
hash we cannot show from receipts alone. The second-stage match is consistent with resubmission and
is all the evidence we have, but the match carries no false-match rate: on a venue where many trades
are identical in token, side, price and size, some of those two-minute matches will be coincidences,
so the counts are upper bounds on the genuine matches.
\subsection{What direction means on a mint-and-merge exchange}
\label{sec:mintmerge}
By count, the exchange settles most taker orders by minting or merging complete sets rather than
against a resting order on the same token: \SI{53.5}{\percent} to \SI{73.8}{\percent} of settled
transactions are mints, \SI{17.9}{\percent} to \SI{27.1}{\percent} complementary matches,
\SI{1.8}{\percent} to \SI{3.7}{\percent} merges and \SI{4.2}{\percent} to \SI{8.1}{\percent} mixed.
Up to \SI{19.2}{\percent} of transactions on one day cannot be classified because a maker leg is on
a token absent from the market metadata; every such case is a token the archive never carried, and
the share is under \SI{0.1}{\percent} on eight of the twelve days. In a mint every leg is a buy and
in a merge every leg is a sell: the maker legs carry the taker's side on the complement token.
Pooling legs by token therefore counts a mint's makers as buyers of the other outcome and
overstates the buy share. The buy share of taker legs is \SI{80.6}{\percent} to
\SI{85.6}{\percent}; over all legs it is \SI{84.7}{\percent} to \SI{88.4}{\percent}.
Pooling every settlement leg in an asset-by-five-second cell and comparing the pooled sign
with the feed's gives agreement on \SI{86.8}{\percent} to \SI{90.9}{\percent} of cells against
pooled legs, below a majority baseline of 0.906 to 0.947 (its balanced accuracy is 0.741 to
0.802, against a chance level of 0.5); against taker legs it agrees on \SI{93.1}{\percent} to
\SI{97.8}{\percent} of cells, and fill by fill the feed's side agrees on \SI{100}{\percent}. The cells
that still disagree against taker legs are not evidence about the feed's side either: the feed side is
timed by receipt and the chain side by block, and a print and its own settlement often fall in
different cells.
\subsection{Classifier comparisons on the same prints}
\label{sec:rules}
\begin{table}[htbp]
\centering\footnotesize\setlength{\tabcolsep}{3.5pt}
\begin{tabular}{lrrrrr}
\toprule
rule & own BA & own cover (\%) & common BA & common accuracy & common MCC \\
\midrule
Lee--Ready & 0.954 & 100.0 & 0.942 & 0.937 & 0.809 \\
quote rule & 0.953 & 99.8 & 0.941 & 0.932 & 0.799 \\
tick test & 0.735 & 96.0 & 0.768 & 0.708 & 0.409 \\
BVC, 5 s bars & 0.670 & 67.8 & 0.670 & 0.652 & 0.260 \\
\bottomrule
\end{tabular}
\caption{Full cached daily histories, scored against hash-matched taker sides. Cells are equal-day means over twelve selected dates. Own-sample scores condition on each rule assigning a sign; own coverage is the percentage of matched prints signed. Common scores use the same intersection of four assigned signs on each day, 16{,}615{,}817 prints in total, covering 57.7--72.2\% of matched prints by day. BVC is a retrospective five-second bar-majority diagnostic. Means and ranges describe selected days; no population test is implied.}
\label{tab:cached-benchmark}
\end{table}
The full daily caches contain 25.0~million prints, of which 24.5~million match settlement
truth. The intersection signed by all four rules contains 16.6~million prints, covering
\SI{57.7}{\percent} to \SI{72.2}{\percent} of matched prints by day. On this identical
sample, equal-day mean balanced accuracy is 0.942 for Lee--Ready, 0.941 for the quote rule,
0.768 for the tick test and 0.670 for the retrospective BVC majority-bar diagnostic
(Table~\ref{tab:cached-benchmark}). Lee--Ready's day range is 0.901 to 0.959, with the lowest
score on 5 August. Its advantage over the tick test is 0.174 on average (0.150 to 0.190 across
days), and over BVC 0.272 (0.207 to 0.309); both differences are positive on all twelve dates.
The quote rule trails Lee--Ready on each date. These are sample descriptions, not tests over
independent randomly selected trading days.
Figure~\ref{fig:cached-direction} places these daily common-sample scores beside the paired
history contrasts discussed below.
\begin{figure}[htbp]\centering
\includegraphics[width=\linewidth]{figures/fig_cached_direction.pdf}
\caption{Cached-tape comparisons on twelve selected dates. Panel~(a) scores every rule on the
same four-rule signed intersection for each day. Panel~(b) holds the scored rows fixed within
each rule and compares full versus settlement-selected sign histories; the vertical scale is
percentage points of balanced accuracy. Connecting segments guide the eye. BVC is a
retrospective bar-majority diagnostic. No interval or interpolation over unobserved days is
implied.}
\label{fig:cached-direction}
\end{figure}
\paragraph{Eligibility matters more than history in this sample.} On each rule's own signed
sample, full-history balanced accuracy is 0.954 for Lee--Ready and 0.735 for the tick test.
The tick test's higher score of 0.768 on the common sample is therefore a scoring-population
change, not an improvement caused by restoring unmatched prints. On identical observations
signed under both histories, the equal-day full-minus-selected change is only 0.000192 for
the tick test and 0.000080 for BVC. The largest daily paired sign-change share is
\SI{0.349}{\percent} for the tick test and \SI{0.377}{\percent} for BVC, and smaller for
the quote-based rules (Table~\ref{tab:cached-history}). Full-tape construction is necessary
for a feed-based benchmark, but its empirical effect here is small. Daily denominators and
truth composition are in Table~\ref{tab:cached-days}.
\paragraph{Timestamp ambiguity and quote timing.} Asset--timestamp print ties affect
\SI{36.9}{\percent} of matched prints. Removing those rows from scoring, while retaining
their history, leaves 9.3~million common observations and raises Lee--Ready's equal-day
balanced accuracy to 0.957. That is a substantial sample change, and cannot establish the
unknown within-tie order or remove its effect on later signs. Keeping one print per matched
hash changes Lee--Ready's common score only to 0.943. On the separate fixed sample where
Lee--Ready and all cached lag variants sign, its balanced accuracy is 0.9538 with the preceding
quote, 0.953 with a 100~ms lag, 0.921 with a 500~ms lag and 0.885 with a one-second lag.
Those figures have a different denominator from the primary four-rule intersection. Quote
lags reveal sensitivity to delivered-quote timing; they do not recover the execution-time
venue book.
\subsection{Trading-cost distortion and the five-minute decomposition}
\label{sec:economic-results}
The accounting sample contains 16{,}448{,}261 prints signed by every rule with valid current
quotes, finite outcomes and sizes, and a five-minute horizon strictly inside the cached day.
It covers \SI{67.03}{\percent} of settlement matches; daily exclusions and fresh-quote counts
appear in Table~\ref{tab:cached-economic-eligibility}. This further selected sample is shared
by all signs and both weightings. Each reported summary averages twelve daily means equally.
Table~\ref{tab:cached-economic-means} gives the levels, and
Table~\ref{tab:cached-economic-distortion} gives paired rule-minus-taker differences.
\begin{table}[htbp]
\centering\footnotesize
\begin{tabular}{lrrr}
\toprule
\multicolumn{4}{l}{\emph{Equal-fill weights}} \\
sign & effective & realised & 5-min impact \\
\midrule
taker & 2.328 & 6.486 & $-4.158$ \\
Lee--Ready & 2.757 & 6.277 & $-3.520$ \\
quote rule & 2.662 & 6.169 & $-3.507$ \\
tick test & 1.475 & 5.279 & $-3.804$ \\
BVC, 5 s bars & 1.204 & 4.692 & $-3.488$ \\
\bottomrule
\end{tabular}
\par\medskip
\begin{tabular}{lrrr}
\toprule
\multicolumn{4}{l}{\emph{Share-volume weights}} \\
sign & effective & realised & 5-min impact \\
\midrule
taker & 2.033 & 1.344 & 0.689 \\
Lee--Ready & 2.432 & 1.708 & 0.724 \\
quote rule & 2.340 & 1.642 & 0.698 \\
tick test & 1.545 & 1.919 & $-0.374$ \\
BVC, 5 s bars & 1.208 & 1.327 & $-0.119$ \\
\bottomrule
\end{tabular}
\caption{Signed accounting quantities in cents of collateral per share, on the identical primary eligible fills for every rule and the taker. Cells average the twelve daily means equally. Each daily equal-fill or share-volume weighting is shared across all signs. Effective equals realised plus five-minute midpoint impact before rounding. The quote is delivered to the collector; these are not fee-adjusted profits or identified causal impact/information estimates.}
\label{tab:cached-economic-means}
\end{table}
\begin{table}[htbp]
\centering\footnotesize
\begin{tabular}{lrrrrrr}
\toprule
\multicolumn{7}{l}{\emph{Equal-fill weights}} \\
rule & $\Delta E$ & $\Delta R$ & $\Delta I$ & flip E & flip R & flip I \\
\midrule
Lee--Ready & 0.429 & $-0.209$ & 0.638 & 0 & 0 & 0 \\
quote rule & 0.334 & $-0.317$ & 0.650 & 0 & 0 & 0 \\
tick test & $-0.853$ & $-1.207$ & 0.353 & 0 & 0 & 0 \\
BVC, 5 s bars & $-1.124$ & $-1.793$ & 0.669 & 0 & 0 & 0 \\
\bottomrule
\end{tabular}
\par\medskip
\begin{tabular}{lrrrrrr}
\toprule
\multicolumn{7}{l}{\emph{Share-volume weights}} \\
rule & $\Delta E$ & $\Delta R$ & $\Delta I$ & flip E & flip R & flip I \\
\midrule
Lee--Ready & 0.399 & 0.365 & 0.035 & 0 & 0 & 0 \\
quote rule & 0.307 & 0.298 & 0.009 & 0 & 0 & 0 \\
tick test & $-0.487$ & 0.575 & $-1.063$ & 0 & 0 & 8 \\
BVC, 5 s bars & $-0.824$ & $-0.016$ & $-0.808$ & 0 & 0 & 6 \\
\bottomrule
\end{tabular}
\caption{Rule minus taker signed means in cents of collateral per share, with the same rows and weights; equal-day means of paired differences. E, R and I denote effective spread, realised spread and five-minute midpoint impact. Flip columns count dates with opposite daily mean signs (out of twelve); absolute means at or below $10^{-12}$ price units are treated as numerical zero. Differences preserve direction rather than taking absolute values. They measure distortion relative to delivered-quote accounting and do not identify changes in market behavior.}
\label{tab:cached-economic-distortion}
\end{table}
Under equal-fill weights, the taker effective spread is 2.328 cents of collateral per share,
realised spread is 6.486 cents and five-minute midpoint impact is $-4.158$ cents. Realised
spread exceeds effective spread because the taker-signed delivered midpoint subsequently moves
in the opposite direction on average. Lee--Ready adds 0.429 cents to effective spread, removes
0.209 cents from realised spread and adds 0.638 cents to midpoint impact. Its effective-spread
increase occurs on every sampled date; the realised-spread difference is negative on eleven.
All rules and the taker retain the same daily mean signs under this weighting. A high score
therefore coexists with a cost distortion, without a daily aggregate sign reversal.
Weighting changes the economic result. With printed-share-volume weights, taker effective
spread is 2.033 cents, realised spread is 1.344 cents and midpoint impact is positive at
0.689 cents. Lee--Ready adds 0.399 cents to effective spread and 0.365 cents to realised
spread, leaving impact close to the taker at 0.724 cents. Tick and BVC instead give negative
mean impact, $-0.374$ and $-0.119$ cents, with paired differences of $-1.063$ and $-0.808$
cents. Their daily impact means have opposite signs to the taker on eight and six dates,
respectively; Lee--Ready and the quote rule have none. This is a statement about a delivered-quote
accounting measure. Neither its positive taker value nor a classifier's negative value identifies
informed trading or causal price impact.
\begin{figure}[htbp]
\centering
\includegraphics[width=\textwidth]{figures/fig_cached_economics.pdf}
\caption{Paired rule-minus-taker distortions in cents of collateral per share on identical primary eligible fills.
Markers are equal-day means; circles use equal-fill weights and diamonds use printed-share-volume weights.
Whiskers span the observed daily minimum and maximum, not confidence intervals. Signing rules and
weights change the three accounting quantities differently; the twelve selected dates do not
support a population-significance interpretation.}
\label{fig:cached-economics}
\end{figure}
The effect is outcome-dependent, rather than a universal attenuation factor
(Figure~\ref{fig:cached-economics}). Lee--Ready's equal-fill effective-spread distortion is
entirely upward; its error-weighted signed effective distance is negative. For realised spread,
positive and negative per-fill differences mostly cancel: the average absolute paired distance
is 6.786 cents although the signed difference between means is only $-0.209$ cents. Accuracy
and absolute per-fill error cannot replace the signed error-mass identity in
Section~\ref{sec:method-signed}. In particular, the effective-spread result does not imply a
similar realised-spread or midpoint-impact distortion under another weighting.
\paragraph{Sensitivity and asset comparisons.}
Table~\ref{tab:cached-economic-sensitivity} reports the volume-weighted impact comparison.
Restricting the current quote to an age of at most one second retains negative tick/BVC
aggregate impact and gives nine and seven daily opposite signs, but Lee--Ready then reverses
one daily mean. One eligible row per hash preserves the primary eight/six counts. Substituting
the on-chain price leaves impact unchanged by construction and changes Lee--Ready's equal-fill
effective distortion only from 0.429 to 0.425 cents.
Removing tied scored prints retains 9.1 million observations and changes the aggregate finding:
taker, tick and BVC volume-weighted impact are all positive, at 0.955, 0.102 and 0.181 cents.
Daily tick/BVC reversals fall to five and two. Thus the primary aggregate sign reversal is
sensitive to this selection. This check retains tied observations in the signing history;
it neither recovers venue order nor establishes which scoring population is preferable.
Exploratory ranks use 49{,}817 asset-day observations with at least fifty common fills.
Table~\ref{tab:cached-economic-ranks} reports daily rank correlations and opposite-sign shares.
Under equal-fill weights, Lee--Ready's realised-spread rank correlation with the taker averages
0.982, compared with 0.612 for tick and 0.519 for BVC. Their asset-level realised-spread means
have opposite signs in \SI{3.31}{\percent}, \SI{21.73}{\percent} and \SI{24.61}{\percent}
of the eligible assets on an average day. The primary absence of aggregate daily sign reversals
therefore does not establish preservation of every asset's economic ordering. These dependent,
selected asset comparisons are descriptive rather than tests on independent market draws.
\section{Discussion}
The exact side and token agreement in Table~\ref{tab:fillmatch} establishes what the feed reports on matched settlements:
the print corresponds to the settling taker order. Researchers retaining this side field can
construct signed measures directly. An archive without trade prices cannot implement the
price-based rules benchmarked here, and an archive without the side field must infer direction.
The distinction is operational: validation of an existing field and reconstruction of a missing
one solve different data problems.
The effective-spread comparison shows why accuracy is an incomplete criterion for choosing a
trade-classification rule. A wrong sign changes $2s(p-m)$ by $-4s(p-m)$, so errors on fills with
large price distances have larger consequences. For valid uncrossed quotes, Lee--Ready turns
off-midpoint negative taker-signed distances into positive ones by construction. The resulting
upward distortion need not resemble the attenuation produced by independent sign noise.
The paired differences in Table~\ref{tab:cached-economic-distortion} measure this distortion
on the same fills as the taker baseline. Its magnitude belongs to this venue's quotes, settlement mechanism and receipt clock;
classification performance on another venue need not imply the same economic distortion.
The binary settlement mechanism also matters for comparisons with previous evidence. Most
settlements in this sample are mints, in which the maker and taker buy opposite outcome tokens.
Counting every leg as a separate direction observation changes the question being asked.
Likewise, scoring each rule only when it assigns a sign changes the eligible population, and
constructing histories after a settlement filter changes the inputs to tick-based rules.
Figure~\ref{fig:cached-direction} separates the common-sample ranking from the paired
history comparison. Actual history effects are small on these cached days, while the common
sample changes the tick test's balanced accuracy materially.
The primary ranking therefore uses one population, while separate own-rule coverage and
paired history comparisons distinguish eligibility from construction.
Receipt-time ordering sets a limit on the economic interpretation. A taker buy below the last
delivered midpoint need not mean that the taker received a favorable execution against the
book at the venue. The quote may have been delivered out of sequence relative to the print;
same-millisecond ties cannot be ordered from these data. The lag sensitivity is informative
about this uncertainty but does not recover the venue clock. Signed costs computed here are
relative to observed delivered quotes. The negative equal-fill and positive share-volume
taker midpoint-impact averages describe different weightings of this tape. Forward quote age
and future bid/ask are unobserved; a finite cached midpoint and a within-day horizon do not
validate the future book. These responses identify neither private information nor causal
price impact. Removing tied scored prints changes the primary aggregate tick/BVC sign reversal,
so that conclusion also depends on the observed scoring population.
The supporting liquidity analyses show further limits to comparisons on this archive
(Appendix~\ref{app:venue-analysis}). The migration has no unaffected market control, and
displayed size is not comparable across it. Ordinary tick changes depend on extreme prices
near resolution, so their event studies describe the lifecycle. The World Cup arms contain
only six and three treated event groups; their descriptive narrowing cannot separate game
progress from the tick. Finally, a health rule based on new listings cannot detect every loss
of older markets that are still trading. Feed--settlement coverage is therefore a separate
check on completeness rather than a substitute for the listings benchmark.
The twelve on-chain days were chosen to span the archive and venue changes. They establish
exact sign agreement within the matched sample and day-specific coverage, with substantial
variation in classifier performance and signed-cost distortion. They are not a probability
sample of trading days. Dependence among contracts of the same market or event and the small
number of days limit inference beyond the observed sample. The results should be read as a
replicable measurement study of one venue and collector, with explicit boundaries on what the
truth labels and quotes observe.
\section{Conclusion}
Polymarket's public trade prints report the settling taker order's side and token on all
24.5~million matched pairs in twelve selected days. Transaction-hash matching and a separate
audit of missing prints and settlements make that result both precise and conditional.
Mint-and-merge settlement explains why pooled chain legs are an inappropriate substitute for
the taker target.
The full-tape common benchmark gives Lee--Ready balanced accuracy of 0.942, but accuracy does
not determine signed-cost distortion. On identical eligible fills, it adds 0.429 cents per
share to effective spread under equal-fill weights. Its realised-spread effect differs in sign
between equal-fill and share-volume weights. Tick and BVC can reverse the sign of volume-weighted
five-minute midpoint impact relative to the taker, although this aggregate reversal disappears
when tied scored rows are excluded. The relevant mechanism is the error-weighted signed price
distance, rather than error frequency alone. Validated signs, common populations, weights and
quote timing are necessary to interpret these accounting quantities. The supporting venue
and archive analyses retain descriptive evidence while leaving their causal roles unresolved.
\clearpage
\section*{Declarations}
\paragraph{Data availability.} The on-chain events (\texttt{OrderFilled}, \texttt{OrdersMatched}
and \texttt{FeeCharged} of the second-generation exchange contracts and their first-generation
counterparts) are public on Polygon and are re-read by the accompanying scraper from any archive
node; the decoded ABIs and the block ranges of the twelve days are versioned. The order-book feed
is a third-party capture of Polymarket's public WebSocket market channel (PMXT, \url{https://archive.pmxt.dev})
in its first and second formats, 1.8~TB of hourly Parquet files with the provider's receipt files;
it is not redistributed here. The derived inputs (the asset-day panel, the tick-change contexts,
the feed prints and on-chain events, and the auxiliary tick-event panel) are publicly deposited with
an archived code snapshot in the replication package \citep{dubach-2026-replication}. The manifest hashes the
derived files and reuses the downloader's SHA-256 receipts for the archive's second-format
hours. First-format hours have no receipts and are recorded by filename, size and modification
time rather than a content hash. Market metadata (categories,
fee schedules, tick sizes, listing and resolution dates) is a snapshot of the venue's public Gamma
API taken on 26 September 2026 and included as raw pages with a manifest.
\paragraph{Code availability.} The analysis is a set of scripts in the repository directory
\texttt{docs/\allowbreak paper-b/\allowbreak v2-revision/}, one per result, each writing a JSON
file from which the tables and figures are generated. A claims map links numeric statements
to results or declared external facts and design choices. The manifest identifies the inputs
and hashes the derived files, result JSONs and scripts; the environment is pinned. Development
provenance and archived specifications are documented separately in the replication notes.
The analysis code, result files and derived inputs are archived in Zenodo version~1.4
at \url{https://doi.org/10.5281/zenodo.23164515}. This release includes the full-tape
common-sample and three-measure analyses reported here, together with the reproduction of
the observed print illustrated in Section~\ref{sec:method-signed}.
\paragraph{Use of automated assistance.} Large language models were used as coding and writing
assistants for the analysis code, the text, the figures and reviews of drafts.
\paragraph{Funding and competing interests.} The author received no funding for this work and
declares no competing interests.
\clearpage
\bibliographystyle{plainnat}
\bibliography{refs}
\clearpage