How we measure
Every estimate is tested against real auction sales the model never saw. This page sets out how.
What the range means
Every estimate is a range, and 9 out of 10 sales land inside it. We measured that on 74,824 sales the model never trained on, Aug 31–Sep 5, 2026. It's a rate across many estimates, not a guarantee for any one vehicle.
Reports also mark a narrower inner range around the estimate. We make no frequency claim for it.
If we change how either range is defined, this page says what changed.
The protocol
- Forward-in-time holdout. The model is trained on sales up to a cutoff date. It is then evaluated on every qualifying sale that happened after that cutoff — lots it never trained on, predicted as they would have been before the sale.
- No look-ahead. The comparable sales the model retrieves for each prediction are restricted to sales that occurred strictly before that lot's own sale date.
- All lots in scope, not a curated subset. Every realized sale above a $100 floor in the window is included.
- Named baselines, same lots. The model is measured next to (a) a naive median of the same year/make/model and (b) the auction house's own per-lot estimate, evaluated on identical lots.
- Every number carries its n and its evaluation window. No exceptions.
What we don't do
- We don't hand-pick examples.
- We don't quote accuracy measured on sales the model trained on.
- We don't quote a figure without its sample size and dates.
- We don't republish raw auction listings.