Blog
How the dataset is built, measured and validated — racing data engineering in the open. No tips, no predictions; just the work. Follow along via RSS.
Putting every horse in the context of its race
7 September 2026
We had 180 point-in-time features describing each runner — and not one that said how a horse stood against the field it actually faced. The context release adds 44: field-relative standing, the gap to the official rating, and a sample-size denominator beside every rate. Here's what we built, and what a full-archive analysis made us change before shipping.
Why building an honest racing rating is harder than it looks
9 August 2026
Our data supplier withdrew the ratings we displayed, so we built our own — and audited, measured and corrected everything on the way. Fabricated dead-heats, a data source that measured itself, a regression broken by the handicapper doing their job, and a rating that inflated until we anchored it. The full story behind VPR.
How we measure racecourses from open geodata
31 July 2026
Track centrelines from OpenStreetMap, terrain from Copernicus EU-DEM, turns from curvature — the complete method behind the open course-topology dataset, including every course we can't yet measure and why.
