Every grade and probability ONYX publishes, checked against what actually happened. The percentages either mean what they say or they don't — this is where you find out. We don't hide the misses.
The ten bats we ranked highest for a home run, and what each one did. Hits and misses at the same size — the top of a board is a probability, not a prediction, and the nights it looks worst are the ones worth publishing.
One slate is one slate — ten bats decide nothing, which is what every chart below this is for. It is here because a rate you cannot check against a game you watched is just a number we typed.
When the model forecast a given HR chance, how often it actually happened. The closer predicted and actual are in each band, the more the percentage on a player's card is worth.
Each dot is one probability band: what the model said against what happened, sized by how many player-games it rests on. The dashed line is perfect calibration — a dot above it beat its own number, below it fell short. Hover any dot for its sample.
Every number above is an average over the whole record, and an average is what a model in trouble hides behind — an overall gap of a point or two can be built from a fortnight too high and a fortnight too low. This is the same figure per slate: what we forecast, and what actually happened.
Hitters we graded EDGE or better to homer — sorted by longest fair-value odds. The model reaching for the dart, and hitting.
The mirror of the catches. These are the plays the model was most confident in that still came up empty — sorted by how loudly we said it. At an average of 31%, missing is the expected outcome roughly 69% of the time. That’s the honest shape of a home-run model, and it’s why we publish this next to the wins.
Each market graded by its own score. Rate should step down from PRIME to COLD every column — that's the model earning its keep.
Every model, scored on its own. AUC is the chance this model rated a hitter who homered above one who didn’t — 0.500 is a coin flip. Measured on the graded record, against a 11.8% base rate. Some of these numbers are not flattering, which is the point of publishing them.
Read this carefully before concluding a model is dead weight. These are standalone numbers. The models are correlated, and one can still earn its weight by sharpening the others in combination even when it ranks poorly alone. Park and Weather also barely vary within a single night, which mechanically depresses their score here. We tested that claim rather than just asserting it: a weight backtest across 26 slates and 6,207 graded player-games found that Power alone — the one model with a strong standalone number — scores the best of any variant on the first half of the data and the worst of any variant on the held-out half (0.6102 then 0.5791, against the live blend’s 0.5874). The weak-looking components are doing real work in combination. Re-weighting to chase these standalone numbers gained 0.0018 AUC out of sample, which is inside the noise, so the live weights stand.
The model’s own top 10% — every play it graded 64 or better — split by what actually happened. These plays homered 19.2% of the time against a 11.8% field, so the grade is doing real work. It is also wrong 81% of the time, and that is the part nobody explains. 232 homered, 976 did not — same grade, opposite night. So what were the nine models each saying in the two piles?
Not one of the nine cleared the bar. On the record so far, the plays that cashed and the plays that missed look alike on every input we have — the biggest gap in the table is smaller than its own error bar.
Which is worth saying plainly, because the instinct it kills is a real one: there is no second filter to apply to the top of this board. Backing only the top plays whose weather looks best, or whose opposing arm looks softest, has not sorted the winners from the losers — those reads are already inside the grade, and past it they are noise. The grade is the read.
This is a description, not a filter. It is measured on the same graded rows it describes, so a gap found here is what already happened rather than a rule that will keep holding — which is exactly why the table leads with z and greys everything under 2 rather than ranking by the gap itself. And it is not an argument for re-weighting. Separation inside a narrow, self-selected band is a different quantity from the standalone predictive power in the table above; POWER in particular is close to constant here by construction, because a high Power grade is most of how a play reached this band at all. A model can be the best in the product and still be no help choosing between its own best plays.
One unit on each of the top 5 HR grades every night, settled at our own fair price. Flat stakes — any sizing plan is a second model, and layering one on would make this a statement about the staking rather than the record.
This is not a profit claim, and it is not closing line value. Nothing here knows what any book offered or closed at — there is no odds feed. Betting a perfectly calibrated model at its own fair price returns exactly zero by construction, so this measures the gap between what the top of the board promised and what it delivered, expressed in units. A real market’s margin comes off the top of whatever is left. It is also a small sample: 250 bets over 50 slates, with a 64-unit worst drawdown along the way. Judge the direction, not the decimal.
Raw log — HR grade and fair-value price going in, result coming out. Green = cashed. No edits.
Each graded slate: bats scored, forecast HR rate, what actually happened.