How the number is made
Every rating on this site is a weighted sum of named factors, converted to a probability by a curve fitted on real outcomes, and checked against what actually happened. This page is the whole of it — the weights, the calibration, the sources, and the parts that did not work.
Model axiom-v1.0.0
1 · The factors and their weights
The home-run rating is the weighted average below. These are read from the engine at render time, not transcribed — a guard holds the scorer, this table and the documentation to the same numbers.
The composite ONYX score blends the three markets: 45% HR · 35% TB · 20% HIT.
2 · Turning a rating into a probability
A 0–100 rating is not a probability. Each market has a logistic curve fitted on the graded record, re-fitted every Monday and kept only when the fresh fit still beats its own base-rate baseline.
p = 1 / (1 + e−(-3.7645 + 3.5345 × score/100))p = 1 / (1 + e−(-0.2302 + 1.3682 × score/100))p = 1 / (1 + e−(-1.2891 + 1.4400 × score/100))Measured against the graded record since 1 July, the home-run rating separates outcomes at an AUC of 0.654 across 1,043 hitter-games. Hitters in the top tenth of the rating homered 22.1% of the time against 6.7% in the bottom tenth — 3.3×.
An AUC of 0.500 is a coin flip and 1.000 is perfect. Numbers in this range are normal for a rare event and are worth stating plainly rather than dressing up: the rating ranks meaningfully, and it is nowhere near an oracle.
The full reliability diagram — predicted against actual, bucketed, with the sample behind each point — is on the track record, computed at render time so it cannot go stale.
3 · The tiers
Cut from the graded record, not chosen. Each band is re-derived before it is changed.
4 · Where the data comes from, and when
Ratings are written before first pitch and re-scored once confirmed lineups post. Results are graded overnight from official boxscores.
5 · Known limitations
Everything below is a measured result that did not go our way. It is published for the same reason the misses are: a model you cannot check is a model you should not trust.
Backtested against outcomes, the intervals on pitcher vulnerability, weather, pitch matchup and bullpen exposure all contain 0.500 — a coin flip. Together they carry roughly a fifth of the home-run rating. An independent rebuild measured weather at 0.508 and park at 0.519 on the same record. Nothing has been re-weighted, because re-weighting on an in-sample comparison is how a model gets worse while looking better, but the reader should know which parts of the number are doing the work.
Crossing a pitcher’s actual pitch usage against a hitter’s performance by pitch type is the single most-cited edge in this category. Built properly and measured, it scores 0.487 — below a coin flip. Two independent measurements agree. It is shown on the player card as a read, never as a projection, and no scorer weights the decomposition.
In the highest band the model has predicted about 28% and delivered about 19%. Three candidate corrections were tested and none cleared the bar to adopt, so none was applied. The weekly recalibration structurally cannot catch this: it is a two-parameter curve, and two parameters set a level and a slope — they cannot bend.
We publish a fair price — the break-even odds implied by our probability — but we do not currently ingest sportsbook lines. That means no line movement, no closing-line value, and no vig-removed comparison against the market. When a book price appears anywhere on this site, you typed it in. Removing vig honestly requires both sides of a market; a one-sided price cannot be de-vigged, and any site that claims otherwise is estimating.
No umpire strike-zone profiles, no defensive positioning or outs-above-average, no catcher framing, and no reliever-level bullpen usage. Strikeout props in particular would need several of those to be modelled honestly, so we do not pretend to model them.
Recent form is computed from game logs rather than from stored 7-, 15- and 30-day windows, which do not yet exist. Form is real but coarser than it should be.