Skip to main content
MLB Simulator logo

About MLB Deserve-to-Win

Baseball is the most over-counted sport in the world and the most under-explained. This site does the opposite: a simulator replays every game to estimate what should have happened, and every number is anchored against the league — so you can read at a glance whether what you just watched was normal or remarkable.

How deserve-to-win works

After every game we ask one question: did the right team win? A 110 mph line drive gets caught; a weak grounder finds a hole. We strip that randomness out in four steps.

Every batted ballexit velo · launch angle
Outcome oddsout · single · double · HR
10,000 replaysre-roll every ball
Deserve-to-win %share of replays won
  1. Collect every plate appearance

    We pull each plate appearance from the game — every batted ball's exit velocity, launch angle, and spray angle, plus the walks, strikeouts, hit-by-pitches, and stolen bases that never become contact.

  2. Estimate what each ball should become

    A model trained on millions of historical batted balls turns that contact into probabilities: how often a ball hit that hard, at that angle, in that direction becomes an out, a single, a double, or a home run.

  3. Simulate the game 10,000 times

    We replay the game ten thousand times, re-rolling each batted ball against those probabilities while keeping the walks, strikeouts, and baserunning fixed. A 110 mph line drive that got caught now falls for a hit in most of the reruns.

  4. Read off a win probability

    The share of those 10,000 reruns each team wins is its simulated win probability. 0.62 means “given how the bats sounded, you should have won 62 out of 100 games like this one.”

Stack those per-game numbers across a season and you get expected wins (xW) — the win total the schedule should have produced. Actual wins minus xW is a team’s luck differential, charted on every team page. Mid-season gaps tend to shrink: a team running five wins above xW in June usually gives some back, and one trailing xW often makes some up.

Why our numbers aren't just averages

A hitter who goes 6-for-10 in the season's first week is not a .600 hitter. Ten trips to the plate is mostly noise — so instead of taking small samples at face value, the site's player numbers come from a hierarchical Bayesian model. That's a formal name for a simple discipline: every player's estimate starts near league average, then moves toward what that player is actually doing as the evidence piles up. Ten plate appearances barely move it; four hundred mostly speak for themselves.

what happenedthe model’s estimate
league average
10 PA100 PA400 PA
The same hot streak, three sample sizes. The model pulls a 10-PA streak most of the way back to league average — and mostly believes a 400-PA one.

This is also why the K% and BB% shown on pitcher pages don't exactly match the raw rates on other stat sites: they are the model's estimate of the pitcher's true rate, not a running tally of what has happened so far. Early in a season the two can differ noticeably; by August they nearly agree.

The headline numbers, written out

  • Hitter productionEB/PA = (estimated bases on contact + walks + HBP) ÷ plate appearances

    Walks and hit-by-pitches count as one base; strikeouts count as zero. Higher is better.

  • Pitcher run preventionxEB/PA = contact share × EB per ball in play + BB% + HBP%

    Same units as EB/PA, built from separately-modeled strikeout, walk, and contact-quality estimates. Lower is better.

  • Expected winsxW = sum of each game's simulated win probability

    A 0.62 deserve-to-win game adds 0.62 wins to the tally — whether the team actually won or lost that night.

  • Luck differentialLuck = actual wins − xW

    Positive: a team has banked more wins than its play earned. Negative: it's owed a few.

See estimated bases livePick an exit velocity and launch angle in the landing explorer and watch where those balls actually land.Try it →

What’s on the site

7 hubs, all built on the same simulated numbers.

Metrics glossary

Every metric on the site, in plain language — with the direction that counts as “better” spelled out.

Estimated bases (EB)
How many bases a batted ball should have been worth
The building block of the whole site. Given a ball's exit velocity, launch angle, and direction, the model estimates its average value in bases — a scorched liner might be worth 1.4 EB even if it's caught, a bloop single only 0.3 even though it fell in.
EB/PA
Estimated bases per plate appearance
A hitter's average bases per trip to the plate: each batted ball's estimated bases, plus one base for every walk and hit-by-pitch, zero for strikeouts, divided by plate appearances. Higher is better. It credits the quality of contact, not where the ball happened to land. Rankings show the model's estimate of this rate rather than the raw tally, so a hot week in a small sample doesn't leapfrog a full season of work.
xEB/PA
Expected bases allowed by a pitcher, per PA
The pitcher's version of EB/PA — same units, same scale, but lower is better. The 'x' marks how it's built: the model estimates a pitcher's strikeout, walk, and hit-by-pitch rates and the contact quality they allow separately, then combines them into expected bases per plate appearance. That composite predicts a pitcher's future results better than simply tallying what happened.
Bases created
A hitter's total estimated bases in a game, contact plus walks
The game-page view of offense: each batted ball's estimated bases stacked with the walks and hit-by-pitches that put the hitter on for free. Higher is better — it credits the quality of contact, not where the ball happened to land.
Bases prevented
How many bases a pitcher suppressed versus an average outing
The pitcher's mirror image of bases created: how many fewer estimated bases the pitcher allowed than a typical pitcher would have over the same batters faced. Higher is better — it rewards weak contact and strikeouts, not lucky defense. Measured against an average pitcher — a stricter bar than the replacement-level baseline used on the best/worst boards.
Excess bases
Bases above or below a replacement-level baseline
Used on the Best/Worst performance boards: a player's estimated bases minus what a readily available fill-in would have produced in the same opportunities. Big positive numbers need both quality and volume — a great night in six trips beats a great night in three. A friendlier bar than the average-player baseline used in the pitch-by-pitch sections: replacement asks “better than a fill-in?”, not “better than typical?”
Batted-ball luck
Actual bases minus deserved bases, per ball or per player
The game-page luck ledger: what a ball actually earned (out = 0, single = 1, ... home run = 4) minus its estimated bases. Positive means the hitter got more than the contact deserved; negative means a well-struck ball died in a glove. Distinct from a team's season-level luck differential.
Upset score
How surprising a result was, given who deserved to win
Combines how strongly the simulator favored the loser with the gap in team quality. A juggernaut out-hitting a rebuilder and still losing scores high; a coin-flip game going either way barely registers.
Deserve gap
How far apart the two teams were on Deserve-To-Win
The distance between the two teams' Deserve-To-Win percentages in a single game — the column of the same name on the Games tab. Near zero means the simulator saw a coin flip whatever the scoreboard said; a wide gap means one side clearly outplayed the other. It says nothing about who actually won: a wide gap in the loser's favor is exactly what marks a game an upset.
Times through the order
How a pitcher fares the 1st, 2nd, and 3rd time facing a lineup
Hitters improve with each look at the same pitcher. The pitcher-page section shows strikeouts, walks, and estimated bases allowed by trip through the lineup — a steep dropoff the third time through is the classic argument for an earlier hook.
BB%
Walk rate per plate appearance
What share of a hitter's at-bats end in a walk. For hitters higher is better (patience and pitch selection). For pitchers it's the opposite — walks given up.
K%
Strikeout rate per plate appearance
What share of plate appearances end in a strikeout. For hitters lower is better (more balls in play). For pitchers higher is better (more dominant stuff).
HR%
Home runs per plate appearance
What share of plate appearances end in a home run. Rises when bat tech, ball construction, or hitter approach favor power; the league-wide rate is a leading indicator of the era's offensive character.
R/G
Runs scored per team per game
The single most legible 'how high-scoring is the era' number. Combines power, contact, and on-base into one rate — the team-level scoreboard, averaged across the league.
Hard-hit %
Share of batted balls hit ≥ 95 mph exit velocity
A power-of-contact stat. Hard-hit balls fall in for hits more often than soft contact, so the percentage is a leading indicator that survives small-sample noise.
FB%
Fly-ball rate — share of batted balls with launch angle 25–50°
How often a hitter (or league) lifts the ball into the air on the productive part of the launch-angle curve. Fly balls turn into doubles, triples, and home runs; ground balls almost never do. League FB% sits around 23% in the Statcast era.
GB%
Ground-ball rate — share of batted balls with launch angle ≤ 10°
How often contact stays on the ground. High-GB% pitchers limit damage by suppressing extra-base hits. League GB% sits around 45%.
Barrel%
Share of batted balls with elite exit velocity + launch angle combo
A 'barrel' is a Statcast-defined sweet spot: at least 98 mph exit velocity and a launch angle that widens as the ball is hit harder. Barreled balls are hits about 80% of the time and homers about half. League Barrel% sits around 8%.
Pull%
Share of batted balls hit to the batter's pull side
A righty pulls the ball to left field, a lefty to right field. Pulled contact is where most home-run power lives, but an extreme pull habit also invites defensive shading. Neither high nor low is 'better' — it describes a hitter's shape, not his quality.
Bat speed
How fast the sweet spot of the bat is moving at contact, in mph.
Measured on competitive swings only — checked swings and bunts are excluded, so a bunt attempt doesn't drag a hitter's number down. Faster isn't better: bat speed trades power against contact, and Luis Arraez is elite while sitting near the bottom of the scale. League-wide tracking starts in 2024, so earlier seasons show no value.
Swing length
How far the sweet spot of the bat travels during the swing, in feet.
A longer swing can build more bat speed but takes more time to get through the zone, so it's a swing-style tradeoff, not a quality score — the longest and shortest swings in the league both belong to good hitters. Measured on competitive swings only, same as bat speed.
Attack angle
The bat's upward or downward tilt through the zone, in degrees.
A steeper attack angle matches a pitch on a downward plane and can help drive the ball into the air; a flatter one keeps the bat in the hitting zone longer. Neither is better on its own — it's a swing-style choice, not a quality score.
BB+HBP%
Walks + hit-by-pitch per plate appearance
The free-base rate: walks the hitter earns plus times he gets plunked. Slightly higher than pure BB% because HBP is folded in, but useful for daily rolling views where separating the two adds noise without insight.
Luck differential
Actual wins minus expected wins
The gap between a team's real win-loss record and how many wins the underlying play would predict. Positive means a team has banked more wins than they've earned; negative means the opposite.
Win %
Share of games won
The simplest 'are they good?' number. Everything else on the Standings page is context for why this number is what it is — luck, schedule strength, run differential, or genuine team quality.
Run differential
Runs scored minus runs allowed
A team can be 24-20 by winning blowouts and losing squeakers (high diff) or vice versa. The gap usually predicts second-half record better than win % alone does — it's the cleanest 'are these wins real?' signal that fans can read directly.
Home/Road split
Record at home versus on the road
Most teams play noticeably better at home. The size of the gap tells you whether a team is travel-fragile or has a real road profile that should hold up in the playoffs.
Playoff probability
Share of simulated seasons where this team makes the playoffs
The simulator runs the rest of the season thousands of times with each team's true quality drawn from its posterior. Playoff probability is how often this team gets in across those simulated seasons.
Projected wins
Most likely final win total, with a 90% range
The simulator's best guess at a team's final win total, plus the range it lands in 90% of the time. A 90-win team with a [82–98] band is less certain than a 90-win team with an [87–93] band — the band width is the uncertainty.
Expected wins
Win total the simulator says a team should have
For each game, the simulator estimates a win probability from the contact quality both teams produced. Expected wins (xW) is the sum of those per-game probabilities across the season — the win total the schedule should have produced. The gap between actual wins and xW is the team's luck differential.
Latent team strength
How good the simulator thinks a team really is
Independent of what the standings say today. Early in the season the standings are noisy; latent strength accounts for who they've played and how lucky they've been, then reports a posterior distribution over true team quality.
Credible interval
A coin-flip range — the model thinks the true value is as likely inside it as outside
Most credible intervals on this site are shown as a 50% probable range: the model thinks the true value has about even odds of landing inside that range versus outside it. It's derived from the model's wider 89% interval, narrowed assuming a roughly bell-shaped posterior — an honest simplification, not a re-run of the model. The playoff win-total band above is constructed differently and still shows its own wider range.

Read more

Long-form write-ups on the methodology, plus short explainer threads on the ideas behind the numbers.

Explainer threads

Connect

See how it’s built, or follow along.

Follow @mlb_simulator on XDaily deserve-to-win verdicts, luck calls, and the explainer threads above — posted as the games happen.

Data downloads

These feeds are stable URLs you can build on — they refresh with every data update, and I’ll keep their shape backward-compatible. Other JSON files the site loads are internal and may change without notice. Rate columns are proportions, not percentages: a bb_pct of 0.1495 is a 14.95% walk rate. The hdi89 columns are 89% credible bounds — wider than the 50% range the ranking charts draw.

  • /downloads/hitter-rankings.csvEvery qualified hitter this season (50+ PA), ranked by estimated bases per plate appearance, with the credible-interval bounds, BB%/K%/HR%/hard-hit rates and archetype. The replacement for the old Streamlit hitter rankings download.
  • /downloads/pitcher-rankings.csvEvery qualified pitcher this season (100+ batters faced), ranked by expected bases allowed per plate appearance, plus ERA, WHIP, innings and the same rate columns.
  • /downloads/team-luck.csvTeam luck rankings — actual vs. expected wins, luck differential, lucky wins, and unlucky losses for all 30 teams. One row per team per season, every season the site can compute (currently 2026 and 2025), newest first — a plain sum across the file double-counts across seasons, since season is column 1. The same table the old Streamlit app exported.
  • /data/standings.jsonCurrent-season standings — wins, losses, home/road splits, run differential, and luck differential per team, as JSON.

Request an idea

Want a chart, a metric, or a view that isn’t here yet? Tell me — it goes straight to my inbox.

max 80
max 120
max 2000

Sources

Game-level data comes from MLB’s public Stats API; player and team totals come from this project’s own data store. Every number on the site traces back to one of those two.