
About MLB Deserve-to-Win
I built a simulator that replays every game ten thousand times to estimate what should have happened. Every chart shows the league’s typical number and its best, so you can tell normal from remarkable.
How deserve-to-win works
After every game, we ask one question: did the right team win? A 110 mph line drive can get caught, and a weak grounder can find a hole. We take that randomness out in four steps.
Collect every plate appearance
We pull every plate appearance from the game: each batted ball's exit velocity, launch angle and spray angle, plus the walks, strikeouts, hit-by-pitches and stolen bases that never become contact.
Estimate what each ball should become
The batted-ball model learned from hundreds of thousands of real batted balls. It turns each ball's contact into odds: how often a ball hit that hard, at that angle, in that direction becomes an out, a single, a double, a triple or a home run.
Simulate the game 10,000 times
We replay the game ten thousand times, re-rolling each batted ball against those probabilities while keeping the walks, strikeouts, and baserunning fixed. A 110 mph line drive that got caught now falls for a hit in most of the reruns.
Read off a win probability
The share of those 10,000 reruns each team wins is its simulated win probability. 0.62 means “given how the bats sounded, you should have won 62 out of 100 games like this one.”
Stack those per-game numbers across a season and you get expected wins (xW), the win total the games should have produced. Actual wins minus xW is a team’s luck differential, shown on every team page and charted on the Teams page.
Why our numbers aren't just averages
A hitter who goes 6-for-10 in the season's first week is not a .600 hitter. Ten trips to the plate are mostly noise, so the site's player numbers come from a hierarchical Bayesian model instead of taking small samples at face value. That's a formal name for a simple idea: every player's estimate starts near the league average, then moves toward what he is actually doing as the evidence piles up. Ten plate appearances barely move it, and four hundred mostly speak for themselves.
This is also why the K% and BB% on pitcher pages don't exactly match the raw rates on other stat sites. They are the simulator's estimate of the pitcher's true rate, not a running tally. Early in a season the two can differ a lot. By the end of a season most pitchers are within a point, though an extreme strikeout pitcher can stay a few points apart.
The headline numbers, written out
- Hitter production
EB/PA = (estimated bases on contact + walks + HBP) ÷ plate appearancesWalks and hit-by-pitches count one base, and strikeouts count zero. Higher is better.
- Pitcher run prevention
xEB/PA = contact share × EB per ball in play + BB% + HBP%It uses EB/PA's units and combines strikeout, walk and contact estimates. Lower is better.
- Expected wins
xW = sum of each game's simulated win probabilityA 0.62 deserve-to-win game adds 0.62 wins to the tally, whether the team won or lost.
- Luck differential
Luck = actual wins − xWPositive means a team won more games than its play earned, and negative means fewer.
What’s on the site
These 8 hubs are all built on the same simulated numbers.
- Games→Every game gets a deserve-to-win verdict, plus luck ledgers, pitching matchups and spray charts drawn over each park's real fences.
- Standings→The playoff picture comes with the storylines behind it: division races, wild-card ladders, projected wins and a full luck ledger for all 30 teams.
- League→League-wide context shows the headline rates and who leads them, how the 30 teams are spread, and season-by-season trend lines you can split by league and division.
- Teams→A quality map shows where every team really stands, with luck halos marking who has won more or fewer games than their play earned.
- Hitters→Every qualified hitter gets leaderboards, skill radars and archetypes, showing who is producing and who has been unlucky.
- Pitchers→Pitchers get the same treatment: expected bases allowed, arsenal and contact profiles, and how each one fares every time through the order.
- AAA→The same estimated-bases ranking runs one level down for every Triple-A hitter with enough plate appearances, each shown with the range his season supports.
- Tools→The batted-ball landing explorer shows where balls land for any exit velocity and launch angle, and Compare players puts two hitters or two pitchers side by side.
Metrics glossary
Here are the metrics the site uses most, in plain language, with the direction that counts as “better” spelled out.
- Deserve-to-Win
- Share of the simulator's 10,000 replays a team wins
- Deserve-to-Win is the share of the simulator's ten thousand replays of a game that a team wins. A 62% means that, given how the ball was hit, the team should win about 62 of 100 games played like this one. Higher means the team played better. It describes the game that was played, not a prediction.
- Estimated bases (EB)
- How many bases a batted ball should have been worth
- The batted-ball model looks at how hard, how high and which way a ball was hit, and estimates what it is worth on average, from 0 bases to 4. Higher is better for the hitter. A 105 mph line drive that gets caught is still worth about 0.7, and a soft ground-ball single only about 0.25.
- EB/PA
- Estimated bases per plate appearance
- EB/PA is a hitter's estimated bases per trip to the plate. Each batted ball counts its estimated bases, each walk or hit-by-pitch counts one, and a strikeout counts zero. Higher is better. It credits the quality of contact, not where the ball landed. Rankings use the simulator's estimate of his true rate, so one hot week can't leapfrog a season.
- xEB/PA
- Expected bases a pitcher allows per plate appearance
- xEB/PA is the pitcher's version of EB/PA, in the same units, but lower is better. The "x" means expected. The simulator estimates his strikeout, walk and hit-by-pitch rates and the contact he allows separately, then combines them into expected bases per plate appearance.
- Bases created
- A hitter's estimated bases plus walks, over a game or season
- Bases created adds up each batted ball's estimated bases plus the walks and hit-by-pitches that put the hitter on base for free, over a game or a whole season. Higher is better. It credits the quality of contact, not where the ball happened to land.
- Bases prevented
- How many bases a pitcher suppressed versus an average outing
- Bases prevented is the pitcher's side of bases created. It counts how many fewer estimated bases he allowed than an average pitcher would have against the same batters. Higher is better, and it rewards weak contact and strikeouts, not lucky defense. An average pitcher is a stricter bar than the replacement level the best and worst boards use.
- Excess bases
- Bases above or below a replacement-level baseline
- Excess bases compares a player's estimated bases with what a readily available fill-in would have produced in the same chances. Higher is better, for a hitter or a pitcher. The best and worst performance boards use it. Big numbers need both quality and volume, so a great night in six trips beats a great night in three.
- Batted-ball luck
- Actual bases minus deserved bases, per ball or per player
- On a game page, batted-ball luck is what a ball actually earned (out 0, single 1, up to 4 for a home run) minus its estimated bases. Positive means the hitter got more than the contact deserved, and negative means a well-hit ball was caught. It is not the same as a team's season luck differential.
- Upset
- A favorite that lost, judged before the game or by Deserve-to-Win
- The Home page's "Upset of the day" is the loss by the night's biggest pre-game favorite, with the odds set from projected wins and home field. The bigger the favorite, the bigger the upset. On the Games tab, the upset badge means something else: the team with the higher Deserve-to-Win lost.
- Deserve gap
- How far apart the two teams were on Deserve-to-Win
- Deserve gap is the distance between the two teams' Deserve-to-Win percentages in one game. Near zero means the simulator saw a coin flip, whatever the scoreboard said. A wide gap means one side clearly outplayed the other. It doesn't say who won, and when the loser had the higher Deserve-to-Win, the Games tab marks the game an upset.
- Times through the order
- How a pitcher fares each time through the lineup
- Hitters tend to improve each time they see the same pitcher in a game. The pitcher-page section shows his strikeouts, walks and estimated bases allowed on each trip through the lineup. Holding steady is better for a pitcher, and a steep drop the third time through is the usual case for pulling a starter earlier.
- BB%
- Walk rate per plate appearance
- BB% is the share of plate appearances that end in a walk. For a hitter higher is better, because it shows patience and pitch selection. For a pitcher lower is better. The league sits around 8 to 9%.
- K%
- Strikeout rate per plate appearance
- K% is the share of plate appearances that end in a strikeout. For a hitter lower is better, because more balls go in play. For a pitcher higher is better. The league sits around 22%.
- HR%
- Home runs per plate appearance
- HR% is the share of plate appearances that end in a home run. For a hitter higher is better, and for a pitcher lower is better. League-wide, it shows how much power an era has, and the league sits around 3%.
- R/G
- Average runs per team per game, counting both teams
- R/G averages the runs each team scored across every game in the view, counting both teams in each game. It is the plainest measure of how high-scoring an era is, so higher means more scoring, not better play. For a division, it describes the games that division played in, so the opponents' runs count too.
- Hard-hit%
- Share of batted balls hit 95 mph or harder
- Hard-hit% is the share of batted balls hit 95 mph or harder. Hard-hit balls fall in for hits far more often than soft contact, so the rate is a useful early read on a hitter. Higher is better for a hitter and lower is better for a pitcher.
- FB%
- Fly-ball rate, the share of balls launched at 25 to 50 degrees
- FB% shows how often a hitter lifts the ball at the angles that produce extra-base hits. Fly balls turn into doubles, triples and home runs, and ground balls almost never do. It describes a hitter's shape, not his quality, so neither high nor low is better. The league sits around 23%.
- GB%
- Ground-ball rate, the share of balls launched at 10 degrees or lower
- GB% shows how often contact stays on the ground. Pitchers with a high GB% limit damage, because ground balls rarely go for extra bases. Like FB%, it describes a shape, not a quality, so neither high nor low is better on its own. The league sits around 45%.
- Barrel%
- Share of batted balls hit at an ideal speed and angle
- A "barrel" is Statcast's sweet spot: at least 98 mph off the bat, at a launch angle window that widens as the ball is hit harder. About three in four barrels go for hits, and about half are home runs. Higher is better for a hitter and lower is better for a pitcher. The league sits around 7%.
- Pull%
- Share of batted balls hit to the batter's pull side
- A right-handed hitter pulls the ball to left field, and a left-handed hitter pulls it to right. Pulled contact is where most home-run power lives, but an extreme pull habit also lets defenses position for it. Neither high nor low is better, because it describes a hitter's shape, not his quality.
- Bat speed
- How fast the bat's sweet spot moves at contact, in mph
- Bat speed counts competitive swings only, so checked swings and bunts don't drag a hitter's number down. Faster isn't better on its own, because bat speed trades power against contact. Luis Arraez swings about 62 mph, near the bottom of the league, and has one of the lowest strikeout rates in baseball. League-wide tracking starts in 2024.
- Swing length
- How far the bat's sweet spot travels during the swing, in feet
- A longer swing can build more bat speed but takes more time to reach the zone, so it is a style tradeoff, not a quality score. Neither longer nor shorter is better. It counts competitive swings only, the same as bat speed.
- Attack angle
- The bat's upward or downward tilt through the zone, in degrees
- Attack angle is the bat's upward or downward tilt at contact. A more upward angle meets the pitch's downward path and helps lift the ball, and a level one keeps the bat in the zone longer. Neither is better on its own, because it is a swing-style choice, not a quality score.
- BB+HBP%
- Walks plus hit-by-pitches per plate appearance
- BB+HBP% is the free-base rate, the walks plus times hit by a pitch per plate appearance. It runs a little higher than BB% because hit-by-pitches are folded in. Higher is better for a hitter.
- Whiff%
- How often a swing at this pitch misses it entirely
- Whiff% is swings that miss, divided by swings taken. It is a per-swing rate, so a pitch nobody offers at cannot inflate it. On the site a foul tip counts as a miss, matching the convention Baseball Savant uses, since the bat did not square it up. High is good for a pitcher and bad for a hitter.
- Chase%
- How often a hitter swings at a pitch outside the strike zone
- Chase% is swings at pitches outside the strike zone, divided by the pitches outside the zone he saw. It is the plainest measure of plate discipline, because a hitter who rarely chases makes pitchers come to him. Lower is better for a hitter, and a pitcher wants it high.
- Archetype
- The kind of player his numbers group him with
- Each player with enough playing time is grouped with the players whose rates look most like his. The groups come from the data, not from anyone's labels. An archetype describes a style, not a level, so two hitters in the same group can be far apart in quality.
- Grade
- A letter grade for his percentile in EB/PA or xEB/PA
- A grade is a letter for a player's percentile. It ranks his established level of EB/PA among hitters with 50+ plate appearances, or of xEB/PA among pitchers with 100+ batters faced. An A+ is 90 and up, then A 80, B+ 70, B 60, C+ 50, C 40, D 30, and F below that.
- Luck differential
- Actual wins minus expected wins
- Luck differential is a team's actual wins minus its expected wins. Positive means the team has won more games than its play earned. Negative means it has won fewer.
- Win %
- Share of games won
- Win % is the simplest "are they good?" number, and higher is better. The rest of the Standings page explains why it is what it is, through luck, schedule, run differential and team strength.
- Run differential
- Runs scored minus runs allowed
- Run differential is runs scored minus runs allowed, and higher is better. A team can sit at 24-20 by winning blowouts and losing close games, or the other way around. Over a long season it tends to predict a team's future record better than win % does, so it is a plain check on whether the wins are real.
- Home/Road split
- Record at home versus on the road
- Home/road split is a team's record at home next to its record on the road. Most teams play a little better at home. The split is not good or bad on its own, but a big gap either way is worth a look.
- Playoff probability
- Share of simulated seasons where this team makes the playoffs
- The simulator plays out the rest of the season thousands of times, drawing each team's true quality from its range of likely values. Playoff probability is how often this team gets in across those seasons, so higher is better.
- Projected wins
- Average final win total across simulated seasons, with a 90% range
- Projected wins is the simulator's best guess at a team's final win total, and higher is better. The range beside it is where the total lands 90% of the time. A projection of 90 with a range of 82 to 98 is less certain than one with a range of 87 to 93.
- Expected wins
- Win total the simulator says a team should have
- For each game, the simulator estimates a win probability from the contact quality both teams produced. Expected wins (xW) is the sum of those probabilities across the season, the win total the games should have produced, so higher is better. Actual wins minus xW is the team's luck differential.
- Team strength
- How good the simulator thinks a team really is
- Team strength is the simulator's estimate of how good a team really is, apart from what the standings say today, and higher is better. Early in the season the standings are noisy. Team strength accounts for who a team has played and how lucky it has been, and it reports a range rather than a single number.
- Credible interval
- A range the true value is as likely inside as outside
- Most ranges on this site are 50% ranges: the simulator thinks the true value is as likely to be inside as outside. A narrower range means more certainty. They are narrowed from the simulator's wider 89% intervals. The playoff win-total range is built differently and stays at 90%.
Fine print
These notes explain the method behind a few charts, for anyone who wants it.
Three kinds of upset
Three marks on the site point at an upset, and each one measures something different. On the Games table, the Upset badge marks any game where the team with the higher Deserve-to-Win lost, however small the gap. The LUCKY pill on Home's game rail is stricter: the losing team had to win more than 15 percentage points more of the simulator's replays than the winner did. "Upset of the day" on Home uses the odds before the game instead, from each team's projected wins plus a home-field edge.
Pitch movement
Rise is induced vertical break, the drop a pitch fights off through backspin with gravity taken out. Break is sideways movement toward the pitcher's arm or glove side, with lefties flipped. Each dot averages his last five outings, weighted by how many of that pitch he threw. The trail is that average at earlier dates and needs six tracked outings. The halo is where ordinary noise puts the dot about two times in three. The hollow ring is the league median, both hands together, among pitchers with 50 or more of that pitch tracked. A pitch marked * has none, because fewer than eight pitchers qualify. On a wider screen, a ring far from its dot carries the pitch's code.
When a pitch counts as changed
A long trail is not proof on its own that a pitch changed. A pitch counts as “changed” when the pitcher's last third of outings sits at least two inches from his first third and the gap beats his normal noise, so it needs nine tracked outings. Before comparing, we take out the league's own drift for that pitch type over the same dates. On average, pitches across the league lose a little break through midsummer and get some of it back by September, and 2024 and 2025 show the same pattern. We correct only the verdict, because the league ring is drawn uncorrected too, so a trail can look long while the verdict says “steady.”
Why the league best is named, not drawn
On the scatter charts, the league best is named above the chart instead of drawn on it. The best hitter or pitcher on a rate usually sits far from everyone else, so stretching the axes to reach him would squeeze the rest of the players into a corner and pile their faces on top of each other. Picking out a player is what these charts are for, so the best gets a name and a number instead of a mark. The same is true of the team charts on the Teams page.
The Triple-A board
The Triple-A estimates are fitted on Triple-A plate appearances only, so they share units with the major-league pages but not a scale. A hitter needs 100 plate appearances to be ranked. Hitters with 25 to 99 plate appearances are listed alphabetically below the board and do not set the average or the best, so one of them can sit past the best mark on a small sample. Age is shown beside every hitter and never used in the ranking. A "No Triple-A" marker means the data has no record of him at this level in the seasons it checked. A veteran back from the majors carries the same marker as a newcomer, so it is not a first-season label.
The landing explorer
The explorer sorts every tracked ball in play the site has on record by where it came down. The sliders keep only the balls hit within 2.5 mph and 3 degrees of the exit velocity and launch angle you pick, so the field shows where one kind of contact tends to go. Each square's estimated bases come from the batted-ball model, the same one behind every game on the site. The smoothed view blurs neighboring squares together, which reads faster but hides where one square ends. The park count compares the landing distance with each park's real wall at that angle and ignores wall height.
The 50% ranges
Most ranges on the site are 50% ranges, narrowed from the 89% intervals the simulator's player estimates come with. Each side of the 89% interval is pulled toward the estimate by the same factor, which is exact when the uncertainty is bell-shaped and an approximation otherwise. That makes it a simplification, not a second run of the simulator. The Triple-A board shows the full 89% interval instead, and the downloadable rankings files carry the 89% bounds in their own columns.
How the luck boards count
Walks and hit-by-pitches sit on neither side of the luck boards, because nothing was put in play, so there was no luck to have. The totals are whole bases over the whole season, so a hitter with more balls in play, or a pitcher who has faced more batters, has had more chances to gain or lose them. Only players with at least 50 balls in play appear. The center line is zero by definition, not a league average. The typical player sits near it but not on it, and the shared image prints the median player's gap, which shows how close to zero the league really sits.
Faces and lines on the league landscape
On the league landscape, only the top players by EB/PA, or by xEB/PA on the pitchers page, get a headshot, and a phone shows fewer faces still. Everyone else is a dot in his team's color, so the picture stays readable instead of turning into a pile of overlapping faces. The dashed reference lines are the walk and home-run rates across the whole league this season, not an average of the players plotted at the moment. They stay put when you change the scope, so a division view is still measured against the whole league.
The league band on the luck trend
The league band on the results vs. contact quality chart comes from full-season rates, one per qualified hitter: his actual bases per ball in play over the whole season. A hot stretch of balls in play pokes above that band more easily than a season line would, and the comparison is built that way on purpose. A point above the band reads as a hitter who has hit like a top-quarter hitter for about three weeks. When his two lines split, the gap is batted-ball fortune, not a change in how well he is hitting the ball.
What the pitch mix chart can and can't say
On the pitch types seen chart, a run of different opposing staffs moves the lines too, so a shift shows what he saw, not always a plan against him. The season-by-season rows are whole-season totals, each compared with that season's own league share. The site's pitch-by-pitch data starts in 2024, so no earlier season appears.
How platoon splits are estimated
Platoon splits are shrunk estimates. On held-out seasons, shrunk splits beat raw splits, most decisively at low plate-appearance or batters-faced counts (relievers especially), and the two converge as playing time builds. For a pitcher, the two splits, weighted by how often he faces each side, average back to his modeled overall EB/PA allowed. That number is not the xEB/PA headline at the top of his page, which folds in strikeouts and walks differently.
Bat tracking and a real swing change
Bat tracking counts competitive swings only. Bunts and check swings under 50 mph are excluded, so the site's bat speeds sit a touch below the ones Baseball Savant publishes, which use a different cut. League tracking starts in 2024, and early 2024 has small gaps, with about 93% of that season's swings carrying a reading. A real change in a swing means the move from his first full month to his latest one is more than three times its own margin of error and big enough to matter, after subtracting the league's own drift. Bat speed rises about half a mile an hour across the league every spring, and that is not a swing change.
How swing timing is measured
Timing numbers are Baseball Savant's season-to-date figures, and tracking starts in mid-2023. A perfect-contact swing is centered (the barrel within 4 inches of the ball horizontally), on time (within 7 milliseconds of ideal timing) and lined up (within 2 inches vertically), all at once, and a swing missing any of the three counts as flawed. Switch hitters have both sides combined, because no side-by-side timing split exists. A real change in timing means the gap between this season and last is more than three times its own margin of error and big enough to matter. Almost no one clears that bar, because a couple of percentage points either way is normal year-to-year wobble.
How the stance diagram is drawn
The stance diagram reconstructs his feet from Statcast's stance measurements (hip position, foot separation and stance angle) about one second before the pitch. It approximates his setup, not his stride or swing, and the average hitter sets up open rather than square. A month needs 25 tracked swings to appear on the trend chart, so an injured or platooned month can be missing. The change line only appears for a sustained month-over-month move that clears normal Statcast noise by a wide margin. On the angle chart, hover a point, or drag on a phone, and the tooltip says open or closed in words.
Read more
These are the long-form write-ups on the method, plus short explainer threads on the ideas behind the numbers.
- Who Deserved to Win? Building an MLB Game Outcome SimulatorRead →
This is the original write-up of how I built the game simulator, why, and what it found.
- Applying Bayesian Hierarchical Methods to MLB Season Win ProbabilitiesRead →
This one rolls the per-game deserve-to-win results up into an estimate of each team's true strength.
Explainer threads
- Player metricsWhat is EB/PA?This thread shows how the simulator estimates a hitter's true production, and why sample size matters.Read thread →
- Run scoringCan a Simple Formula Predict Baseball Scores?This thread shows why the Poisson distribution fails for MLB runs, and what works better.Read thread →
- ProjectionsUnderstanding Player ProjectionsThis thread shows how multi-year Bayesian projections handle aging, uncertainty and small samples.Read thread →
- Small samplesHow Much Should You Trust the First Week?Five games tell you almost nothing. This thread shows when the estimates start to settle.Read thread →
Connect
See how it’s built, or follow along.
Follow @mlb_simulator on XThe account posts daily deserve-to-win verdicts, luck calls and the explainer threads above as the games happen.FollowData downloads
These feeds are stable URLs you can build on. They refresh with every data update, and I’ll keep their shape backward-compatible. Other JSON files the site loads are internal and may change without notice. Rate columns are proportions, not percentages: a bb_pct of 0.1495 is a 14.95% walk rate. The hdi89 columns are 89% credible bounds, wider than the 50% range the ranking charts draw.
- /downloads/hitter-rankings.csvThe file ranks every qualified hitter this season (50 or more plate appearances) by estimated bases per plate appearance, with the 89% bounds, BB%, K%, HR%, hard-hit rate and archetype. It replaces the old Streamlit hitter rankings download.
- /downloads/pitcher-rankings.csvThe file ranks every qualified pitcher this season (100 or more batters faced) by expected bases allowed per plate appearance, with ERA, WHIP, innings and the same rate columns.
- /downloads/team-luck.csvThe file holds team luck rankings for all 30 teams: actual and expected wins, luck differential, lucky wins and unlucky losses. There is one row per team per season, for every season the site can compute (currently 2026 and 2025), newest first. Because
seasonis column 1, a plain sum across the file double-counts. It is the same table the old Streamlit app exported. - /data/standings.jsonThe file holds current-season standings as JSON, with wins, losses, home and road splits, run differential and luck differential for each team.
Request an idea
Want a chart, a metric or a view that isn’t here yet? Tell me, and it goes straight to my inbox.
Sources
Game-level data comes from MLB’s public Stats API. Player and team totals come from this project’s own data store, and every number on the site traces back to one of those two.