Every team that finishes ahead of its run differential acquires a mythology by Thanksgiving: the bullpen built for close games, the manager’s buttons, the clubhouse that knew how to win. The claim under all of it is testable. If beating your Pythagorean record is a skill — anything a front office actually owns — the teams that do it should tend to do it again. Skills persist; luck doesn’t. So I built every team’s Pythagorean residual — actual wins minus the wins its runs scored and allowed predict — for all 120 team-seasons in the bundled 2021–2024 standings, paired each season with the same franchise’s next one, and correlated this year’s residual with next year’s across all 90 pairs.

The answer is r = +0.003. Not small — absent. Knowing a team beat its run differential by six wins tells you nothing about which side of the ledger it lands on next year. And before you file that under “the data was mute,” the same 90 pairs speak loudly when asked the control question: this year’s luck strongly predicts next year’s win change, at almost exactly minus one win per lucky win. Pythagorean luck doesn’t repeat. It un-happens.

+0.003correlation between a team's Pythagorean residual and its next-year residual, 90 pairs (2021–2024), 95% CI [−0.20, +0.21]
−0.98next-season win change per residual win — the luck pays itself back almost exactly 1-for-1 (r = −0.32)
3.0mean size of a season's residual, in wins; 57% of team-seasons land within ±3, and 14 of 120 beyond ±6
2 of 30teams on the same side of their run differential all four seasons — coin flips would produce about 3.8

The question three neighbors left open

This site has walked all the way around this result without stepping on it. The luckiest and unluckiest teams of 2023 named one season’s outliers and called the gap mostly luck — asserted, not tested. The run-differential piece showed Pythagorean win% out-predicts actual win% (0.57 vs 0.54), which implies the gap is noise, but never measured the gap’s own carryover. And the .500 gravity well measured how hard records regress without being able to say how much of the snap-back was evaporating luck, because its decade file has no run columns. This piece runs the direct test all three point at: is the residual a trait, or a coin flip?

Method, in one paragraph

The data is the bundled multiseason file: W, L, runs scored, and runs allowed for all 30 teams, 2021 through 2024. (The site’s longer 2015–2024 standings file carries wins and losses only — no RS/RA — so four seasons and 90 consecutive-season pairs is the honest maximum window; more on that below.) Expected wins are the Pythagorean formula with the exponent fit to this data rather than taken on faith: least squares over all 120 team-seasons gives 1.74, in line with the 1.69 this site fit to 2023 alone and below the textbook 1.83. A team-season’s residual is actual wins minus expected wins; across the file the residuals average −0.1 (centered, as they should be), the typical one is 3.0 wins in either direction, and the tails are real: 14 of 120 team-seasons missed their expectation by six or more wins. Then the test: pair every residual with the same franchise’s residual one year later, and correlate.

The flat line

Two scatter panels built from 90 consecutive-season MLB team pairs, 2021 to 2024. Left panel: Pythagorean residual in year N against residual in year N plus 1, a shapeless cloud with a perfectly flat fit line, correlation plus 0.003, slope 0.00; the 2021 Mariners are labeled at plus 14 wins and the 2021 Diamondbacks at minus 10.1. Right panel: residual in year N against the change in wins the following season, a clearly downward-sloping cloud with fit slope minus 0.98 and correlation minus 0.316; the 2021 Diamondbacks are labeled improving by 22 wins and the 2021 Mariners changing by zero.
Left: this year’s Pythagorean residual against next year’s — the persistence test, and the line is flat (r = +0.003). Right: the same residual against next season’s change in wins — the regression test, and the line falls at −0.98 wins per residual win. Data: MLB Stats API standings + run totals 2021–2024, bundled as data_layer/standings_multiseason.json (retrieved 2026-07-09), charted by charts/chart_pythag_luck_persistence.py.

The left panel is the headline: r = +0.003, 95% CI [−0.204, +0.210], slope indistinguishable from zero. The group version says it even more plainly than the correlation does. The 15 luckiest team-seasons in the file averaged +5.7 wins over expectation; the following year they averaged −0.9. The 15 unluckiest averaged −6.7; the following year, −0.8. Both groups — the blessed and the cursed — landed in the same place: zero, give or take a rounding error. Whatever put them in the tails did not travel with them.

The file’s single largest residual makes the case by itself. The 2021 Mariners went 90–72 while being outscored by 51 runs — 14.0 wins above expectation, the biggest overshoot in the file by five clear wins. If beating your differential is a skill, the 2021 Mariners are its Mozart. In 2022 their residual was +1.8. (They won 90 again — but honestly, with a +67 differential. The luck didn’t repeat; the team got better instead.)

The placebo: luck doesn’t repeat, it un-happens

A null on 90 pairs invites the objection that the test just can’t see anything. So here is the control. If these residuals are luck, they should predict one thing extremely well — next year’s regression: a team that won 88 games on 82-win run production should fall, not because it got worse but because the phantom wins evaporate. That prediction holds at full strength: residual versus next-season win change correlates at −0.32 (95% CI [−0.49, −0.12]), with an OLS slope of −0.98 — each win of Pythagorean luck predicts almost exactly one win handed back. The 15 luckiest teams declined by 5.1 wins on average; the 15 unluckiest improved by 6.1. The 52–110 Diamondbacks of 2021, the file’s unluckiest team, gained 22 wins in 2022; the 2022 Rangers, second-unluckiest at −9.5, gained 22 in 2023 and won a World Series. Same data, same method, same n — a strong signal exactly where luck says one should be, a dead flat line exactly where skill would have to be. That contrast, not the null alone, is the finding.

One number keeps the bookkeeping honest, because the residual and the win change share a term (this year’s wins appear in both). Correlate this year’s residual against next year’s winning percentage instead — no shared term — and you get −0.018. Luck predicts the change because it’s subtracted back out; it contains nothing at all about where the team ends up. That is the precise sense in which regression to the mean is not a force pulling teams down — it’s the accounting correction for wins that were never earned.

The eight biggest departures, and what happened next

Every residual of 6.5 wins or more in the bundled 2021–2024 file, with the same team’s residual one year later. Expected wins use the fitted exponent 1.74. Data: MLB Stats API, data_layer/standings_multiseason.json.
Team-seasonRecordRun diffResidualNext year’s residual
2021 Mariners90–72−51+14.0+1.8
2021 Diamondbacks52–110−214−10.1−3.3
2022 Rangers68–94−36−9.5−5.4
2023 Padres82–80+104−9.4+3.1
2023 Marlins84–78−57+8.8+0.2
2024 White Sox41–121−306−8.5(2025 not in file)
2023 Orioles101–61+129+7.8+1.8
2024 Cardinals83–79−47+6.8(2025 not in file)

Six of the eight have a next season in the file, and none of the six repeated anything close to their departure — the largest follow-up is the Rangers’ −5.4, and the 2023 Padres flipped sign entirely, from −9.4 to +3.1 (and from 82 wins to 93). The 2023 rows agree with the season-specific piece, which quoted the Marlins at +9 and the Padres at −10 using the textbook exponent; the fitted 1.74 shaves the decimals, not the story.

Did anyone stay lucky three straight years?

The strongest form of the skill claim is a streak, so I checked every one. Seven of the 30 franchises kept a same-signed residual three or more consecutive seasons (five positive, two negative) — and cointossing would produce more: with independent 50/50 signs, about 11 of 30 teams would show a three-streak somewhere in four seasons. Two franchises ran the table on the plus side all four years, the Pirates (+1.5, +3.2, +4.3, +2.4) and the Tigers (+1.7, +2.1, +4.9, +0.8), against roughly 3.8 expected by chance. The Pirates are the best skill candidate the file can offer, and even they are less streaky than a random-number generator’s output. The sign test is a heuristic, not a proof (residual signs aren’t perfectly 50/50 or independent), but it lands where the correlation landed: nothing here outruns chance.

Who’s “due” right now

Per the flagship’s bundled live snapshot (MLB Stats API, retrieved 2026-08-05, teams averaging 113.9 games — a partial-season read, not a projection), the current residual leaders are the Rays at +6.8 wins over expectation (67–46, +37 differential) and the Reds at +4.0 (54–58 despite being outscored by 62); the trailer is the Tigers at −8.8 — 55–58 with a +71 run differential. Yes: the one franchise that stayed on the lucky side all four bundled seasons is currently the unluckiest team in baseball, which is about as pointed a summary of this article as the league could have arranged. What the 90 pairs license you to say is exactly this much: nothing about anyone’s residual next season, and a lean — not a promise — that the phantom wins and losses get paid back.

Limitations, stated plainly

Four of them. First, n = 90 pairs from four seasons, and the confidence interval [−0.20, +0.21] is wide — this test would miss a modest persistence of 0.1, under half a win of carryover on a typical residual. “No detectable skill, and any undetected one is tiny” is the defensible claim; “exactly zero” is not. Second, the window is 2021–2024 because it has to be: the site’s decade file (standings_2015_2024.json) carries W/L only, no runs — I verified — so the multiseason file with RS/RA is the whole usable universe. The classic studies on longer histories find the same null; comforting, but that’s their result, not mine. Third, the exponent doesn’t matter: rerun everything at the textbook 1.83 and persistence comes out r = +0.016 [−0.19, +0.22], same flat line. Fourth, this test can’t distinguish pure luck from a real skill that churns completely year to year (a bullpen edge that dissolves with each winter’s roster turnover would also show no persistence) — but non-persistence is precisely what the luck hypothesis predicts and what the skill narrative can’t survive, and the residual’s best-known ingredients, one-run records and cluster timing, are the league’s two most reliable noise generators.

The bottom line

Beating your run differential is not a skill this data can find. The residual correlates with its own future at +0.003; the luckiest fifteen teams and the unluckiest fifteen both came back to zero; the only teams that stayed on one side of the ledger did so less often than coin flips would; and the single thing a residual reliably predicts is its own disappearance, at a rate of one win back per lucky win. When your team beats its Pythagorean record, enjoy every stolen one-run game — those wins are real and they count. Just don’t book them again next April. The standings gave them; the standings take them back.

Reproduce it

The file ships in the site’s data_layer/ (provenance in data_layer/SOURCE.txt), and both correlations fit in twenty lines:

import json, math, statistics as st
from collections import defaultdict

K = 1.737   # least-squares fit over all 120 team-seasons
D = json.load(open("data_layer/standings_multiseason.json", encoding="utf-8"))
by = defaultdict(dict)
for t in D["teams"]:
    by[t["season"]][t["team_id"]] = t

def resid(t):
    e = t["RS"]**K / (t["RS"]**K + t["RA"]**K)
    return t["W"] - e * (t["W"] + t["L"])

S = sorted(by)
pairs = [(resid(by[a][i]), resid(by[b][i]), by[b][i]["W"] - by[a][i]["W"])
         for a, b in zip(S, S[1:]) for i in by[a]]

def corr(x, y):
    mx, my = st.mean(x), st.mean(y)
    return sum((a-mx)*(b-my) for a, b in zip(x, y)) / math.sqrt(
        sum((a-mx)**2 for a in x) * sum((b-my)**2 for b in y))

r0, r1, dw = zip(*pairs)
print(len(pairs), "pairs")
print("luck vs next year's luck:       r = %+.3f" % corr(r0, r1))
print("luck vs next year's win change: r = %+.3f" % corr(r0, dw))

# output:
#   90 pairs
#   luck vs next year's luck:       r = +0.003
#   luck vs next year's win change: r = -0.316

The two-panel exhibit is charts/chart_pythag_luck_persistence.py: it refits the exponent, rebuilds all 90 pairs, and picks the labeled extremes by computed value at run time. Confidence intervals in the text are standard Fisher-z with n = 90.

Sources & Further Reading