The Formulas

you will still be fiddling with your [flipping] formulas

Pete Alonso to Mets president David Stearns, on how his Hall of Fame case would be judged

He was not wrong. This is the fiddling, published in full.

Everything below is rendered directly from the code that runs the site. There is one copy of each number, so what is published here and what the site computes with cannot drift apart. If a threshold changes, this page changes with it.


1. The decline thresholds

Six signals, each scored STRONG, WATCH or DECLINING from the latest snapshot. Nothing is stored as a judgment; the chips are computed from these numbers every build.

Decline thresholds for each of the six signals
SignalFieldStrongWatchDeclining
Exit velocityavgEV>= 93.592 to 93.4< 92
Hard-hit ratehardHitPct>= 5044 to 49.9< 44
ProductionwrcPlus>= 125105 to 124< 105
Plate disciplinebbPct>= 108 to 9.9< 8
Strikeout ratekPct< 25.825.8 to 28.6>= 28.7
Durabilitygames, streak, ILstreak live, or >= 95%80% to 95%< 80%, or IL > 30d

Strikeout rate is the one band this site does not fix in advance. It is computed from Alonso's own completed seasons: a baseline of 23.0% across 7 seasons, with a standard deviation of 2.85 points about it. WATCH begins one standard deviation above the baseline and DECLINING at two, so the numbers in the row above move as his record grows. A hitter who has always struck out a lot is not declining, and the absolute line this replaced said as much in its own rationale.

Why these numbers

Exit velocity
The 92 mph line is this site’s calibration, not an outside standard. It is drawn where a drop is large enough to read as lost contact quality rather than as a slow month, and exit velocity is the signal this site expects to move first.
Hard-hit rate
Calibrated so the ZiPS base case crosses into WATCH late in the contract rather than at signing.
Production
Set so the published ZiPS glide path (112 OPS+ by 2030) reads WATCH around 2029, matching what the scenario model implies.
Plate discipline
Aging hitters who hold or improve walk rate age best; a collapsing walk rate tends to follow bat-speed loss.
Strikeout rate
Read against his own record rather than an absolute line. The baseline is the mean of his completed seasons and the bands are set at one and two of his own year-to-year swings above it, so a hitter who has always struck out a lot is not flagged for it. There is no absolute backstop on purpose: if a high rate is costing him, production carries that, and production is its own signal.
Durability
Injury is the most common cliff trigger in the comp set, so durability is scored on availability rather than on any rate stat. The 80-95% band without a live streak reads WATCH.

These cut points are this site’s own. None of them is an outside standard. Five are fixed in advance, set so the published ZiPS glide path crosses into WATCH late in the contract rather than at signing, and drawn where a fall reads as lost quality rather than a slow month. The sixth, strikeout rate, is derived from the player’s own completed seasons instead, so it moves as his record grows. The five fixed bands are absolute rather than age-adjusted, which means a 140 wRC+ reads the same at 35 as at 31 even though the first is the better result. That is a known limitation, kept for now because an auditable rule beats a clever one.

What each production band went on to produce

The line above is a judgement about a rate. This is what happened next to the 162 hitters the comp rule admits, grouped by where they sat at age 30 and measured across the same five ages this contract buys. The pool averaged 10.3 wins.

What hitters in each production band produced across ages 31 to 35
Band at 30HittersMean wins, 31-35Under five wins
STRONG9312.423%
WATCH547.148%
DECLINING1510.440%

So the STRONG line is doing real work, and it is less impressive than the word. It separates cleanly, and clearing it by a point is not the same bet as clearing it by forty: 125 to 139 produced 9.7, 140 to 159 produced 13.2, 160 and up produced 18.7. A hitter who clears the line and no more has done a little worse than the pool average.

The bands below STRONG cannot be checked this way, and are not predictions. DECLINING reads higher than WATCH above, which is a selection effect rather than a finding: every hitter in the pool cleared a weighted OPS+ of 115 through 30, so a low figure in that one season means a hitter who dipped, not a weak hitter. Those lines are definitions of what this site will call decline. Only the STRONG line is measured here.

Two limits on all of it. The pool carries OPS+ and this threshold is wRC+, which are close cousins and not the same statistic, so this checks the shape of the line rather than its exact value. And only production and durability can be checked at all: exit velocity, hard-hit rate and walk rate are absent from the comp data, and two of them exist only from 2015. Those three remain judgement, which the table above cannot fix and this page should not pretend it does.


2. The verdict rule

Four verdicts, in order of severity. Which one the site publishes is decided by a rule written down before the answer became inconvenient, so a verdict change is never a matter of taste.

  1. 1NOT YETNo decline signal has broken down. The profile still matches the hitters who lasted.
  2. 2CRACKS SHOWINGDeterioration is visible but not decisive. Could be noise, could be the start.
  3. 3DECLININGThe answer to this site’s question has flipped. Multiple signals have broken down.
  4. 4CLIFFNot a gradual fade. Production and either durability or contact quality gone at once.

The rule

Let D be the number of signals reading DECLINING, and W the number reading WATCH.

CLIFF           if production is DECLINING
                and (durability or exit velocity) is DECLINING
DECLINING       else if D >= 2
CRACKS SHOWING  else if D == 1, or W >= 3
NOT YET         otherwise
Every combination of declining and watching signals, and the verdict each one produces. With no signal declining the verdict is NOT YET until three read WATCH. From one declining signal onward the watch count no longer changes the answer: one is CRACKS SHOWING and two or more is DECLINING, whatever else is on watch. Combinations needing more than 6 signals cannot occur and are left blank. The table below carries the same grid. NOT YETCRACKSSHOWINGDECLININGCLIFFonly six signals,so these cannot happennow01234560123456D, signals reading DECLININGW, reading WATCH
The rule reads two counts, and the picture shows what the ordered list beside it cannot: the watch count stops mattering the moment anything reads DECLINING. Every state with one declining signal is CRACKS SHOWING whether nothing else is on watch or all five are, and every state with two is DECLINING on the same terms. W decides the answer in one column of seven. The site currently reads D = 0, W = 0, marked above. From there CRACKS SHOWING is 3 more signals on watch away, or one signal dropping to DECLINING. Two routes, and the grid shows both as the same short distance. CLIFF is the exception, and it is not a count. It fires when production reads DECLINING and either durability or exit velocity reads DECLINING with it, which is a statement about which signals broke rather than how many. Two axes of counts have nowhere to put that, so it can supersede the answer anywhere in the DECLINING band, and it is unavoidable only in the corner where every signal has gone.
Show the diagram as a table
The verdict for every reachable pair of counts, computed by the same function the site publishes with. A blank cell is a state the six signals cannot produce.
W \ D0123456
6CRACKS SHOWING
5CRACKS SHOWINGCRACKS SHOWING
4CRACKS SHOWINGCRACKS SHOWINGDECLINING
3CRACKS SHOWINGCRACKS SHOWINGDECLININGDECLINING
2NOT YETCRACKS SHOWINGDECLININGDECLININGDECLINING
1NOT YETCRACKS SHOWINGDECLININGDECLININGDECLININGDECLINING
0NOT YETCRACKS SHOWINGDECLININGDECLININGDECLININGDECLININGCLIFF

The persistence rule

A computed verdict does not publish immediately. Moving up the severity ladder takes 2 consecutive snapshots. Moving back down takes 3.

The two counters. A verdict moves up the severity ladder after 2 consecutive snapshots computing the harsher reading, and moves back down only after 3. to escalate2 consecutiveto walk back3 consecutive
Being slow to declare a decline guards against noise. Being slower to withdraw one guards against wishful reading. The counter is consecutive, so a single snapshot that agrees with the published verdict empties it and the next disagreement starts again from one.

This rule exists because of a real event. Alonso's strikeout rate by completed month in 2026 ran 23.0% in April, 21.0% in May, then 28.0% in June, per the MLB Stats API. The band in force at the time was the absolute one, and 28.0% was exactly its DECLINING cut point. On that month alone the scorecard would have read one signal DECLINING and this site would have published CRACKS SHOWING. It came back to 24.0% in July. Without persistence the site would have declared a decline in June and withdrawn it in July, in its first season, on noise. August is still being played and is left out rather than quoted as a month that is finished.

That absolute line is gone. A25 replaced it with the derived band in the table above, where WATCH begins at 25.8% and DECLINING at 28.7%, so the same June now reads WATCH: one W, no D, and the verdict does not move. Both rules were needed and they fix different things. Persistence stops a verdict flickering on one month. The derived band stops a hitter being flagged for a level he has carried his whole career. Keeping this example against the rule that was actually in force is the point, because the site is meant to be checkable against what it did at the time.


3. The durability model

Injury is the most common cliff trigger in the comp set, so demonstrated durability shifts probability away from the downside. It shifts the odds between scenarios. It does not change what any scenario is worth.

Let t run from 0 (average slugger health) to 1 (iron man).

How demonstrated durability shifts the odds between the three scenarios. As the durability dial runs from 0 to 1, the early-cliff share falls from 32.0% to 14.0% and the Lasted-pace share rises from 14.0% to 26.0%. The table below carries the same figures. formula defaultEarly cliffZiPS glide pathLasted pace050100t = 0t = 1
Durability moves probability between the three scenarios and never changes what a scenario is worth. The three shares always sum to 100, which is why they are drawn as one stack rather than three lines. Scorecard and comp track are held neutral here so the dial's own effect is visible; at the formula's default of t = 0.8 the split is 23.6% Lasted pace, 58.8% ZiPS glide path and 17.6% early cliff, at 154 games a year. The site does not publish at that default. Today's snapshot implies t = 1.00, or 162 games a year, which moves the same three shares to 26.0%, 60.0% and 14.0%.
Show the diagram as a table
The same curve at five points on the dial, computed from the same function.
tgames/yearLasted paceZiPS glide pathEarly cliff
0.0012014.0%54.0%32.0%
0.2513117.0%55.5%27.5%
0.5014120.0%57.0%23.0%
0.7515223.0%58.5%18.5%
1.0016226.0%60.0%14.0%
P(downside) = 32 - 18t - 22 x lean      floored at 3
P(upside)   = 14 + 12t + 30 x lean
P(base)     = 100 - P(downside) - P(upside)

lean        = (scorecard - 0.5) + (comp track - 0.5)
games/year  = 120 + 42t

Two readings move the odds besides durability, and neither changes what a scenario is worth. Scorecard is the six signals scored 1 for STRONG, 0.5 for WATCH, 0 for DECLINING. Comp track is where his production sits between the two group curves at his age: 1 on the Lasted curve, 0 on the Collapsed one.

Both are anchored at 0.5, so a reading that says nothing leaves the durability-only formula exactly as it was first published. The base case stays the ZiPS glide path. This site departs from the market's view only as far as its own measurements disagree with it.

The odds as published, and as the current evidence sets them
ReadingLasted paceZiPS glideEarly cliff
Durability only, both readings neutral26.0%60.0%14.0%
Today: scorecard 1.00, track 1.0056.0%41.0%3.0%

At t = 1.00, about 162 games a year. That is the reading the newest snapshot implies, not a fixed default: it moves with his availability.

Why the comp track is trusted here and would not be at 26: ages 30 to 31, career seasons 7 to 8, is exactly where the two groups separate. The collapsed composite has already fallen to 92 by career season 8 while the lasted composite is still at 137. A reading taken at that moment discriminates; the same reading taken five years earlier does not. The limit is that both groups were selected, so the gap between the curves is probably wider than a full population would show, which inflates confidence near the extremes. That is why cliff risk is floored at 3% and never allowed to vanish.


4. The contract valuation

A win costs more every year, so a fixed salary buys more wins late in a deal than early. That is why higher inflation helps a long contract.

$/WAR(year)     = 8.5 x (1 + g) ^ (year - 2024)
scenario value  = sum over years of WAR(year) x $/WAR(year)
expected value  = sum over scenarios of P(scenario) x value(scenario)
surplus         = expected value - guarantee
How the three scenario values become one expected value and then a gap against the guarantee. Each scenario is drawn at what it is worth, with the share the odds buy filled in: Lasted pace $200.6M at 56.0%, ZiPS glide path $118.5M at 41.0%, Early cliff $76.3M at 3.0%. The three filled shares total $163.2M, against a $155M guarantee, a gap of +$8.2M. The table below carries the same figures. guarantee $155MLasted paceillustrative$200.6M x 56.0% = $112.33MZiPS glide pathpublished$118.5M x 41.0% = $48.58MEarly cliffillustrative$76.3M x 3.0% = $2.29MExpected value$112.33M + $48.58M + $2.29M = $163.2M+$8.2M against the guarantee$0M$50M$100M$150M$200M
The odds move value between scenarios and never change what a scenario is worth, which is why every bar keeps its full length and only the filled share changes. The three filled shares are the same ink as the expected-value bar below them, rearranged rather than recalculated. Drawn at the assumptions the published verdict uses, so the total here is the fair value The Contract reports: the 2026 column anchored on what has actually been banked rather than on what was projected for it, and the odds weighted by the scorecard and comp-track readings from section 3. Every figure is computed from the same functions, not restated from the prose beside it.
Show the diagram as a table
The same chain as arithmetic. The contributions are printed to two decimals because the total is their sum, so the column adds up on screen.
ScenarioBasisValueOddsContribution
Lasted paceillustrative$200.6M56.0%$112.33M
ZiPS glide pathpublished$118.5M41.0%$48.58M
Early cliffillustrative$76.3M3.0%$2.29M
Expected value$163.2M
Guarantee$155M
Gap+$8.2M

The two inputs the diagram uses

The formula above is the whole of it, but a formula is only as good as what is fed into it, and this site feeds it two things that are worth stating outright rather than leaving inside the code.

The season in progress is anchored on what has happened. A projection published last December is a worse estimate of 2026 than 2026 itself is. So the first year of the path is replaced by the WAR actually banked, plus the projected rate applied to whatever is left of the season. That leaves it counted as 129 of 162 games played. Later years are untouched, because nothing has happened in them yet. Without this the model would still be pricing the current season off a forecast that has already been overtaken.

The odds are the ones section 3 arrives at, weighted by the scorecard and the comp track, not the durability-only figures. Those two readings are the site's own measurement of whether decline has begun, and declining to use them here would mean publishing a valuation that ignores the evidence the rest of the site is built on.

Both are why the figure above differs from what the bare formula returns on the unanchored paths at durability alone. That version prices this deal near $30M under water, and it is not a number this site publishes anywhere, because it answers the question with two of the site's own measurements switched off.

The base price is $8.5M per WAR in 2024. The growth rate g is adjustable from 0 to 10 percent, defaulting to 6, which sits inside a historical market range of roughly 5 to 7.

Neither of those two numbers is sourced. The cost of a win and its growth rate both come from this project's own working document, which states them without saying where they came from, and every dollar figure on this site is computed from them. They are the weakest evidence here and they carry the most weight. The Contract lists them beside the projections, with the same caveat.

Grade bands

  • >= +$20M STRONG VALUE
  • >= $0M FAIR VALUE
  • >= $-15M MARKET RISK
  • below that OVERPAY RISK

WAR paths

The base path is the published ZiPS projection. Upside and downside are comp-study estimates around it, and are not projections in the same sense.

Orioles 2026-30, $155M guaranteed, 5 yrs · ages 31-35
Scenario20262027202820292030
upside4.24.03.83.43.0
base3.62.92.01.40.7
downside2.81.91.00.40.0
Mets offer (declined, 2023), $158M offered, 7 yrs · ages 29-35
Scenario2024202520262027202820292030
upside3.43.94.24.03.83.43.0
base3.43.93.62.92.01.40.7
downside3.43.92.81.91.00.40.0

The Mets offer covers 2024 to 2030, and its first two seasons actually happened. Those are locked to the same values across all three scenarios, because a scenario cannot change the past.


5. What a club should have paid

The valuation above answers what a win costs. This answers what a player is worth to a particular club, which is a different question. Fair value is the guarantee at which the club breaks even: offer less and it captures surplus, offer more and it is paying for wins it does not expect to get.

fair value = sum over scenarios of P(scenario) x production value(scenario)
gap        = fair value - guarantee        negative is an overpay

A guarantee is also a revealed opinion. Holding cliff risk where the evidence puts it, the price a club actually paid can be solved backwards for the confidence in the Lasted pace that would have made it indifferent. Comparing that against the confidence the measurements support is the honest form of the question: not "our number differs from theirs" but "here is what they had to believe, and here is what can be defended".

Two things deliberately left out

The cost of a win is derived by dividing real contracts by real WAR, so it already contains everything the market prices on average. Both of these were in an earlier version of this model and both were charging twice.

No incumbent is subtracted
WAR is Wins Above Replacement, so the replacement baseline is already inside the metric, and the cost of a win prices wins relative to that same baseline. Subtracting a notional incumbent charges for the baseline a second time. The expected alternative at a positionis replacement level: for every one-win stopgap there is a negative-win one, or one who is hurt in May and replaced by someone worse. Only a committed, genuinely good player left with nowhere to play is real, and that is a fact about a roster rather than a default. It ships at zero.
Leverage is a deviation, not an absolute
The same wins are worth more to a club near the postseason cut line, because the odds curve is steepest there. But clubs that sign free agents are mostly contenders, so that leverage is already inside the market rate. Only the difference from a typical signing club counts, taken here as 85 wins.

What leverage is actually worth

Differenced properly it is a small term and usually a negative one. The odds curve is symmetric, so the gain from a few wins peaks near 86.5 wins and the reference sits just below it. A club has to be unusually placed to beat a typical signer.

Leverage over three seasons adding 3.0, 2.5 and 2.0 marginal wins, against a $25M berth
Club baselineLeverage
65 wins-$10.90M
75 wins-$8.82M
85 wins+$0.00M
87 wins+$0.49M
92 wins-$3.31M
100 wins-$9.50M

Both club adjustments ship zero-effect, so the default view is the plain cost-of-a-win calculation and every departure from it is something a reader set on The Contract. That is deliberate: this site has no sourced figure for either club's baseline wins, its alternative at first base, or what a berth is worth to it, and inventing one would be the kind of number this page exists to refuse.


6. Freshness

The last-updated stamp is derived from the newest snapshot in the log. It is never a field anyone maintains by hand, because a field maintained by hand is a field that can be wrong.

The build warns past 140 days. Past 190 days the site tells you how old its data is rather than presenting it as current.


7. Where the numbers come from

Each update pulls from several sources.

Baseball Savant
avgEV, hardHitPct, xwOBA, kPct, bbPct, pullPct, oppoPct, ranks.avgEV
FanGraphs
wrcPlus, war, ranks.wrcPlus, zipsWar
ESPN
games, teamGames, hr
Baseball-Reference
streakGames, streakActive, ilDaysThisSeason, ranks.streakGames
MLB Stats API
monthly kPct, for the persistence example on the Formulas page

8. What this does not prove

The limits of the evidence belong in the same place as the evidence, at the same size.

  • The base scenario is the published ZiPS projection from December 2025. The upside and downside WAR paths, and the group aging curves, are illustrative estimates derived from a small and deliberately selected comp sample. They are not a controlled study.
  • Statcast bat speed data exists only from 2024 onward. Historical bat-speed decay is inferred from proxies: exit velocity, slugging, and strikeout trends.
  • Nelson Cruz and David Ortiz aged as designated hitters. The cleanest right-handed first-base comps in the set are Konerko and Abreu, one graceful and one cliff. Miguel Cabrera is not one of the fourteen. Since A33 he is one of the 164 the inclusion rule admits, so he is drawn on the curve chart and counted in the quartile bands, which is a change from what this caveat used to say.
  • Career-season alignment is confounded by debut age and era. It complements the calendar-age view rather than replacing it.
  • The 2023 shift restrictions reduce a headwind that damaged several of the pull-heavy comps, so their decline curves may overstate the risk facing a pull-heavy hitter today.
  • Alonso figures for 2026 are partial-season. Anything drawn from them is provisional and is marked as such.
  • The two curated groups were chosen to illustrate two archetypes, not sampled at random. That bias was assumed to make the two paths look more distinct than a full population would; measured against the 164 the rule admits, it ran the other way. The curated pair spans 6.9 to 17.0 wins and the quartiles of the pool span 0.4 to 22.6, so the hand-picked set understated both ends.