Polar Bear
Projections

you will still be fiddling with your [flipping] formulas

Pete Alonso to Mets president David Stearns, on how his Hall of Fame case would be judged

He was not wrong. This is the fiddling, published in full.

Everything below is rendered directly from the code that runs the site. There is one copy of each number, so what is published here and what the site computes with cannot drift apart. If a threshold changes, this page changes with it.


1. The decline thresholds

Six signals, each scored STRONG, WATCH or DECLINING from the latest snapshot. Nothing is stored as a judgment; the chips are computed from these numbers every build.

Decline thresholds for each of the six signals
SignalFieldStrongWatchDeclining
Exit velocityavgEV>= 93.592 to 93.4< 92
Hard-hit ratehardHitPct>= 5044 to 49.9< 44
ProductionwrcPlus>= 125105 to 124< 105
Plate disciplinebbPct>= 108 to 9.9< 8
Strikeout ratekPct< 2222 to 27.9>= 28
Durabilitygames, streak, ILstreak live, or >= 95%80% to 95%< 80%, or IL > 30d

Why these numbers

Exit velocity
The 92 mph line is the published early-warning threshold, not a new claim. Exit velocity is the best single predictor of retained power.
Hard-hit rate
Calibrated so the ZiPS base case crosses into WATCH late in the contract rather than at signing.
Production
Set so the published ZiPS glide path (112 OPS+ by 2030) reads WATCH around 2029, matching what the scenario model implies.
Plate discipline
Aging hitters who hold or improve walk rate age best; a collapsing walk rate tends to follow bat-speed loss.
Strikeout rate
The 28% line is the published early-warning threshold. The WATCH band starts at 22% because Alonso sits above what the graceful agers carried; this is a level judgment, not a trend judgment.
Durability
Injury is the most common cliff trigger in the comp set, so durability is scored on availability rather than on any rate stat. The 80-95% band without a live streak reads WATCH.

These cut points are calibrated, not derived. Exit velocity and strikeout rate use thresholds published before this site existed. The other four are set so the published ZiPS glide path crosses into WATCH late in the contract rather than at signing. They are absolute rather than age-adjusted, which means a 140 wRC+ reads the same at 35 as at 31 even though the first is the better result. That is a known limitation, kept for now because an auditable rule beats a clever one.


2. The verdict rule

Four verdicts, in order of severity. Which one the site publishes is decided by a rule written down before the answer became inconvenient, so a verdict change is never a matter of taste.

  1. 1NOT YETNo decline signal has broken down. The profile still matches the hitters who lasted.
  2. 2CRACKS SHOWINGDeterioration is visible but not decisive. Could be noise, could be the start.
  3. 3DECLININGThe answer to this site’s question has flipped. Multiple signals have broken down.
  4. 4CLIFFNot a gradual fade. Production and either durability or contact quality gone at once.

The rule

Let D be the number of signals reading DECLINING, and W the number reading WATCH.

CLIFF           if production is DECLINING
                and (durability or exit velocity) is DECLINING
DECLINING       else if D >= 2
CRACKS SHOWING  else if D == 1, or W >= 3
NOT YET         otherwise

The persistence rule

A computed verdict does not publish immediately. Moving up the severity ladder takes 2 consecutive snapshots. Moving back down takes 3.

The asymmetry is deliberate. Being slow to declare decline guards against noise, and being slower to withdraw it guards against wishful reading.

This rule exists because of a real event. Alonso's strikeout rate by month in 2026 ran 23.0% in April, 21.0% in May, then 28.0% in June, which is exactly the cut point above. On that month alone the scorecard would have read one signal DECLINING and this site would have published CRACKS SHOWING. It came back to 24.0% in July and 22.5% in August. Without persistence the site would have declared a decline in June and withdrawn it in July, in its first season, on noise.


3. The durability model

Injury is the most common cliff trigger in the comp set, so demonstrated durability shifts probability away from the downside. It shifts the odds between scenarios. It does not change what any scenario is worth.

Let t run from 0 (average slugger health) to 1 (iron man).

P(downside) = 32 - 18t - 22 x lean      floored at 3
P(upside)   = 14 + 12t + 30 x lean
P(base)     = 100 - P(downside) - P(upside)

lean        = (scorecard - 0.5) + (comp track - 0.5)
games/year  = 120 + 42t

Two readings move the odds besides durability, and neither changes what a scenario is worth. Scorecard is the six signals scored 1 for STRONG, 0.5 for WATCH, 0 for DECLINING. Comp track is where his production sits between the two group curves at his age: 1 on the Lasted curve, 0 on the Collapsed one.

Both are anchored at 0.5, so a reading that says nothing leaves the durability-only formula exactly as it was first published. The base case stays the ZiPS glide path. This site departs from the market's view only as far as its own measurements disagree with it.

The odds as published, and as the current evidence sets them
ReadingLasted paceZiPS glideEarly cliff
Durability only, both readings neutral23.6%58.8%17.6%
Today: scorecard 0.92, track 1.0051.1%45.9%3.0%

At t = 0.8, about 154 games a year.

Why the comp track is trusted here and would not be at 26: ages 30 to 31, career seasons 7 to 8, is exactly where the two groups separate. The collapsed composite has already fallen to 92 by career season 8 while the lasted composite is still at 137. A reading taken at that moment discriminates; the same reading taken five years earlier does not. The limit is that both groups were selected, so the gap between the curves is probably wider than a full population would show, which inflates confidence near the extremes. That is why cliff risk is floored at 3% and never allowed to vanish.


4. The contract valuation

A win costs more every year, so a fixed salary buys more wins late in a deal than early. That is why higher inflation helps a long contract.

$/WAR(year)     = 8.5 x (1 + g) ^ (year - 2024)
scenario value  = sum over years of WAR(year) x $/WAR(year)
expected value  = sum over scenarios of P(scenario) x value(scenario)
surplus         = expected value - guarantee

The base price is $8.5M per WAR in 2024. The growth rate g is adjustable from 0 to 10 percent, defaulting to 6, which sits inside the historical market range of roughly 5 to 7.

Grade bands

  • >= +$20M STRONG VALUE
  • >= $0M FAIR VALUE
  • >= $-15M MARKET RISK
  • below that OVERPAY RISK

WAR paths

The base path is the published ZiPS projection. Upside and downside are comp-study estimates around it, and are not projections in the same sense.

Orioles 2026-30, $155M guaranteed, 5 yrs · ages 31-35
Scenario20262027202820292030
upside4.24.03.83.43.0
base3.62.92.01.40.7
downside2.81.91.00.40.0
Mets offer (declined, 2023), $158M guaranteed, 7 yrs · ages 29-35
Scenario2024202520262027202820292030
upside3.43.94.24.03.83.43.0
base3.43.93.62.92.01.40.7
downside3.43.92.81.91.00.40.0

The Mets offer covers 2024 to 2030, and its first two seasons actually happened. Those are locked to the same values across all three scenarios, because a scenario cannot change the past.


5. What a club should have paid

The valuation above answers what a win costs. This answers what a player is worth to a particular club, which is a different question. Fair value is the guarantee at which the club breaks even: offer less and it captures surplus, offer more and it is paying for wins it does not expect to get.

fair value = sum over scenarios of P(scenario) x production value(scenario)
gap        = fair value - guarantee        negative is an overpay

A guarantee is also a revealed opinion. Holding cliff risk where the evidence puts it, the price a club actually paid can be solved backwards for the confidence in the Lasted pace that would have made it indifferent. Comparing that against the confidence the measurements support is the honest form of the question: not "our number differs from theirs" but "here is what they had to believe, and here is what can be defended".

Two things deliberately left out

The cost of a win is derived by dividing real contracts by real WAR, so it already contains everything the market prices on average. Both of these were in an earlier version of this model and both were charging twice.

No incumbent is subtracted
WAR is Wins Above Replacement, so the replacement baseline is already inside the metric, and the cost of a win prices wins relative to that same baseline. Subtracting a notional incumbent charges for the baseline a second time. The expected alternative at a positionis replacement level: for every one-win stopgap there is a negative-win one, or one who is hurt in May and replaced by someone worse. Only a committed, genuinely good player left with nowhere to play is real, and that is a fact about a roster rather than a default. It ships at zero.
Leverage is a deviation, not an absolute
The same wins are worth more to a club near the postseason cut line, because the odds curve is steepest there. But clubs that sign free agents are mostly contenders, so that leverage is already inside the market rate. Only the difference from a typical signing club counts, taken here as 85 wins.

What leverage is actually worth

Differenced properly it is a small term and usually a negative one. The odds curve is symmetric, so the gain from a few wins peaks near 86.5 wins and the reference sits just below it. A club has to be unusually placed to beat a typical signer.

Leverage over a five-year deal adding 3.0, 2.5 and 2.0 marginal wins, against a $25M berth
Club baselineLeverage
65 wins-$10.90M
75 wins-$8.82M
85 wins+$0.00M
87 wins+$0.49M
92 wins-$3.31M
100 wins-$9.50M

Both club adjustments ship zero-effect, so the default view is the plain cost-of-a-win calculation and every departure from it is something a reader set on The Contract. That is deliberate: this site has no sourced figure for either club's baseline wins, its alternative at first base, or what a berth is worth to it, and inventing one would be the kind of number this page exists to refuse.


6. Freshness

The last-updated stamp is derived from the newest snapshot in the log. It is never a field anyone maintains by hand, because a field maintained by hand is a field that can be wrong.

The build warns past 60 days. Past 90 days the site tells you how old its data is rather than presenting it as current.


7. Where the numbers come from

Each update pulls from several sources.

Baseball Savant
avgEV, hardHitPct, xwOBA, kPct, bbPct, pullPct, oppoPct, ranks.avgEV
FanGraphs
wrcPlus, war, ranks.wrcPlus, zipsWar
ESPN
games, teamGames, hr
Baseball-Reference
streakGames, streakActive, ilDaysThisSeason, ranks.streakGames

8. What this does not prove

The limits of the evidence belong in the same place as the evidence, at the same size.

  • The base scenario is the published ZiPS projection from December 2025. The upside and downside WAR paths, and the group aging curves, are illustrative estimates derived from a small and deliberately selected comp sample. They are not a controlled study.
  • Statcast bat speed data exists only from 2024 onward. Historical bat-speed decay is inferred from proxies: exit velocity, slugging, and strikeout trends.
  • Nelson Cruz and David Ortiz aged as designated hitters. The cleanest right-handed first-base comps are Konerko, Cabrera and Abreu, and those three split one graceful, one average-only, one cliff.
  • Career-season alignment is confounded by debut age and era. It complements the calendar-age view rather than replacing it.
  • The 2023 shift restrictions reduce a headwind that damaged several of the pull-heavy comps, so their decline curves may overstate the risk facing a pull-heavy hitter today.
  • Alonso figures for 2026 are partial-season. Anything drawn from them is provisional and is marked as such.
  • The comp groups were chosen to illustrate two archetypes, not sampled at random. Selection bias runs in the direction of making the two paths look more distinct than the full population would.