Vintage baseball research graphic for Putting WAR Into Practice, a collector-focused guide to comparing hitters using Wins Above Replacement.

Putting WAR Into Practice: How Collectors Can Compare Hitters

Part 1 of the Putting WAR Into Practice Series 

In our first WAR post, What Is WAR in Baseball? A Collector-Friendly Guide, we broke down what the statistic measures and why collectors may find it useful. 

This time, we are moving past the definition and into the fun part: putting the number to work. 

In this post, we will take a look at Baseball Reference's top 10 overall WAR leaders from the 2025 season. We will then break the numbers down using actual questions and comparisons we made when first seeing this list. 

The 2025 WAR Leaderboard

Rank Player 2025 bWAR
1 Aaron Judge 9.7
2 Cristopher Sánchez 8.1
3 Paul Skenes 8.0
4 Shohei Ohtani 7.7
5 Cal Raleigh 7.2
6 Bobby Witt Jr. 7.1
7 Geraldo Perdomo 7.1
8 Julio Rodríguez 6.6
9 Tarik Skubal 6.4
10 Matt Olson 6.2

bWAR = Baseball-Reference Wins Above Replacement. Rankings reflect overall 2025 regular-season WAR.

For readers seeing a WAR leaderboard for the first time, some of the results may seem surprising. How did Aaron Judge finish so far ahead of Cal Raleigh despite Raleigh leading in home runs and RBIs? How did Bobby Witt Jr. and Geraldo Perdomo arrive at the exact same WAR total? And why did Raleigh finish with more position-player WAR than Shohei Ohtani?

Those are exactly the kinds of questions that make WAR useful. The metric gives us a common framework for comparing players, but its greatest value may be how often it encourages us to look beyond the final number.

To show what we mean, we will walk through those three comparisons, examine how each player built his WAR total, and consider what collectors can learn from the differences.


1) Aaron Judge vs. Cal Raleigh: 2025 Regular Season - How did Judge finish so much higher than Raleigh?

The first thing that stood out to us from looking at this chart was the difference between Judge (9.7 bWAR) and Raleigh (7.2 bWAR). Sure, we knew that Judge hit over .400 for the first 2 months of the year and still had a .396 average in Mid June. But Raleigh ultimately finished with more home runs and RBIs, and the two were widely viewed as the leading candidates for the American League MVP Award. So why was the gap in their WAR totals so large?

The most obvious place to begin is with their offensive statistics. Below are their regular season totals for 2025:

Player G R HR RBI BB SO AVG OBP SLG OPS OPS+ bWAR
Aaron Judge 152 137 53 114 124 160 .331 .457 .688 1.144 213 9.7
Cal Raleigh 159 110 60 125 97 188 .247 .359 .589 .948 168 7.2

G = Games; R = Runs; HR = Home Runs; RBI = Runs Batted In; BB = Walks; SO = Strikeouts; AVG = Batting Average; OBP = On-Base Percentage; SLG = Slugging Percentage; OPS = On-Base Plus Slugging; OPS+ = Adjusted OPS; bWAR = Baseball-Reference WAR.

First read of the chart: Both players clearly had incredible seasons. They appeared in nearly every game, exceeded 50 home runs, and reached triple digits in both runs and RBIs. Raleigh’s batting average was substantially lower, but his record-setting power and run production appeared to keep the comparison close.

That is also how many fans naturally read a stat line. We first look at games played for context, then move quickly to home runs, RBIs, and batting average. Those statistics are popular for good reason: they can tell us a great deal about a player’s season.

The problem begins when we use them to fill in the entire picture. Once assumptions start replacing information, we risk misunderstanding how much value a player actually created. In this case, WAR immediately alerted us that there was more separation between Judge and Raleigh than the headline statistics suggested.

Now that we have identified our natural bias toward home runs—and, by extension, RBIs—let’s temporarily remove those categories from the comparison.

Judge finished with a significantly higher batting average, hitting .331 compared with Raleigh’s .247. To put that difference into perspective, Judge recorded 32 more hits despite having 55 fewer official at-bats. He also played seven fewer games.

That immediately tells us that Judge was producing hits at a much higher rate. But batting average is only the beginning of the difference between their offensive seasons.

Judge also posted the higher on-base percentage (.457 to .359), drew more walks (124 to 97), and finished well ahead in slugging percentage (.688 to .589), OPS (1.144 to .948), and OPS+ (213 to 168).

Those advantages helped Judge score 27 more runs than Raleigh over the course of the season. While runs scored are influenced by the teammates surrounding a player, Judge still had to put himself in position to be driven home. His .457 on-base percentage shows that he did so at a substantially higher rate than Raleigh.

This is where WAR becomes especially useful to collectors. If we looked only at home runs and RBIs late in the season, it would have been easy to conclude that Raleigh was leading the race and Judge needed to close the gap. However, WAR was already telling a different story. By accounting for a broader range of contributions, it estimated that Judge had already created substantially more overall value, while Raleigh was the one trying to close the gap.

The lesson is not that one statistic should replace another. It is that WAR can help us see the full shape of a season before we decide how much weight to give the headline numbers—or how we view the player behind the cards.

Takeaway: Headline statistics can be misleading. When two players appear close in categories such as home runs and RBIs, a meaningful gap in WAR can alert us that there is more to the story. The next step is determining what created that gap.

2) Bobby Witt Jr. vs. Geraldo Perdomo: 2025 Regular Season - Was Perdomo really as valuable as Bobby Witt Jr. last season?

Bobby Witt Jr. and Geraldo Perdomo both finished the 2025 season with 7.1 Baseball-Reference WAR.

Admittedly, this comparison caught our attention as Arizona baseball fans. We watched Perdomo play throughout the season and knew he was having a phenomenal year. But was his total contribution really comparable with Bobby Witt Jr.? And should that cause collectors to reconsider where Perdomo fits in the card market?

Because both players were shortstops, neither gained an advantage from playing a more valuable position. Defense still could have separated them, but Baseball-Reference credited each player with three fielding runs. That allows us to focus more closely on how their offensive and baserunning contributions differed.

We will start with their 2025 batting statistics:

Player G PA H HR RBI BB SO SB AVG OBP SLG OPS OPS+ bWAR
Bobby Witt Jr. 157 687 184 23 88 49 125 38 .295 .351 .501 .852 137 7.1
Geraldo Perdomo 161 720 173 20 100 94 83 27 .290 .389 .462 .851 137 7.1

G = Games; PA = Plate Appearances; H = Hits; HR = Home Runs; RBI = Runs Batted In; BB = Walks; SO = Strikeouts; SB = Stolen Bases; AVG = Batting Average; OBP = On-Base Percentage; SLG = Slugging Percentage; OPS = On-Base Plus Slugging; OPS+ = Adjusted OPS; bWAR = Baseball-Reference WAR.

At first glance, the two seasons look remarkably similar. Witt hit .295 with 23 home runs, 88 RBIs, and an .852 OPS. Perdomo hit .290 with 20 home runs, 100 RBIs, and an .851 OPS. They also finished with the exact same 137 OPS+.

The similarities become more interesting once we look at how they produced those results.

Witt recorded 11 more hits, three more home runs, a higher slugging percentage, and 11 more stolen bases. Perdomo, however, drew nearly twice as many walks and struck out 42 fewer times. That difference in plate discipline helped Perdomo reach base at a .389 rate compared with Witt’s .351.

According to Baseball-Reference, Perdomo produced 33 batting runs compared with 27 for Witt. So how did Witt close the gap and finish with the same overall WAR?

The answer was baserunning.

How Their 7.1 bWAR Was Built

Player Batting Runs Baserunning Runs Fielding Runs Position Runs bWAR
Bobby Witt Jr. 27 6 3 9 7.1
Geraldo Perdomo 33 3 3 10 7.1

Batting, baserunning, fielding, and positional runs are Baseball-Reference estimates measured relative to league and positional context.

Baseball-Reference credited Witt with six baserunning runs and Perdomo with three. Witt’s 11 additional stolen bases help illustrate that advantage, but the calculation goes beyond steals alone. It also considers caught stealing, advancement on hits and outs, double-play avoidance, and other ways runners gain or lose value.

Perdomo held an estimated six-run advantage through batting, while Witt gained three runs back through baserunning. Their fielding estimates were identical, and their positional adjustments were nearly the same. After Baseball-Reference incorporated the remaining components of the calculation and converted those estimated runs into wins, both players finished at 7.1 WAR.

This is another reason collectors should avoid treating WAR like a ranking that ends the discussion. Witt and Perdomo finished with the same total, but that does not mean they were interchangeable players. It means Baseball-Reference estimated that they created similar overall value through different combinations of skills.

WAR gives us the destination. The individual statistics show us the path each player took to get there.

For collectors, this comparison also demonstrates why a strong WAR total should prompt additional research rather than an immediate purchase.

Perdomo and Witt created similar estimated on-field value in 2025, but their card markets were built on very different foundations. Witt entered the season with a longer track record of star-level production, stronger national recognition, greater hobby visibility, and an established collector following. His combination of power, speed, prospect pedigree, personality, and market presence had already positioned him as one of the faces of the sport.

Perdomo’s 2025 season does not mean his cards should suddenly command Witt-level prices. It does suggest that his performance may deserve more attention than his hobby profile currently receives.

The next question is whether he can sustain it. Witt has established a multi-year record as an elite player, while Perdomo must show that his breakout was not an isolated season. Continued production could strengthen his recognition among collectors, but consistency, demand, card availability, and broader star appeal will still shape his market.

Takeaway: Two players can finish with the same WAR while creating value in different ways. Equal WAR does not mean equal players, equal career outlooks, or equal card markets—it tells us where deeper research should begin.

3) Aaron Judge, Shohei Ohtani, and Cal Raleigh: How Does Ohtani’s Two-Way Status Affect the Equation?

At the risk of stating the obvious here, Shohei Ohtani presents a unique challenge for almost any player-evaluation system.

Since we already compared Judge vs Raleigh, we can again use their stats as a point of reference alongside Ohtani's 2025 regular season totals:

Player Primary Role G R H HR RBI BB SO SB AVG OBP SLG OPS OPS+ Position-Player bWAR Total bWAR
Aaron Judge OF / DH 152 137 179 53 114 124 160 12 .331 .457 .688 1.144 213 9.7 9.7
Shohei Ohtani DH / Pitcher 158 146 172 55 102 109 187 20 .282 .392 .622 1.014 179 6.6 7.7
Cal Raleigh Catcher / DH 159 110 147 60 125 97 188 14 .247 .359 .589 .948 168 7.2 7.2

G = Games; R = Runs; H = Hits; HR = Home Runs; RBI = Runs Batted In; BB = Walks; SO = Strikeouts; SB = Stolen Bases; AVG = Batting Average; OBP = On-Base Percentage; SLG = Slugging Percentage; OPS = On-Base Plus Slugging; OPS+ = Adjusted OPS; bWAR = Baseball-Reference WAR.

Position-player bWAR reflects hitting, baserunning, fielding, positional adjustment, and playing time. Ohtani’s 7.7 total bWAR also includes the value he created as a pitcher.

What is interesting here is that Ohtani’s offensive statistics more closely resemble Judge’s than Raleigh’s, yet his final WAR total was much closer to Raleigh’s.

Ohtani finished with 55 home runs, a 1.014 OPS, and a 179 OPS+. Those numbers were comfortably ahead of Raleigh’s .948 OPS and 168 OPS+, but Raleigh still finished with the higher position-player WAR: 7.2 compared with Ohtani’s 6.6.

The difference begins with position.

When Ohtani was not pitching, he spent almost all of his time at designated hitter. Because a designated hitter does not contribute defensively, WAR applies a negative positional adjustment to reflect the fact that offensive production is generally easier to replace at that position.

Raleigh occupied the opposite end of the positional spectrum. Catcher is one of the most demanding positions in baseball, and WAR gives players credit for handling that responsibility. Raleigh was not only producing 60 home runs at the plate; he was doing so while playing a position that carries substantial defensive and positional value.

That helps explain how Raleigh finished with the higher position-player WAR despite Ohtani being the more productive hitter overall.

Ohtani’s two-way status then changes the comparison again. His 6.6 position-player WAR accounts only for the value he created as a hitter, baserunner, and designated hitter. Once his pitching contribution is included, his total rises to 7.7 WAR, moving him back ahead of Raleigh.

This is what makes Ohtani such an unusual case. His position as a designated hitter limits his WAR on one side of the calculation, while his ability to pitch adds value through an entirely separate path.

For collectors, this comparison demonstrates why WAR cannot be read as an offensive leaderboard. Ohtani was the stronger hitter overall, Raleigh finished with more position-player WAR in part because of the defensive and positional value associated with catching, and Ohtani ultimately finished ahead by adding value on the mound.

The final numbers bring those contributions together, but the breakdown explains what made each season valuable. 

What These Comparisons Tell Us

Taken together, these examples show why WAR is most useful when it leads us back to the individual statistics.

Judge and Raleigh showed that headline statistics can hide a larger difference in total performance. Witt and Perdomo demonstrated that identical WAR totals can be built through different combinations of skills. Ohtani’s comparison with Judge and Raleigh added positional context, showing why WAR should not be read as an offensive leaderboard.

In each case, WAR gave us a reason to look closer. It did not provide the entire explanation by itself.

Takeaway: WAR tells us that value was created and that a player’s numbers may deserve a closer look. The breakdown tells us how that value was created—and whether it matches the way collectors currently view the player.

How We Use WAR at Arthur’s

If there is one thing we hope you take away from this article, it is that we view WAR as a research tool.

We often begin with the league leaders, then move through the broader rankings looking for players whose totals are higher or lower than we expected. Those surprises are frequently where the metric becomes most useful.

When WAR challenges our initial impression, we revisit the player’s statistics, position, career trajectory, recent performance, and broader role. From there, we compare what we learned with the card market.

This process helps us identify players worth following, reevaluate players already in our collection, and better understand whether hobby perception matches on-field performance.

It Helps Us Understand What Is Driving Performance

The Judge–Raleigh comparison helps us identify which parts of a player’s performance deserve the closest attention going forward.

Judge’s production was distributed across several offensive categories, while Raleigh’s public profile was tied more closely to his historic home-run total. A player whose value is supported by several strengths may be better positioned to remain productive when one category declines. A player whose reputation is tied heavily to one standout skill may experience a larger shift in hobby interest if that skill falls back.

That does not tell us whether to buy or avoid Raleigh’s cards. It tells us what to monitor: whether the power returns, whether his broader offensive production remains strong, and how collectors respond if home runs no longer dominate the conversation.

It Helps Us Find Players the Market May Be Overlooking

The Witt–Perdomo comparison illustrates another use.

Perdomo’s 7.1 WAR did not tell us that his cards should be valued like Bobby Witt Jr.’s. It told us that his on-field performance deserved more attention than his hobby profile may have suggested.

The next step would be to look for signs that his performance was becoming sustainable and that the market was beginning to recognize it. Was he repeating the production? Was he receiving more national attention and recognition from collectors outside Arizona? Was activity in his card market increasing?

Witt already had an established track record, elite prospect pedigree, national recognition, hobby visibility, and what we think of as superstar presence. Perdomo’s breakout narrowed the performance gap, but it did not erase the difference in collector demand.

WAR helped identify the player. The rest of the research would determine whether the collecting opportunity was real.

It Helps Us Challenge Our Assumptions

WAR also helps keep us grounded when our perception of a player begins to outpace the details of his actual role.

Shohei Ohtani is a good example. Because he contributes as both a hitter and a pitcher, it is easy to think of him as doing everything on the field. But when he is not pitching, he is generally limited to designated hitter and does not add defensive value in the way an everyday outfielder, shortstop, or catcher would.

We do not believe that meaningfully hurts Ohtani’s card value. His two-way ability, historic production, international popularity, and place in baseball history extend well beyond a positional adjustment. Still, WAR reminds us that even the game’s most extraordinary players have tradeoffs in how their value is created.

The same comparison gave us a greater appreciation for Raleigh’s season. Catching carries an enormous physical burden over the course of a year, and Raleigh handled that responsibility across nearly a full season while hitting 60 home runs.

In this case, WAR did more than compare players. It corrected part of our own perspective.

Final Collector Takeaways

WAR is not a magic number that tells us who the best player is, who should win MVP, or whose cards are worth collecting.

Its real value lies in helping us break down performance and compare the many different ways players contribute.

These examples show how WAR can help collectors look beyond home runs, RBIs, and batting average to better understand the complete value a hitter created. Just as importantly, they show why the final number should lead us back to the individual statistics rather than replace them.

Coming Up Next

Position players are only one side of the calculation. Pitcher WAR introduces a different set of questions, including workload, run prevention, defensive support, ballpark effects, and the differences between Baseball-Reference and FanGraphs.

How can a starting pitcher who takes the mound only once every five days finish ahead of hitters who appear nearly every day? And how did Cristopher Sánchez finish ahead of Paul Skenes?

We will examine those questions in our next WAR feature, which will focus specifically on pitchers.

Join the Conversation

How much weight do you give WAR when evaluating the player behind the card?

Let us know which comparison surprised you most, whether you think the market overlooks any of these players, or which hitters you would like us to compare next.

Back to blog

Leave a comment