Why a Single Digit Can’t Hold a 300-Hour RPG

I’ve been reading, writing, and arguing about game reviews for two decades. In that stretch, I’ve watched a sprawling 300-hour RPG get flattened into an 8.7, a brilliant indie oddity dismissed with a 6.5, and a technically broken blockbuster handed a 9 because the box had the right logo. The number sits there, bold and definitive, pretending to wrap up everything that matters. It doesn’t. The single-digit review score isn’t just an oversimplification—it’s a structural failure that distorts how we talk about games, how we buy them, and how they get made.

The Illusion of Precision

When a site stamps a 7.8 on a game, it borrows the language of measurement. That decimal whispers, “we ran the tests, we crunched the data.” But nobody crunched anything. I’ve been in editorial meetings where a score climbed from 7.5 to 8.0 because the art director liked the color palette. I’ve watched a game shed half a point because the reviewer was hungry during the final boss fight. The number projects objectivity where none exists.

Look at Death Stranding. At launch, scores ran from 3.5 to 10. Same game, same mechanics, same lonely trudges across a shattered America. One critic found it meditative and profound; another called it a glorified fetch quest. Both takes are legitimate. The game is a strange, ambitious experiment that some people will adore and others will loathe. But the score system shoves that reality into a false binary: masterpiece or failure. The truth gets buried under the weight of a number.

How Scores Strip Away Context

A review score peels off every qualifier. It won’t tell you that Cyberpunk 2077 got a 9 from a critic running it on a high-end PC before the console versions imploded. It won’t mention that the 7.5 for Days Gone came from someone who only saw the first ten hours, missing the narrative payoff that redeems its sluggish opening. The number floats free of its origins, hardening into a permanent brand on the game’s reputation.

I remember reading a 9.5 review for Red Dead Redemption 2 and feeling genuinely confused when I finally played it. The controls felt like steering a drunk marionette through molasses. The mission design was so rigid that stepping two feet off the prescribed path triggered a fail state. The world was stunning, the writing exceptional, but the moment-to-moment interaction was frequently miserable. That 9.5 didn’t prepare me for any of that. It just told me the game was excellent, full stop. A number can’t communicate that a game is simultaneously a technical marvel and a mechanical frustration.

Person holding a game controller in a dimly lit room, reflecting the solitary nature of forming a review opinion
A single number can’t capture the hours of solitary experience that shape a reviewer’s perspective.

The Metacritic Problem

Metacritic didn’t invent review scores, but it perfected their tyranny. By aggregating scores into a single weighted average, it manufactures the illusion of consensus. A game with an 85 Metascore is “great,” while a 74 is merely “mixed.” That eleven-point gap might represent the difference between a studio receiving a bonus and laying off staff. Publishers literally tie developer compensation to Metacritic thresholds. Obsidian Entertainment missed a bonus for Fallout: New Vegas by one Metacritic point—84 instead of the contractually required 85. One point. That’s the difference between a reviewer who had a good lunch and one who didn’t.

Metacritic also flattens the diversity of critical voices. A thoughtful, detailed review from a small outlet gets reduced to a number and averaged in with scores from publications that may have entirely different standards. The site’s color-coded system—green for good, yellow for mixed, red for bad—trains readers to scan for colors instead of engaging with arguments. I’ve watched friends decide against buying a game solely because its Metacritic score was yellow. They never read a single review.

The 7–10 Scale Collapse

Most review outlets operate on a de facto 7–10 scale. Anything below 7 is treated as garbage. A 5 should mean average, but in practice, a 5 is a disaster. This inflation isn’t accidental—it’s structural. Publishers withhold early review copies from outlets that score too low. Advertising relationships create soft pressure to be generous. And readers themselves punish low scores with outrage, because they’ve already pre-ordered the game and need their purchase validated.

Consider Starfield. Pre-release hype was astronomical. When reviews landed with 7s from some major outlets, the backlash was immediate and vicious. A 7 is a good score. It means the game has substantial merits but also significant flaws. But in the current climate, a 7 reads as a condemnation. The discourse around that game wasn’t about its ambitious scope or its uneven execution—it was about whether it “deserved” an 8 or a 9. The number became the entire conversation.

Close-up of a gaming keyboard with colorful backlighting, representing the technical complexity of modern games
Modern games are layered systems of mechanics, narrative, and technology—none of which can be honestly reduced to a single digit.

What Scores Actually Measure

If we’re being honest, review scores measure three things: production values, brand expectations, and the reviewer’s personal tolerance for frustration. A game with high-budget cutscenes and a licensed soundtrack starts at an 8 before anyone touches the controller. A sequel to a beloved franchise gets a built-in cushion because the reviewer already likes the world. And a game that respects the player’s time—no grinding, no padding—gets a boost that has nothing to do with artistic merit.

Hades earned near-universal acclaim, and deservedly so. But part of its high scores came from the fact that it never wasted your time. You die, you’re back in the hub in seconds, new dialogue waiting. Compare that to Persona 5, which I love but which also features a tutorial that lasts roughly fifteen hours. If a reviewer docked Persona 5 for that, they’d be correct. But the score wouldn’t tell you that. It would just say 9.3, and you’d assume it’s flawless.

Genre Bias in Numerical Ratings

Review scores carry an unspoken genre tax. Strategy games, simulation games, and complex RPGs consistently score lower than action-adventure titles with comparable quality. Crusader Kings III is one of the most sophisticated strategy games ever made, a dynastic simulator of staggering depth. Its Metacritic score sits at 91. That’s excellent, but compare it to God of War Ragnarök at 94. Is Ragnarök really three points better, or does it simply benefit from being a cinematic action game that reviewers can finish in thirty hours and rate confidently?

The answer is obvious. Reviewers are human beings with deadlines. A game that takes sixty hours to understand will rarely get the same thorough evaluation as one that delivers its pleasures in the first five. The score doesn’t account for this. It pretends all genres compete on a level field.

The Player Score Counter-Mess

User scores are supposed to be the antidote—the voice of the people against the corrupt critic elite. In practice, they’re even worse. User scores on Metacritic and similar platforms are routinely bombed by coordinated campaigns. A game ships with a minor performance issue and gets 0s from people who haven’t played it. A game makes a political statement someone dislikes and gets review-bombed into oblivion. The user score for The Last of Us Part II dropped to 3.4 within hours of release, before anyone could have possibly finished its 25-hour campaign. That number tells you nothing about the game. It tells you about a culture war.

Steam’s binary thumbs-up/thumbs-down system is marginally better because it avoids the false precision of a 100-point scale. But it still reduces complex opinions to a single bit. A game can earn 95% positive reviews and still have fundamental problems that a significant minority finds game-breaking. The percentage hides the distribution.

Person sitting in front of multiple monitors displaying game analytics and charts
The obsession with scores and metrics turns layered experiences into data points on a chart.

What We Lose When We Reduce Games to Numbers

The most insidious effect of review scores is how they shape the games themselves. Developers know that certain features reliably boost Metacritic scores: high-fidelity graphics, cinematic set-pieces, emotional story beats that photograph well in trailers. Systems that are deep but hard to demonstrate—complex AI behaviors, emergent gameplay, subtle mechanical interactions—don’t move the needle. So they get cut. Budgets flow toward the score-boosters, and the medium narrows.

Look at the evolution of the Dragon Age series. Origins was a tactical RPG with pause-and-play combat, extensive party management, and branching dialogue trees that could lock you out of entire questlines. Inquisition streamlined everything into an action-oriented open-world checklist. The Veilguard reportedly strips away even more complexity. Each iteration chased broader appeal and higher review scores, and each lost something of what made the original distinctive. The numbers rewarded the changes. The art suffered.

Alternatives That Already Exist

Some outlets have abandoned scores entirely. Eurogamer switched to a system of Essential, Recommended, and Avoid badges in 2015. Kotaku stopped scoring reviews years ago. Rock Paper Shotgun never used scores. These publications force readers to actually engage with the text—to weigh the critic’s descriptions of what works and what doesn’t against their own preferences. A game described as “punishingly difficult but fair” might appeal to a Soulslike fan while warning off someone looking for a relaxing experience. A number can’t make that distinction.

Other outlets have experimented with multi-axis scoring. Instead of one number, you get separate ratings for gameplay, story, visuals, sound, and value. This is better—it at least acknowledges that a game can excel in one area while failing in another. But it still reduces each category to a number, and it still invites the reader to average them into a single figure in their head.

A Modest Proposal for Readers

If you’re still reading reviews that end with a score, here’s what I suggest: ignore the number. Read the text. Pay attention to the specific complaints and the specific praise. If a reviewer says the combat feels weightless, ask yourself whether that matters to you. If they praise the writing but criticize the pacing, consider your own tolerance for slow burns. The information you need is in the words, not the digit at the bottom of the page.

Better yet, find critics whose tastes you understand. Follow them over time. Learn what they value and what they dismiss. A reviewer who hates crafting systems will always give survival games a lower score than they deserve—for you, if you love crafting. Once you know a critic’s biases, their reviews become useful regardless of the number attached. The score becomes irrelevant; the perspective becomes everything.

Frequently Asked Questions

Why do review sites still use scores if they’re so flawed?

Because scores drive traffic. A number is easy to aggregate, easy to compare, and easy to argue about. Metacritic and similar sites have built entire business models around turning reviews into data points. Publications that drop scores often see an initial dip in readership, even if their long-term credibility improves. The economic incentive to keep scoring is powerful, even when editors know the system is broken.

Are user reviews on Steam more reliable than critic scores?

Steam reviews have different problems. The binary recommend/not recommend system avoids false precision, but it’s vulnerable to review bombing, meme reviews, and the fact that players with hundreds of hours often leave the most critical reviews—because they’re the ones who engaged deeply enough to see the flaws. A game with 95% positive reviews might still have a late-game balance issue that only the most dedicated players encounter. Read the actual reviews, not just the percentage.

What should I look for in a review instead of a score?

Look for specific descriptions of systems and experiences. A good review tells you what the game asks you to do, how it feels to do it, and what kind of player would enjoy it. Phrases like “the inventory management is tedious” or “the dialogue trees have genuine consequences” give you actionable information. A number just tells you whether the reviewer liked it, which may have nothing to do with whether you will.

Do review scores affect which games get made?

Absolutely. Publishers use Metacritic scores to determine bonuses, greenlight sequels, and allocate marketing budgets. A game that scores below 80 is often considered a commercial disappointment regardless of sales. This creates a powerful incentive for developers to design games that score well rather than games that are interesting. Safe, polished, focus-tested products get higher scores than risky, innovative ones. The number shapes the art.

The next time you see a big, bold score at the end of a review, remember what it actually represents: one person’s rough estimate, rounded to look authoritative, stripped of all the context that might actually help you decide whether to spend your money and your time. Games are too complex, too varied, and too personal to fit inside a single digit. The sooner we stop pretending otherwise, the better our conversations about this medium will become.