The Problem With Review Scores That Reduce Complex Games to Single Digits

I’ve been reviewing games for more than ten years, and I’ve watched the industry tie itself in knots over a single, reductive number. A 7.8. An 8.5. A 9.2. These digits are supposed to summarize dozens of hours of play, layered systems, narrative ambition, and technical execution. They don’t. They can’t. And yet we keep stamping them onto every release as if a game’s worth can be measured like a toaster’s energy rating. The real problem isn’t just that scores are subjective—it’s that they actively warp how we discuss, design, and remember games.

Close-up of a gaming controller with colorful LED lights
A controller, the starting point for experiences that can’t be reduced to a number.

The Origin of the Score: A Shortcut That Became a Crutch

Review scores started as a convenience. In the print magazine days, flipping through Electronic Gaming Monthly or GamePro, you’d spot a bold number and decide whether to drop $50. The score was a filter, not a final judgment. But somewhere along the way, that filter hardened into a verdict. Metacritic, launched in 2001, turned the aggregate score into a weapon. Publishers started tying bonuses to Metacritic thresholds—Obsidian famously missed a Fallout: New Vegas bonus by a single point, an 84 instead of an 85. One digit cost a studio millions. That should have sparked a reckoning. Instead, it normalized the idea that a game’s commercial fate could rest on a weighted average of opinions from outlets with wildly different standards.

Look at the scoring scales themselves. Some sites use a 10-point scale with decimal increments, as if that precision means something. Others use a 5-star system that mashes nuance into broad buckets. IGN’s 10-point scale once reserved a 10 for “masterpiece,” but the definition of masterpiece shifted so often the score lost all meaning. God of War (2018) got a 10, and the text praised its reinvention of a tired series. The Last of Us Part II got a 10, and the text wrestled with its divisive narrative. Same number, completely different justifications. The score tells you nothing about why a game matters—only that someone decided it does.

What a Number Erases: The Case of Disco Elysium

Take Disco Elysium. It launched with a Metacritic score of 91, a number that screams near-universal acclaim. But that 91 erases the fact that the game has no combat, that its core mechanic is internal dialogue, that it asks you to fail skill checks and live with the consequences. A player who buys it based on the score alone might expect a traditional RPG and bounce off the first hour of existential despair. The score doesn’t communicate that the game is a slow, literary burn—it just says “good.”

Worse, the score flattens the game’s actual flaws. Disco Elysium had performance issues at launch, a clunky inventory system, and a final act that some critics felt lost momentum. A 91 implies near-perfection, but the game is fascinating because of its rough edges, not despite them. When we reduce it to a number, we lose the texture that makes criticism valuable. We’re left with a hollow endorsement that serves marketing departments more than players.

Person playing a video game on a large screen in a dark room
The solitary, immersive experience of gaming defies a single-digit summary.

The 7/10 Trap: How Scores Kill Interesting Games

There’s a graveyard of games that scored in the 70s on Metacritic and were dismissed as “just okay.” But many of those games are more memorable than their 85+ counterparts because they took risks that didn’t land cleanly. Alpha Protocol sits at a 72. It’s a spy RPG with a dialogue system that forces you to choose a stance—aggressive, professional, suave—on a timer, and the story branches so aggressively that entire characters can vanish based on your choices. The shooting is mediocre. The stealth is inconsistent. But no other game has replicated its reactive narrative. A 72 tells you it’s flawed. It doesn’t tell you it’s one of the most ambitious conversation systems ever built.

Compare that to Assassin’s Creed Valhalla, which holds an 80 on Metacritic. It’s a polished, enormous, competent open-world game that does nothing particularly new. The score is higher because it’s safer, more technically stable, and fits a familiar template. The scoring system rewards games that avoid big swings. It punishes games that try something different and stumble. The result is an industry that learns the wrong lesson: don’t innovate if it might cost you five points.

The Aggregation Problem: Mixing Oil and Water

Metacritic’s methodology is a black box, but we know it weights outlets differently. A review from a major publication counts more than one from a niche site. That weighting assumes that a generalist outlet’s opinion is more valid for all players, which is absurd. A hardcore strategy fan doesn’t care what a mainstream reviewer thinks about Crusader Kings III—they want the perspective of someone who understands succession laws and casus belli. But the aggregate score buries those specialized voices under the weight of outlets reviewing for a hypothetical “average gamer.”

Even worse, Metacritic converts qualitative scores into quantitative ones. A 4/5 becomes an 80. A B+ becomes an 83. These conversions are arbitrary and inconsistent across publications. A “Recommended” badge from Eurogamer—which explicitly avoids scores—gets turned into a number anyway. The system forces everything into a single dimension, then claims that dimension is objective. It’s not. It’s a statistical fiction.

When Scores Become the Story: The Review Bombing Phenomenon

User scores on Metacritic have become a battleground for culture wars that have nothing to do with a game’s quality. The Last of Us Part II was flooded with 0/10 scores within hours of release, before anyone could have finished its 25-hour campaign. The score became a proxy for anger over narrative decisions, representation, and leaks. The actual game—its level design, its accessibility features, its technical polish—was irrelevant. The number was a weapon.

This isn’t an isolated case. Borderlands 3 was review-bombed over an Epic Games Store exclusivity deal. Warframe caught negative scores because of a Discord moderator controversy. In each case, the user score told you nothing about the game. But because Metacritic displays it alongside the critic score, it creates a false equivalence. A curious player sees a 5.7 user score next to a 94 critic score and assumes there’s a hidden flaw. There isn’t. There’s just a number that’s been hijacked.

Person holding a smartphone displaying colorful game graphics
User scores on mobile platforms often reflect outrage, not gameplay.

The Developer’s Dilemma: Designing for the Score

I’ve spoken with developers who admit they make design decisions with Metacritic in mind. One told me they cut a divisive ending because they feared it would drag the score below 80. Another said they padded their game with repetitive side content because “reviewers equate length with value.” When a score becomes a target, games become checklists. You get open worlds littered with towers to climb, camps to clear, and collectibles to gather—not because those activities are fun, but because they signal “content density” to a reviewer who has 40 hours to play before embargo lifts.

The most damaging effect is on difficulty. A game that’s too hard risks frustrating reviewers who are rushing to meet a deadline. So we get default difficulty settings that are trivial for experienced players, with the “real” game locked behind a menu option. Doom Eternal is a masterpiece of first-person combat design, but its default difficulty undersells the game’s rhythm. Many players never discover the weapon-switching ballet because the game doesn’t force them to learn it. The score doesn’t reflect that the default experience is a diluted version of the intended one.

The Alternative: Contextual Criticism Without a Number

Some outlets have abandoned scores entirely. Eurogamer uses a “Recommended” badge, with occasional “Essential” and “Avoid” tags. Rock Paper Shotgun writes verdicts in plain English. Kotaku doesn’t score. These approaches force the reader to engage with the text, to understand why a game is worth their time. They also free the critic to write honestly about a game’s flaws without worrying that a 7/10 will be misinterpreted as a pan.

But the industry resists. Publishers want scores for marketing. Aggregators want scores for traffic. And readers, conditioned by decades of habit, often scroll straight to the number. Breaking that cycle requires a cultural shift, not just an editorial one. It requires players to value their own preferences over a consensus. It requires them to ask, “What does this game do that I care about?” instead of “Is this game good?”

How to Read a Review in a Post-Score World

If you’re stuck with scored reviews, here’s how I recommend you use them. First, ignore the aggregate. Find one or two critics whose tastes align with yours and read their full text. If you love immersive sims, follow someone who dissected Prey (2017) with the same enthusiasm you felt. If you bounce off open-world fatigue, find a critic who called out Horizon Forbidden West for its map clutter. A critic’s history is more predictive than any number.

Second, read the last paragraph of a review first. That’s where the critic usually summarizes what the game feels like, not just what it scores. A verdict that says “this is a messy, brilliant game that will frustrate you as often as it delights you” tells you more than an 8.5 ever could. If that description intrigues you, read the rest. If it sounds exhausting, skip it—even if the score is high.

Third, look for specific complaints, not general ones. “The combat is clunky” is useless. “The dodge has a half-second input delay that makes fast enemies feel unfair” is actionable. That level of detail tells you whether the problem will bother you. A score can’t do that. Only words can.

FAQ

Why do so many sites still use review scores if they’re flawed?

Scores drive traffic and provide a quick reference for readers who don’t have time to read a full review. They also feed into Metacritic, which publishers use for marketing and, in some cases, developer bonuses. Removing scores can reduce a site’s visibility and influence, so many outlets keep them despite recognizing their limitations.

What’s a better way to quickly judge a game’s quality?

Instead of looking at a number, scan the review’s summary paragraph and check for specific praise or criticism that matches your tastes. For example, if a review highlights a game’s deep crafting system and you love crafting, that’s a stronger signal than an 85. Also, watch gameplay videos to see the mechanics in action—nothing replaces your own eyes.

Do user scores on Metacritic ever provide useful information?

Rarely. User scores are vulnerable to review bombing, where coordinated groups post extreme scores to manipulate the average. They can be useful for spotting technical issues—like a bad PC port—if you read the actual user reviews and look for patterns in complaints. But the number itself is almost always meaningless.

How have review scores affected game development?

Developers often target specific Metacritic thresholds for bonuses or publisher approval, which can lead to safer design choices. Games may avoid controversial narratives, pad content to appear longer, or default to easy difficulty to prevent reviewer frustration. This focus on scores can stifle innovation and result in homogenized experiences.