I caught myself staring at a number again last night. A 7. Perfectly respectable, stamped at the bottom of a review for a game I’d spent sixty hours pulling apart. That 7 is supposed to stand for labyrinthine level design, a combat system that rewards patience over twitch reflexes, and a narrative that deliberately refuses to hand you closure. But it doesn’t. It flattens all of that into a digit that sits right next to a 7 for a competent but forgettable racing sim from the same outlet. The two games share nothing except the number, and that’s exactly the problem.
Review scores have become the default shorthand for quality—a numerical anchor readers skim for and publishers obsess over. But when a game is built on contradictions—intentionally frustrating mechanics that serve a thematic point, or a story that sacrifices pacing for atmosphere—a single digit erases the very texture that makes it worth talking about. I’m not arguing against criticism or evaluation. I’m arguing that the method we’ve settled on is intellectually bankrupt for the medium we claim to love.
The Metacritic Effect: When 74 Becomes a Failure
Metacritic didn’t invent review scores, but it perfected their tyranny. By aggregating dozens of outlets into a single weighted average, it created a pseudo-objective benchmark that publishers now treat as scripture. Bonuses get tied to Metacritic thresholds. Studios get shuttered when a game lands at 74 instead of 75. The difference between those two numbers isn’t a measurable decline in quality—it’s the opinion of one reviewer who used a slightly different internal rubric.
Consider Alpha Protocol, Obsidian’s 2010 espionage RPG. It sits at a 72 on Metacritic. That number tells you nothing about its revolutionary dialogue system, where conversations happen in real time and choices lock you into responses before you’ve fully processed the situation. It doesn’t capture how the game’s buggy shooting mechanics are almost irrelevant if you build a character who talks their way through every mission. The 72 implies mediocrity. In reality, it’s one of the most ambitious reactive narrative systems ever shipped, buried under a score that compares unfavorably to dozens of polished but forgettable shooters.
The flattening effect gets worse when scores are stripped of context. A 7 from a critic who values mechanical depth means something entirely different from a 7 awarded by someone who prioritizes accessibility. But on Metacritic, they’re identical. The number becomes a blunt instrument that erases the conversation it supposedly represents.
The Genre Bias Embedded in Every Scale
Review scales are not neutral. They carry baked-in assumptions about what makes a good game, and those assumptions overwhelmingly favor specific genres. Action games benefit from criteria like responsiveness, visual spectacle, and moment-to-moment excitement. A slow-burn strategy game or an experimental walking simulator gets measured against the same yardstick, and the results are predictably absurd.
Take Pathologic 2, Ice-Pick Lodge’s survival horror remake. It’s a game about futility, about being a doctor in a town ravaged by plague where you cannot save everyone. Resources are scarce. Time is your real enemy. The game is deliberately punishing, and that punishment is the point—it’s a thematic device, not a design flaw. Yet many reviews docked it for being “frustrating” or “unfair,” applying standards from power-fantasy RPGs to a work that explicitly rejects power fantasies. The score didn’t reflect a failure of design. It reflected a failure of the scoring system to accommodate the game’s actual goals.
This genre bias extends to indie and experimental titles. Kentucky Route Zero is a magical-realist point-and-click adventure that abandons traditional puzzle design in favor of atmospheric storytelling and thematic resonance. It’s widely celebrated now, but at launch, some outlets struggled to score it because it didn’t fit their rubric. How do you assign a number to a game that intentionally subverts the concept of player agency? The answer, too often, is to penalize it for not being something it never tried to be.

The Technical Dimension That Scores Ignore
Games are not static texts. They’re software that runs differently across hardware configurations, receives patches, and evolves over time. A review score assigned at launch captures a snapshot that may be unrecognizable six months later. Cyberpunk 2077 is the obvious cautionary tale—some outlets scored it based on pre-release PC code, others on broken last-gen console versions. The scores ranged from 9 to 4, not because critics disagreed about its artistic merits, but because they were reviewing fundamentally different products.
But the problem runs deeper than high-profile launch disasters. No Man’s Sky launched in 2016 to scores averaging around 6. Today, after dozens of free expansions that added base building, multiplayer, underwater exploration, and a full narrative campaign, it’s a dramatically different game. Those original scores remain frozen in time, permanently affixed to the product like a scar. A new player researching whether to buy it sees a 6 and reasonably assumes mediocrity, unaware that the number describes a version of the game that no longer exists.
Even without catastrophic launches, technical performance varies wildly across systems. A game that runs at a locked 60fps on a high-end PC might stutter on a base PS4. The reviewer’s experience is not universal, but the score is presented as if it were. This is particularly damaging for PC strategy games and simulators, where performance is heavily dependent on hardware and where the community often resolves issues through mods and configuration tweaks that a reviewer on deadline never touches.
When Ambition Gets Punished
There’s a perverse incentive structure at work. A game that takes no risks and executes a familiar formula competently will reliably score between 7 and 8. A game that reaches for something unprecedented but stumbles in places might score a 6. The scoring system, as currently implemented, rewards safety and punishes ambition. This isn’t a theoretical concern—it shapes what gets funded and what gets greenlit.
Look at Deadly Premonition, Swery’s surreal open-world murder mystery. It’s technically atrocious. The controls are clunky, the graphics look a generation behind, and the PC port was borderline non-functional at launch. It also features one of the most memorable casts of characters in gaming, a genuinely unpredictable plot, and a tone that oscillates between horror and absurdist comedy in ways no other game has managed. Scores ranged from 2 to 10. The 2s focused on the technical failures. The 10s focused on the artistic achievement. Neither number alone captures the experience, but the aggregate score—around 68 on Metacritic—leans negative, effectively warning players away from something singular.
The same dynamic played out with Vampire: The Masquerade – Bloodlines in 2004. Troika’s RPG shipped in a broken state, and review scores reflected that. It sold poorly, Troika closed, and the game was only salvaged years later by a dedicated fan patch. Today, it’s rightly regarded as a masterpiece of atmosphere and writing. The scores that helped kill it were accurate about the bugs but blind to everything else.

The Subjectivity That Numbers Pretend to Erase
Numbers create an illusion of objectivity. An 8.5 feels precise, scientific, the result of careful measurement. But the process behind that number is anything but. Reviewers weigh different elements according to personal preference, then perform a kind of mental arithmetic to arrive at a final digit. One critic might assign 40% weight to story, 30% to gameplay, 20% to graphics, and 10% to sound. Another might invert those priorities. The resulting numbers look comparable but are built on incompatible foundations.
This pseudo-precision becomes absurd when you examine the sub-scores some outlets provide. A game receives a 9 for graphics, an 8 for sound, a 7 for gameplay, and a 9 for story, then gets an overall score of 8.5. What formula produced that average? Was it weighted? If so, how? The reader is never told, but the decimal point implies rigor. It’s a magic trick—a way of dressing up subjective judgment in the clothes of empirical measurement.
The problem intensifies with games that defy easy categorization. Disco Elysium has no combat. Its entire gameplay consists of walking, talking, and failing skill checks—and the failures are often more interesting than the successes. How do you score “gameplay” for a game that deliberately rejects the conventions of gameplay? Some outlets gave it a separate score for “narrative” and averaged that with a lower “gameplay” score, producing a final number that penalized the game for its own design philosophy. Others simply gave it a 10 and moved on. The inconsistency isn’t between games—it’s between the frameworks applied to them.
The Player’s Burden: Decoding What a 7 Actually Means
As a reader, I’ve developed an internal translation layer for review scores. An 8 from Edge magazine means something very different from an 8 from IGN. Edge uses the full 1-10 scale, where 5 is average and 7 is genuinely good. IGN’s scale, like many mainstream outlets, effectively runs from 6 to 10, where anything below 7 signals significant problems. But this translation requires institutional knowledge that most consumers don’t possess. A casual browser sees two 8s and assumes equivalence.
This decoding burden falls heaviest on players looking for niche experiences. If you love dense military simulators, a 6 from a critic who specializes in the genre carries more weight than a 9 from someone who clearly bounced off the learning curve. But aggregators don’t surface that context. They present a single number as if it represents a consensus, when it often represents a collision of incompatible perspectives.
The solution isn’t to abolish criticism—it’s to abandon the pretense that complex evaluations can be reduced to digits. Some outlets have already moved in this direction. Eurogamer dropped scores in 2015, replacing them with tags like “Essential,” “Recommended,” and “Avoid.” This system forces readers to engage with the text rather than skipping to the number. It’s not perfect—tags can still be reductive—but it’s a meaningful step away from the false precision of a 7.8.
The Publisher’s Weapon: How Scores Get Deployed Against Developers
Review scores don’t just mislead consumers—they get weaponized internally. Developers have told me, off the record, that publishers use Metacritic scores to justify withholding bonuses, canceling sequels, or reassigning teams. A game that scores 84 instead of the contractually specified 85 can cost a studio millions. That one-point gap might reflect nothing more than a few reviewers who didn’t connect with the art style, but the financial consequences are real and devastating.
This creates a chilling effect on design. If your bonus depends on hitting an 85, you’re going to sand off every rough edge that might alienate a reviewer. You’re going to prioritize broad appeal over distinctive vision. You’re going to make the game that scores well rather than the game that matters. The scoring system, in this context, becomes a mechanism for homogenization—a force that pushes the medium toward the safe, the familiar, and the forgettable.
Obsidian again provides the case study. After Alpha Protocol’s 72, the studio narrowly survived and eventually produced Fallout: New Vegas, which scored an 84 on Metacritic—one point short of the 85 threshold in their Bethesda contract. That single point allegedly cost them a substantial bonus. The game is now considered one of the finest RPGs of its generation, but the score-based contract turned a triumph into a financial disappointment. The number mattered more than the legacy.

What Should Replace the Number?
I’m not naive enough to suggest that scores will disappear. They’re too convenient for marketing departments, too embedded in consumer habits, too profitable for aggregators. But I can describe what a better system would look like, and I can point to the outlets that are already building it.
A useful review should answer three questions. First, what is this game trying to do? This requires the critic to engage with the work on its own terms, not against an abstract checklist. Second, how well does it achieve that goal? This is where technical analysis and design critique belong. Third, who is this game for? A horror game that’s too intense for casual players might be perfect for genre veterans, and the review should make that distinction explicit.
Some outlets structure their reviews around these questions without ever assigning a number. Rock Paper Shotgun’s reviews often conclude with a paragraph that synthesizes the critic’s experience and recommends the game to specific audiences. Kotaku’s reviews sometimes include a “should you play this” section that addresses different player profiles. These approaches treat readers as individuals with distinct tastes rather than as a monolith that needs a single verdict.
For outlets that insist on keeping scores, the minimum improvement would be transparency. Show the rubric. Explain the weighting. Acknowledge the platform and hardware used for testing. Update the score when major patches release. These steps wouldn’t solve the fundamental problem, but they would at least make the number’s limitations visible rather than hiding them behind a veneer of authority.
The Games That Defy Scoring
Some works are so resistant to numerical evaluation that they expose the entire system’s bankruptcy. The Stanley Parable is a game about games, a meta-commentary on choice and agency that deliberately frustrates completionist impulses. Assigning it a score feels like missing the joke. Outer Wilds is a knowledge-based exploration game where progression depends entirely on what the player understands, not what they’ve unlocked. A score can’t convey that the entire experience changes once you know certain things—and that the game is designed to be played exactly once, with full attention, like a novel you can’t reread.
Then there’s Dwarf Fortress, a game so procedurally deep and aesthetically impenetrable that it’s essentially its own genre. A review score for Dwarf Fortress is meaningless unless it’s accompanied by a detailed explanation of what the game simulates and why that simulation matters. The number tells you nothing. The stories the game generates—of fortress collapses, were-lizard infestations, and cats dying of alcohol poisoning after walking through spilled beer and cleaning themselves—are the actual review.
These games aren’t outliers. They’re the leading edge of a medium that’s still discovering what it can do. As games continue to diversify—into autobiographical experiences, generative storytelling, and forms we haven’t yet named—the single-digit score will look increasingly absurd. It’s a relic of a time when games were simpler products, evaluated like toasters on a consumer reports scale. That time is over.
Frequently Asked Questions
Why do review scores still exist if they’re so flawed?
Review scores persist because they serve multiple entrenched interests. Aggregators like Metacritic build their business model around them. Marketing departments use high scores in promotional materials. Consumers, overwhelmed by choice, use scores as a quick filtering mechanism. The system is self-reinforcing, even though everyone involved—critics, developers, and readers—acknowledges its limitations in private conversations.
Don’t some games deserve a simple numerical rating?
Some games are straightforward enough that a number feels adequate. A basic match-three puzzle game or a competent but unambitious sports title might not suffer much from being reduced to a digit. The problem is that the same system gets applied to everything, from those simple games to sprawling narrative experiments. A tool that works for one type of product fails catastrophically for another, and the industry uses the same tool across the board.
What can I do as a reader to get better information about games?
Read the full review text, not just the score. Follow critics whose tastes you understand, even if they don’t always align with yours—knowing a critic’s preferences helps you interpret their judgments. Seek out outlets that have dropped scores entirely or that provide detailed context about their evaluation process. And when a game sounds interesting despite a middling score, investigate further. Some of the most rewarding experiences in gaming sit at 72 on Metacritic, waiting for players who look past the number.
Have any major outlets successfully moved away from scores?
Yes. Eurogamer abandoned numerical scores in 2015 in favor of a recommendation system. Kotaku has experimented with score-free reviews. Rock Paper Shotgun has never used scores. These outlets demonstrate that it’s possible to build an audience without reducing criticism to digits. Their success suggests that the industry’s dependence on scores is more about habit and institutional inertia than genuine necessity.