Iâve been reading, writing, and arguing about game reviews for two decades. In that stretch, Iâve watched a sprawling 300-hour RPG get flattened into an 8.7, a brilliant indie oddity dismissed with a 6.5, and a technically broken blockbuster handed a 9 because the box had the right logo. The number sits there, bold and definitive, pretending to wrap up everything that matters. It doesnât. The single-digit review score isnât just an oversimplificationâitâs a structural failure that distorts how we talk about games, how we buy them, and how they get made.
The Illusion of Precision
When a site stamps a 7.8 on a game, it borrows the language of measurement. That decimal whispers, âwe ran the tests, we crunched the data.â But nobody crunched anything. Iâve been in editorial meetings where a score climbed from 7.5 to 8.0 because the art director liked the color palette. Iâve watched a game shed half a point because the reviewer was hungry during the final boss fight. The number projects objectivity where none exists.
Look at Death Stranding. At launch, scores ran from 3.5 to 10. Same game, same mechanics, same lonely trudges across a shattered America. One critic found it meditative and profound; another called it a glorified fetch quest. Both takes are legitimate. The game is a strange, ambitious experiment that some people will adore and others will loathe. But the score system shoves that reality into a false binary: masterpiece or failure. The truth gets buried under the weight of a number.
How Scores Strip Away Context
A review score peels off every qualifier. It wonât tell you that Cyberpunk 2077 got a 9 from a critic running it on a high-end PC before the console versions imploded. It wonât mention that the 7.5 for Days Gone came from someone who only saw the first ten hours, missing the narrative payoff that redeems its sluggish opening. The number floats free of its origins, hardening into a permanent brand on the gameâs reputation.
I remember reading a 9.5 review for Red Dead Redemption 2 and feeling genuinely confused when I finally played it. The controls felt like steering a drunk marionette through molasses. The mission design was so rigid that stepping two feet off the prescribed path triggered a fail state. The world was stunning, the writing exceptional, but the moment-to-moment interaction was frequently miserable. That 9.5 didnât prepare me for any of that. It just told me the game was excellent, full stop. A number canât communicate that a game is simultaneously a technical marvel and a mechanical frustration.

The Metacritic Problem
Metacritic didnât invent review scores, but it perfected their tyranny. By aggregating scores into a single weighted average, it manufactures the illusion of consensus. A game with an 85 Metascore is âgreat,â while a 74 is merely âmixed.â That eleven-point gap might represent the difference between a studio receiving a bonus and laying off staff. Publishers literally tie developer compensation to Metacritic thresholds. Obsidian Entertainment missed a bonus for Fallout: New Vegas by one Metacritic pointâ84 instead of the contractually required 85. One point. Thatâs the difference between a reviewer who had a good lunch and one who didnât.
Metacritic also flattens the diversity of critical voices. A thoughtful, detailed review from a small outlet gets reduced to a number and averaged in with scores from publications that may have entirely different standards. The siteâs color-coded systemâgreen for good, yellow for mixed, red for badâtrains readers to scan for colors instead of engaging with arguments. Iâve watched friends decide against buying a game solely because its Metacritic score was yellow. They never read a single review.
The 7â10 Scale Collapse
Most review outlets operate on a de facto 7â10 scale. Anything below 7 is treated as garbage. A 5 should mean average, but in practice, a 5 is a disaster. This inflation isnât accidentalâitâs structural. Publishers withhold early review copies from outlets that score too low. Advertising relationships create soft pressure to be generous. And readers themselves punish low scores with outrage, because theyâve already pre-ordered the game and need their purchase validated.
Consider Starfield. Pre-release hype was astronomical. When reviews landed with 7s from some major outlets, the backlash was immediate and vicious. A 7 is a good score. It means the game has substantial merits but also significant flaws. But in the current climate, a 7 reads as a condemnation. The discourse around that game wasnât about its ambitious scope or its uneven executionâit was about whether it âdeservedâ an 8 or a 9. The number became the entire conversation.

What Scores Actually Measure
If weâre being honest, review scores measure three things: production values, brand expectations, and the reviewerâs personal tolerance for frustration. A game with high-budget cutscenes and a licensed soundtrack starts at an 8 before anyone touches the controller. A sequel to a beloved franchise gets a built-in cushion because the reviewer already likes the world. And a game that respects the playerâs timeâno grinding, no paddingâgets a boost that has nothing to do with artistic merit.
Hades earned near-universal acclaim, and deservedly so. But part of its high scores came from the fact that it never wasted your time. You die, youâre back in the hub in seconds, new dialogue waiting. Compare that to Persona 5, which I love but which also features a tutorial that lasts roughly fifteen hours. If a reviewer docked Persona 5 for that, theyâd be correct. But the score wouldnât tell you that. It would just say 9.3, and youâd assume itâs flawless.
Genre Bias in Numerical Ratings
Review scores carry an unspoken genre tax. Strategy games, simulation games, and complex RPGs consistently score lower than action-adventure titles with comparable quality. Crusader Kings III is one of the most sophisticated strategy games ever made, a dynastic simulator of staggering depth. Its Metacritic score sits at 91. Thatâs excellent, but compare it to God of War Ragnarök at 94. Is Ragnarök really three points better, or does it simply benefit from being a cinematic action game that reviewers can finish in thirty hours and rate confidently?
The answer is obvious. Reviewers are human beings with deadlines. A game that takes sixty hours to understand will rarely get the same thorough evaluation as one that delivers its pleasures in the first five. The score doesnât account for this. It pretends all genres compete on a level field.
The Player Score Counter-Mess
User scores are supposed to be the antidoteâthe voice of the people against the corrupt critic elite. In practice, theyâre even worse. User scores on Metacritic and similar platforms are routinely bombed by coordinated campaigns. A game ships with a minor performance issue and gets 0s from people who havenât played it. A game makes a political statement someone dislikes and gets review-bombed into oblivion. The user score for The Last of Us Part II dropped to 3.4 within hours of release, before anyone could have possibly finished its 25-hour campaign. That number tells you nothing about the game. It tells you about a culture war.
Steamâs binary thumbs-up/thumbs-down system is marginally better because it avoids the false precision of a 100-point scale. But it still reduces complex opinions to a single bit. A game can earn 95% positive reviews and still have fundamental problems that a significant minority finds game-breaking. The percentage hides the distribution.

What We Lose When We Reduce Games to Numbers
The most insidious effect of review scores is how they shape the games themselves. Developers know that certain features reliably boost Metacritic scores: high-fidelity graphics, cinematic set-pieces, emotional story beats that photograph well in trailers. Systems that are deep but hard to demonstrateâcomplex AI behaviors, emergent gameplay, subtle mechanical interactionsâdonât move the needle. So they get cut. Budgets flow toward the score-boosters, and the medium narrows.
Look at the evolution of the Dragon Age series. Origins was a tactical RPG with pause-and-play combat, extensive party management, and branching dialogue trees that could lock you out of entire questlines. Inquisition streamlined everything into an action-oriented open-world checklist. The Veilguard reportedly strips away even more complexity. Each iteration chased broader appeal and higher review scores, and each lost something of what made the original distinctive. The numbers rewarded the changes. The art suffered.
Alternatives That Already Exist
Some outlets have abandoned scores entirely. Eurogamer switched to a system of Essential, Recommended, and Avoid badges in 2015. Kotaku stopped scoring reviews years ago. Rock Paper Shotgun never used scores. These publications force readers to actually engage with the textâto weigh the criticâs descriptions of what works and what doesnât against their own preferences. A game described as âpunishingly difficult but fairâ might appeal to a Soulslike fan while warning off someone looking for a relaxing experience. A number canât make that distinction.
Other outlets have experimented with multi-axis scoring. Instead of one number, you get separate ratings for gameplay, story, visuals, sound, and value. This is betterâit at least acknowledges that a game can excel in one area while failing in another. But it still reduces each category to a number, and it still invites the reader to average them into a single figure in their head.
A Modest Proposal for Readers
If youâre still reading reviews that end with a score, hereâs what I suggest: ignore the number. Read the text. Pay attention to the specific complaints and the specific praise. If a reviewer says the combat feels weightless, ask yourself whether that matters to you. If they praise the writing but criticize the pacing, consider your own tolerance for slow burns. The information you need is in the words, not the digit at the bottom of the page.
Better yet, find critics whose tastes you understand. Follow them over time. Learn what they value and what they dismiss. A reviewer who hates crafting systems will always give survival games a lower score than they deserveâfor you, if you love crafting. Once you know a criticâs biases, their reviews become useful regardless of the number attached. The score becomes irrelevant; the perspective becomes everything.
Frequently Asked Questions
Why do review sites still use scores if theyâre so flawed?
Because scores drive traffic. A number is easy to aggregate, easy to compare, and easy to argue about. Metacritic and similar sites have built entire business models around turning reviews into data points. Publications that drop scores often see an initial dip in readership, even if their long-term credibility improves. The economic incentive to keep scoring is powerful, even when editors know the system is broken.
Are user reviews on Steam more reliable than critic scores?
Steam reviews have different problems. The binary recommend/not recommend system avoids false precision, but itâs vulnerable to review bombing, meme reviews, and the fact that players with hundreds of hours often leave the most critical reviewsâbecause theyâre the ones who engaged deeply enough to see the flaws. A game with 95% positive reviews might still have a late-game balance issue that only the most dedicated players encounter. Read the actual reviews, not just the percentage.
What should I look for in a review instead of a score?
Look for specific descriptions of systems and experiences. A good review tells you what the game asks you to do, how it feels to do it, and what kind of player would enjoy it. Phrases like âthe inventory management is tediousâ or âthe dialogue trees have genuine consequencesâ give you actionable information. A number just tells you whether the reviewer liked it, which may have nothing to do with whether you will.
Do review scores affect which games get made?
Absolutely. Publishers use Metacritic scores to determine bonuses, greenlight sequels, and allocate marketing budgets. A game that scores below 80 is often considered a commercial disappointment regardless of sales. This creates a powerful incentive for developers to design games that score well rather than games that are interesting. Safe, polished, focus-tested products get higher scores than risky, innovative ones. The number shapes the art.
The next time you see a big, bold score at the end of a review, remember what it actually represents: one personâs rough estimate, rounded to look authoritative, stripped of all the context that might actually help you decide whether to spend your money and your time. Games are too complex, too varied, and too personal to fit inside a single digit. The sooner we stop pretending otherwise, the better our conversations about this medium will become.