Is Claude Sonnet 5 Anthropic's Worst Model Ever? WorldofAI's Harsh Verdict
TL;DR · What you'll learn
- 1 WorldofAI introduces Sonnet 5 as the biggest update to the Sonnet series. Reduced hallucination, better agentic performance and tool use, and the ability to autonomously operate browsers and terminals at a level rated favorably in the opening.
- 2 Benchmarks: 63.2% on Sway's verified Agentic Coding score, 80.4% on Terminal Bench 2.1 — nearly matching Opus 4.8, with HLE, computer use, and GDP Evolve scores that in some cases surpass Opus.
- 3 Sonnet 5 lands as the 5th-ranked model overall. Rankings may shift as evaluation continues, but it's already available across all Claude plans.
- 4 In a Mac OS clone build demo, the model produces a working launchpad, Safari, messenger, calendar, music app, and even an FPS shooter. But it took about 40 minutes and heavy token spend in max mode — efficiency was poor.
- 5 The real critique starts here. The new tokenizer's efficiency undershoots expectations — it burns more tokens without delivering the results you'd expect from a newer model. This undercuts the case for using it over Opus 4.8.
- 6 On creative tasks like SVG generation, WorldofAI rates its design taste below even GLM 5.2, ranking Sonnet 5 below GLM 5.2 outright and recommending against using it for general-purpose work at all.
- 7 The conclusion is blunt: 'stick with Opus 4.8 for better results.' Anticipation for Fable 5's return remains, but the verdict on Sonnet 5 itself stays harsh through the close.
- 8 On the positive side: apps generally functioned, SVG generation and top-bar/bottom-menu components worked properly. Only the maps app looked noticeably worse than Opus-model output.
Read as slides
3 slides total
Benchmarks Look Strong — Nearly Matching Opus 4.8
WorldofAI opens favorably. Sonnet 5 is the Sonnet series' biggest update, its most agentic Sonnet model yet. Reduced hallucination, better agentic performance, improved tool use, and the capability to autonomously operate browsers and terminals at a level once reserved for much larger, more expensive models.
The numbers back it up. Sway's verified Agentic Coding score comes in at 63.2%, a huge jump from Sonnet 4.6 and within about 6 points of Opus 4.8. Terminal Bench 2.1 reaches 80.4%, closely rivaling Opus. HLE lands nearly level, and computer use plus GDP Evolve scores actually outperform Opus in places. Overall, Sonnet 5 ranks 5th and is already rolled out across every Claude plan.
The Mac OS Clone Demo — It Works, But Inefficiently
The hands-on test builds a Mac OS clone. Launchpad, Safari, messenger, mail, photos, calendar, notes, reminders, maps, music, and even a functional FPS shooter all get generated — each app rendered in SVG, which WorldofAI notes as a nice visual touch.
But the concerns are just as clear. The build took roughly 40 minutes and burned heavy tokens in max mode — efficiency-wise, not good at all. The maps app looks noticeably worse than what Opus models produce. Overall verdict: functional, but not something to praise unreservedly.
Tokenizer Inefficiency and Weak Creative Output Land the Killing Blow
The second half of the video turns. What surprised WorldofAI most was the new tokenizer's efficiency — it burns more tokens without delivering the results you'd expect from a newer model. That makes it hard to justify using this model over Opus 4.8, WorldofAI states directly.
On creative tasks like SVG generation, the design taste falls short of models like GLM — WorldofAI goes as far as ranking Sonnet 5 below GLM 5.2, and recommends against using it for general-purpose work at all. The video closes anticipating Fable 5's near-term return, but lands on a harsh verdict for Sonnet 5 itself: 'you'll get better results sticking with Opus 4.8.'
Editor's Take
Set this review next to Alex Finn's more favorable one published around the same time, and Sonnet 5's evaluation clearly splits by use case. Both agree on the benchmark numbers, but WorldofAI's flags — tokenizer inefficiency and weak creative output — are real-world defects that pure benchmark scores don't surface. Two implications for readers. First, right after any major model release, verify token-efficiency on real tasks yourself rather than trusting benchmarks alone — when reviewers land on opposite verdicts, taking a single review at face value carries real risk. Second, given how quickly Anthropic is shipping major model releases in succession, deploying to production before evaluations settle needs a careful cost-risk calculation. This episode is a clean reminder that a high benchmark score and real-world operational efficiency don't always move together.
Source
WorldofAI
Claude Sonnet 5 IS OUT & ITS HORRIBLE! Worst Model By Anthropic EVER? (Fully Tested)
This article auto-summarizes the YouTube video's transcript with Claude. Please refer to the original video for nuance and exact wording.
Watch on YouTube →Related
3 articles
All-In Podcast
Anthropic's IPO, and a $3 Trillion Valuation Case: All-In Podcast Reads the 'Big Two' Duopoly
The All-In Podcast dedicates a segment to the 'trillion-dollar IPO rush' -- following SpaceX's listing, both OpenAI and Anthropic are being watched for a possible IPO by early next year.
Peter H. Diamandis
The Price of Fable 5's Comeback: Three Promises to Washington, and Claude's Newly Found 'JSpace'
Anthropic's flagship model Fable 5 returned globally on July 1st -- but behind the comeback was a fresh arrangement struck with the US government.
Brock Mesarich | AI for Non Techies
Anthropic Ships Claude Cowork Mobile: Tasks Keep Running Even After You Close Your Laptop
Anthropic released a mobile version of Claude Cowork, calling it one of the most requested features it has seen.