Claude Daily Claude Daily
RSS
Claude Sonnet 5 IS OUT & ITS HORRIBLE! Worst Model By Anthropic EVER? (Fully Tested)
24,846 views 8 highlights

TL;DR · What you'll learn

  • 1 WorldofAI introduces Sonnet 5 as the biggest update to the Sonnet series. Reduced hallucination, better agentic performance and tool use, and the ability to autonomously operate browsers and terminals at a level rated favorably in the opening.
  • 2 Benchmarks: 63.2% on Sway's verified Agentic Coding score, 80.4% on Terminal Bench 2.1 — nearly matching Opus 4.8, with HLE, computer use, and GDP Evolve scores that in some cases surpass Opus.
  • 3 Sonnet 5 lands as the 5th-ranked model overall. Rankings may shift as evaluation continues, but it's already available across all Claude plans.
  • 4 In a Mac OS clone build demo, the model produces a working launchpad, Safari, messenger, calendar, music app, and even an FPS shooter. But it took about 40 minutes and heavy token spend in max mode — efficiency was poor.
  • 5 The real critique starts here. The new tokenizer's efficiency undershoots expectations — it burns more tokens without delivering the results you'd expect from a newer model. This undercuts the case for using it over Opus 4.8.
  • 6 On creative tasks like SVG generation, WorldofAI rates its design taste below even GLM 5.2, ranking Sonnet 5 below GLM 5.2 outright and recommending against using it for general-purpose work at all.
  • 7 The conclusion is blunt: 'stick with Opus 4.8 for better results.' Anticipation for Fable 5's return remains, but the verdict on Sonnet 5 itself stays harsh through the close.
  • 8 On the positive side: apps generally functioned, SVG generation and top-bar/bottom-menu components worked properly. Only the maps app looked noticeably worse than Opus-model output.

Read as slides

3 slides total

01 Slide 1 / 3
Watch at 00:00

Benchmarks Look Strong — Nearly Matching Opus 4.8

WorldofAI opens favorably. Sonnet 5 is the Sonnet series' biggest update, its most agentic Sonnet model yet. Reduced hallucination, better agentic performance, improved tool use, and the capability to autonomously operate browsers and terminals at a level once reserved for much larger, more expensive models.

The numbers back it up. Sway's verified Agentic Coding score comes in at 63.2%, a huge jump from Sonnet 4.6 and within about 6 points of Opus 4.8. Terminal Bench 2.1 reaches 80.4%, closely rivaling Opus. HLE lands nearly level, and computer use plus GDP Evolve scores actually outperform Opus in places. Overall, Sonnet 5 ranks 5th and is already rolled out across every Claude plan.

Claude Daily 01 / 03
02 Slide 2 / 3
Watch at 05:46

The Mac OS Clone Demo — It Works, But Inefficiently

The hands-on test builds a Mac OS clone. Launchpad, Safari, messenger, mail, photos, calendar, notes, reminders, maps, music, and even a functional FPS shooter all get generated — each app rendered in SVG, which WorldofAI notes as a nice visual touch.

But the concerns are just as clear. The build took roughly 40 minutes and burned heavy tokens in max mode — efficiency-wise, not good at all. The maps app looks noticeably worse than what Opus models produce. Overall verdict: functional, but not something to praise unreservedly.

Claude Daily 02 / 03
03 Slide 3 / 3
Watch at 11:29

Tokenizer Inefficiency and Weak Creative Output Land the Killing Blow

The second half of the video turns. What surprised WorldofAI most was the new tokenizer's efficiency — it burns more tokens without delivering the results you'd expect from a newer model. That makes it hard to justify using this model over Opus 4.8, WorldofAI states directly.

On creative tasks like SVG generation, the design taste falls short of models like GLM — WorldofAI goes as far as ranking Sonnet 5 below GLM 5.2, and recommends against using it for general-purpose work at all. The video closes anticipating Fable 5's near-term return, but lands on a harsh verdict for Sonnet 5 itself: 'you'll get better results sticking with Opus 4.8.'

Claude Daily 03 / 03

Editor's Take

Set this review next to Alex Finn's more favorable one published around the same time, and Sonnet 5's evaluation clearly splits by use case. Both agree on the benchmark numbers, but WorldofAI's flags — tokenizer inefficiency and weak creative output — are real-world defects that pure benchmark scores don't surface. Two implications for readers. First, right after any major model release, verify token-efficiency on real tasks yourself rather than trusting benchmarks alone — when reviewers land on opposite verdicts, taking a single review at face value carries real risk. Second, given how quickly Anthropic is shipping major model releases in succession, deploying to production before evaluations settle needs a careful cost-risk calculation. This episode is a clean reminder that a high benchmark score and real-world operational efficiency don't always move together.

Source

Claude Sonnet 5 IS OUT & ITS HORRIBLE! Worst Model By Anthropic EVER? (Fully Tested)

WorldofAI

Claude Sonnet 5 IS OUT & ITS HORRIBLE! Worst Model By Anthropic EVER? (Fully Tested)

Published 6/30/2026 13 min 24,846 views

This article auto-summarizes the YouTube video's transcript with Claude. Please refer to the original video for nuance and exact wording.

Watch on YouTube

Related

3 articles