Claude Sonnet 5 in the Wings, a Post-Mythos Model, Sakana's Fugu, and GPT-5.6 Pro This Week
TL;DR · What you'll learn
- 1 Anthropic appears to be prepping the fifth Claude Sonnet — its model slug surfaced through a partner provider. Early tests suggest a solid upgrade with strong outputs.
- 2 Also reported: Anthropic has trained a model more capable than Mythos. Pushing past an already-banned controversial model — handling of the release will be closely watched.
- 3 OpenAI is set to launch GPT-5.6 Pro this week (Tuesday or Thursday). The first major voice-model refresh since GPT-4o, called BDI, ships alongside it.
- 4 Japan's new Sakana AI lab unveiled Fugu Ultra Systems, claiming Claude Fable 5 / Mythos 5-level performance. A closer look says the claims are optimistic — but the model is genuinely capable.
- 5 The unannounced Anthropic model has stronger reasoning, sharper incentive coding, improved planning, and more reliable execution on massive tasks. Open question: ship via Project Last Sync, fold into a future public model, or keep internal to accelerate the next generation.
- 6 Hands-on with GPT-5.6 Pro: with the right prompts, it can build basically anything. Front-end quality and design taste have stepped up substantially. The new BDI 1 voice model has an August 2025 knowledge cutoff.
- 7 Head-to-head: Sakana Fugu Ultra vs Claude Opus 4.8 Ultra Code on a 3D Crossy Road-style game generation test. Opus wins on final quality; Fugu wins on speed, cost, efficiency, and difficulty scaling.
- 8 By the numbers: Opus took 79 minutes, ~940K tokens, $37.85, and got stuck twice. Fugu Ultra finished in 22 minutes on ~90K tokens for $7.32 — roughly 1/5 the cost. Quality vs efficiency is now an active choice.
Read as slides
3 slides total
Signs of Anthropic's next generation — Sonnet 5 and a post-Mythos model
It's a busy week for model launches and leaks. Anthropic is prepping the fifth Claude Sonnet, with the model slug surfacing through a partner provider. WorldofAI's early testing reports a solid upgrade with impressive outputs; some examples come later in the video.
More striking is the report of a model trained that's more capable than Mythos — pushing past the already-banned, already-controversial model. The release path matters: ship through Project Last Sync, fold into a future public model, or hold internally to accelerate the next generation. Anthropic's choice will say a lot about how it now reads the regulatory environment.
OpenAI's week and Sakana's arrival — GPT-5.6 Pro, BDI, and Fugu
The competitor side is loud too. OpenAI is set to launch GPT-5.6 Pro this week — Tuesday or Thursday looks likely. Alongside it ships BDI (GPT BDI 1), the first major voice-model refresh since GPT-4o, with an August 2025 knowledge cutoff. Some users already see it rolling out in the ChatGPT app.
Japan also entered the frontier conversation. Sakana, a new AI lab, unveiled Fugu Ultra Systems and claimed Claude Fable 5 / Mythos 5-level performance on multiple benchmarks. WorldofAI's own look at the results says the claims are optimistic and the gap is real — but the model is still genuinely capable. A non-US challenger to the frontier oligopoly is worth tracking on its own terms.
The head-to-head — Sakana Fugu Ultra vs Claude Opus 4.8 Ultra Code
The video runs an actual head-to-head: Sakana Fugu Ultra vs Claude Opus 4.8 Ultra Code, on a 3D Crossy Road-style game generation test. The split was striking.
Opus 4.8 Ultra Code won on final design quality, functionality, and polish. But it took 79 minutes, burned around 940K tokens, cost $37.85, and got stuck twice along the way. Fugu Ultra had real rough edges — inverted controls, a wonky camera, missing SFX — but finished in 22 minutes on ~90K tokens for $7.32 (roughly 1/5 the cost), and its difficulty scaling was actually better. The summary: Opus wins quality; Fugu wins speed, cost, and efficiency. That's now a real trade-off, not a theoretical one.
Editor's Take
The real signal this week is category fragmentation. Anthropic chases absolute capability with Sonnet 5 and the post-Mythos model; OpenAI grabs the high-frequency end-user surface with GPT-5.6 Pro and the new BDI voice; Japan's Sakana enters on a completely different axis — quality traded for speed, cost, and efficiency. Fugu Ultra losing to Opus 4.8 on quality but finishing at one fifth the cost isn't a footnote. It means usage segmentation is now real: Opus for the critical path where final quality matters, Fugu for high-volume iteration and exploration. Combined with the Mythos ban story, the case against single-vendor lock-in gets stronger every week. Multi-model operations — picking the right model per task on availability, cost, and quality dimensions independently — is starting to look like the default operating mode for next quarter.
Source
WorldofAI
Claude Sonnet 5, Mythos 6 ALREADY?, GPT-5.6 This Thursday, Sakana Fugu Beats Mythos, & More! AI NEWS
This article auto-summarizes the YouTube video's transcript with Claude. Please refer to the original video for nuance and exact wording.
Watch on YouTube →Related
3 articles
PIVOT 公式チャンネル
Building a Personal AI Agent for a TV Anchor in 4 Hours: How Claude Code Turned 21 Hours of Research Into 10 Minutes
Japanese business show PIVOT filmed a live segment building a personal AI agent for anchor Tsuyoshi Nojima using Claude Code, from scratch, in four hours.
AI Revolution
China Warns of a 'Backdoor' in Claude Code, Demanding Restrictions Over Alleged Location Data Leaks
China is starting to lock down its domestic AI industry -- Alibaba and ByteDance reportedly plan to shut down user-created agents on July 15th.
AI News & Strategy Daily | Nate B Jones
Claude Fable 5 Bossed 20 Cheap AI Agents to Build a Whole Website for $8 -- and Caught Its Own Cheating Along the Way
A swarm of roughly 20 cheap AI agents, orchestrated by Fable, rebuilt the creator's wife's website in one hour -- better than six days of hands-on work with Codex the month before.