Fireship Takes On the 'Is Claude Conscious' Debate: Swap a Spider's Thought for an Ant's, and the Answer Changes
TL;DR · What you'll learn
- 1 Anthropic's paper, 'A Global Workspace and Language Models,' is described as the most philosophically cursed research ever published by a company that also happens to sell API tokens.
- 2 The paper conveniently pushes the AGI-and-singularity narrative forward, though not everyone is buying it, the video cautions.
- 3 Researchers pinpointed where Claude stores 'private thoughts' and swapped one for another -- the model's entire chain of reasoning obediently followed the substituted thought.
- 4 They also deleted the region entirely, and, surprisingly, Claude kept speaking fluently and confidently in English.
- 5 The most important surprise: nobody designed this structure -- it emerged on its own through training, strikingly similar to the 1988 'global workspace theory.'
- 6 Asked how many legs a web-spinning animal has, 'spider' lit up in J-space right before the answer 'eight' -- swap it for 'ant,' and the answer changes to six.
- 7 The video underscores that Anthropic itself states in the paper this doesn't tell us whether Claude is conscious, while still finding it fascinating that a 'scratchpad for thoughts' like this can emerge spontaneously.
Read as slides
3 slides total
'Philosophically Cursed,' Per Fireship
Anthropic's paper, 'A Global Workspace and Language Models,' is described by Fireship as the most philosophically cursed research ever published by a company that also happens to sell API tokens. The claim -- that buried deep inside Claude's incomprehensible pile of matrices is a small, organized set of neural patterns called J-space -- conveniently pushes the AGI-and-singularity narrative forward, though not everyone is buying it, the video notes, before diving into Claude's weird brain.
Swapping a Thought, and Speaking Fluently After Deleting the Region
Anthropic researchers pinpointed exactly where 'private thoughts' get stored inside Claude's brain. Swap one thought for another, and the model's entire chain of reasoning obediently follows the lie. For good measure, they deleted the region entirely -- and, surprisingly, Claude kept speaking fluently and confidently in English.
The most important surprise: nobody designed this structure. It emerged on its own through training -- strikingly similar to the 'global workspace theory' proposed by Bernard Baars in 1988, the idea that the brain runs countless functions automatically in the background while consciousness is a single, brightly lit stage.
Swap 'Spider' for 'Ant,' and Eight Legs Become Six
In one experiment, Claude was asked how many legs a web-spinning animal has. Just before answering 'eight,' the concept 'spider' lit up in J-space. Researchers then surgically swapped that internal concept for 'ant' -- with neither the prompt nor the output edited -- and Claude's answer changed to 'six.'
Fireship stresses that Anthropic itself states in the paper that none of this tells us whether Claude is conscious, while still finding it fascinating that a 'scratchpad for thoughts' like this can emerge spontaneously given enough data and the right linear algebra. Personally unconvinced Claude is conscious, the video closes on a wry note.
Editor's Take
Even covering the same J-space research, the official announcement and Fireship's take strike noticeably different tones. Where the official framing emphasizes safety benefits from visibility into thought, Fireship layers in a wink at the conflict of interest -- a company selling API tokens conveniently bolstering the AGI narrative -- a good example of how the same research can land very differently depending on who's telling the story. The spider-to-ant swap changing a numeric answer conveys the technical persuasiveness of the research clearly, but given how easily secondary coverage of this topic tends to overreport 'consciousness discovered' while ignoring Anthropic's own explicit disclaimer, readers should treat this as a reminder to read primary sources carefully.
Source
Fireship
Claude is definitely not conscious…
This article auto-summarizes the YouTube video's transcript with Claude. Please refer to the original video for nuance and exact wording.
Watch on YouTube →Related
3 articles
Anthropic
What's at the Center of Claude's Mind? Anthropic Maps the Line Between Conscious and Unconscious Processing
To test whether AI models have anything like the human divide between conscious and unconscious processing, Anthropic borrowed methods from neuroscience.
Nate Herk | AI Automation
How Claude Is Creating a New Generation of Millionaires -- Without Writing a Line of Code
Vulcan, a company building software for government agencies, looks like the work of a hundred engineers -- but most of its team can't write a single line of code and built it all by describing what they wanted to Claude.
Moon
'Anthropic Is Completely Done For': A Critique Imagining Claude's Takeover of Every Profession
A critique-and-near-future-scenario video that opens from a provocative premise: invoking Mythos, Fable, and Opus 4.8, it claims Anthropic has overtaken OpenAI on nearly every front in revenue and valuation to become 'the most powerful AI company in the world.'