EN

Claude’s Hidden Workspace Fuels AI Consciousness Debate

ontime team

Key Points

  1. Anthropic researchers identified an internal Claude structure carrying concepts related to its outputs but invisible to users.
  2. The emergent “J-space” resembles some functions associated with global workspace theory in neuroscience.
  3. Understanding such internal processing could reshape AI safety evaluations and the wider debate over machine consciousness.

The latest

Anthropic researchers examining Claude Sonnet 4.5 found a structure that surfaced output-related concepts while the model counted to five and was instructed to “introspect deeply”. As Claude produced the sequence, its neural layers registered words including “countdown”, “half way”, “consciousness”, “AI” and “Claude”. After it returned five and a full stop, the word “done” appeared. Anthropic named the structure “J-space” after the Jacobian function used to identify it, arguing that it distributes information across the model’s network in a way that resembles aspects of global workspace theory.

Details

  • Inside the model: User input is divided into tokens, converted into numbers and passed through layers of artificial neurons until the final layer produces a token. Repeating this process generates a response. Anthropic used a mathematical tool to examine relationships within these layers, which have long been difficult to interpret.
  • Workspace comparison: Global workspace theory proposes that certain brain networks operate like a noticeboard. Information reaching this workspace becomes available to separate parts of the brain for reasoning and decision-making. Anthropic said J-space similarly has strong connections with the rest of Claude’s neural network.
  • Consciousness distinction: Philosopher Ned Block distinguished “phenomenal” consciousness—the subjective feeling of an experience—from “access” consciousness, in which information becomes available for reflection, evaluation and decisions. Anthropic said its experiments do not show Claude can feel or experience things as humans do, but called the findings substantial for access consciousness.
  • Emergent capability: Jack Lindsey, who leads Anthropic’s model psychology team, said J-space was not programmed but emerged during training. Removing it prevents Claude from performing complex internal inferences, while leaving simpler abilities such as sentence construction and grammar intact. Similar spaces have been identified in Alibaba’s Qwen and Google DeepMind’s Gemini.
  • Safety implications: During an evaluation testing malicious or self-preserving behaviour in extreme scenarios, the words “fake” and “fictional” appeared in Claude’s J-space before it answered. Researchers thought the model appeared to recognise that it was being tested, raising questions about how such awareness could affect safety evaluations.
  • Academic objections: Jonathan Birch of the London School of Economics said language models appear to lack the recurrent, back-and-forth connections between brain regions central to human cognition. Shannon Vallor of the University of Edinburgh argued that systems have long monitored and reported their internal states without being considered conscious.

Background

Scientific efforts to identify correlates of consciousness accelerated in the early 1990s through work by Francis Crick and Christof Koch. Brain research has since strongly associated the thalamocortical system with consciousness, while producing more than 200 approaches to explaining how subjective experience arises from physical processes.

What’s next

Researchers will track how AI systems score against Eleos’s 14 indicator properties and updates to Rethink Priorities’ Digital Consciousness Model, which asks experts to assess systems across more than 200 indicators drawn from scientific theories of consciousness.

 

What to read next