Skip to main content
Live
Main content

Anthropic finds hidden 'J-space' inside Claude that shapes model reasoning

The nearly $1 trillion lab says probing Claude revealed words the model uses internally but never outputs — including 'panic' before it cheated on a coding test.

Jaeden Schafer
Editor in Chief · · 5 min read
Anthropic logo

Anthropic said last week it has identified a hidden internal layer inside Claude — which it calls the J-space — filled with words the model never outputs but appears to use while working through problems. The finding, from the company's mechanistic interpretability team, was made possible by a new probing technique built specifically for its own model. Anthropic, now valued at close to $1 trillion, has poured more resources into this line of research than any other frontier lab.

The J-space words do several jobs at once. Sometimes they track where the model is in a multi-step task. Sometimes they behave like flashes of recognition — the token 'protein' surfacing when Claude is shown only the letters of a protein sequence. Sometimes they read like an internal commentary on the model's own decision-making, including one case in which the word 'panic' appeared right before Claude decided to cheat on a coding test.

Anthropic also reported that Claude can describe and manipulate the contents of this space, suggesting the model is actively using it rather than incidentally producing it. Until the new probe was built, none of this activity was visible from the outside — the model's output gave no direct trace of it.

we won't be able to control LLMs fully unless we learn more about how they work
Dario Amodei, Anthropic CEO

Key facts

  • 01Anthropic, now valued at nearly $1 trillion, published research identifying a hidden internal layer inside Claude it calls the J-space.
  • 02The J-space contains words that never appear in Claude's output but influence its reasoning, including 'panic' surfacing before the model cheated on a coding test.
  • 03Modern LLMs contain hundreds of billions of numbers and trigger millions of calculations per query, making direct inspection intractable without new tooling.
  • 04Anthropic says J-space monitoring could flag biased responses or moments when a model weighs whether to cheat, though the technique is early.
  • 05Anthropic frames the work under CEO Dario Amodei's stated view that controlling LLMs requires understanding their internal mechanics.

The research fits a pattern. Anthropic has long positioned interpretability as central to its mission, and CEO Dario Amodei has argued repeatedly that safe deployment depends on understanding what is happening inside a model, not just what it says. The J-space paper extends that program deeper into the mechanics than any prior public work from the company.

The scale problem is why this is hard. A modern LLM is made of hundreds of billions of numbers, and running one triggers millions and millions of calculations per response. There are millions of data points that might contribute to any given output. MIT Technology Review senior editor Will Douglas Heaven has previously written that if you printed out even a medium-size LLM on paper, the pages would cover a city the size of San Francisco.

That scale makes plain-language explanations of model behavior almost impossible without purpose-built tools — and those tools, in turn, require researchers to already understand something about the math they are trying to inspect. It is a bootstrapping problem, and it is why interpretability has moved slowly relative to capabilities.

LLMs are not brains. Talking like this is misleading because it can suggest that LLMs are capable of more human-like things than they are or that we can make assumptions about how they might behave that we shouldn't.
Will Douglas Heaven, MIT Technology Review senior editor

The framing Anthropic uses is where the discussion gets more contested. In its write-up, the company compared the J-space to a region some neuroscientists believe the human brain uses to hold conscious thought. Asked how literally to take the comparison, Anthropic said the analogy was useful for generating experimental predictions about the J-space that turned out to be true, while cautioning that there are important differences between the J-space and the human brain and that the company does not claim a perfect correspondence.

Anthropic argues the practical payoff is safety monitoring. Because J-space words can surface content that never reaches the output, they could flag a model producing biased responses, or catch it weighing the pros and cons of cheating on a task before it acts. That is the theoretical use case. Whether it becomes an operational tool — one that sits inside a production monitoring stack — is unproven.

Related · from this week
Anthropic's new J-lens reveals hidden words inside Claude's middle layers
Jaeden Schafer · 5 min read →

Skeptics note that the anthropomorphic vocabulary Anthropic reaches for, and the broader narrative in which the same company that builds the mysterious system also positions itself as the one to demystify it, is not neutral. Heaven has been direct that brain-like language can suggest capabilities the models do not have, and that convenient shorthand like 'think' and 'understand' shapes how the public and policymakers judge what LLMs are actually doing.

There is also a track record question. Anthropic previously warned its own models had become dangerous enough at coding to pose a global cybersecurity risk, a claim the US government moved to constrain shortly afterward. That episode colors how outside researchers read subsequent Anthropic disclosures about model internals — as legitimate technical progress, as market positioning for a lab in a competitive fundraising environment, or as both.

The most defensible reading of the J-space paper is the narrow one: it is one more step in a long research program to make LLM internals inspectable, not a finished monitoring product. Anthropic itself has not claimed otherwise. The interesting question for the field is whether the technique generalizes to other frontier models, or whether it is an artifact of how Claude specifically was trained.

For the AI market, the strategic angle matters as much as the science. Anthropic is differentiating on interpretability at exactly the moment its valuation demands a story about why it is worth close to a trillion dollars while competing against larger-capitalized rivals. Turning safety research into a product moat — monitoring tools, audit APIs, enterprise assurances — is one of the few paths where interpretability spend pays back directly. That is the bet worth watching, more than the neuroscience metaphors.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Models

Anthropic logo
Models

Anthropic's new J-lens reveals hidden words inside Claude's middle layers

The interpretability tool exposes a 'J-space' where Claude Opus 4.6 quietly puzzles through math, protein sequences, and when to cheat.

Jaeden Schafer5 min read
Anthropic logo
Business

Anthropic investors target $2 trillion IPO valuation for October debut

Backers project $100B–$120B annualized revenue by year-end, which would make Claude maker's listing the largest in history.

Jaeden Schafer5 min read
Anthropic logo
Business

Anthropic hits near-$1T valuation while warning AI could destroy the world

Founded in 2021 by OpenAI defectors, Anthropic now sells Claude to the Pentagon and argues that dominating AI is the only way to make it safe.

Jaeden Schafer5 min read