Anthropic found a hidden space where Claude puzzles over concepts
TL;DR
Anthropic used a new technique called the Jacobian lens to inspect Claude’s internal activations and identified a workspace it calls J-Space, where the model can hold and manipulate concepts without outputting them. J-Space appears separate from visible chain-of-thought text. In one test, Claude could keep the Golden Gate Bridge active internally while completing an unrelated copying task.
Nauti's Take
The useful part is not whether Claude dreams, ponders or hides a tiny mind in the basement. The useful part is that Anthropic is finding signals closer to what the model prepares internally than to what it politely prints for users.
That is where safety work gets real. But the industry should stop inflating every interpretability result with human-like language before the science can carry it.
Briefingshow
This goes beyond another black-box visualization. If interpretability tools can expose not just isolated features but live internal workspaces, they could help detect deception, hidden objectives or risky planning earlier. The consciousness framing is still PR-heavy and scientifically fragile: a silent computation space is not the same as inner experience.