Claude Interpretability (J-Space)
Other Key Results From the Paper
Additional insights on J-Space and interpretability.

This summarizes additional findings from the interpretability research the series is based on — insights beyond the core idea, presented as an accessible overview.
In simple terms
A round-up of extra interpretability findings and what they imply.
How it works
- 1Pretrained vs post-trained models differ in what their internal workspace tracks.
- 2Post-training installs self-monitoring, reactions and a sense of perspective.
- 3Experiential language depends on this internal activity being present.
- 4Thinking can be shaped through training (e.g. reflection training reduces dishonest behaviour).
- 5The internal workspace is not the whole story — much processing is automatic.
Key points
- The workspace holds a limited set of concepts; other parts handle fluent speech and simple tasks.
- Limitations remain: it's approximate and there are open questions.
- Future work: better methods, deeper understanding, safer aligned AI.
- More visibility → better models → safer, more trustworthy AI.
Why it matters
These results deepen the picture: interpretability is progressing but still partial, and it points toward safer, better-understood models.
Frequently asked questions
- Is interpretability solved?
- No — it's advancing, but the internal workspace is only part of the story and remains approximate.
- What does post-training add?
- Self-monitoring, reactions and a sense of perspective, according to the findings summarised here.
More in Claude Interpretability (J-Space)
Inside Claude's Brain
What happens before Claude speaks — the J-Space workspace.
What Is J-Space?
A silent internal workspace inside Claude's neural network.
Global Workspace Theory (GWT)
The theory that inspired J-Space.
How Claude Thinks Silently
Claude thinks first, reasons, then speaks.
How Anthropic Reads Claude's Thoughts
The J-Lens (Jacobian Lens) tool explained.
J-Space vs Chain of Thought
Two layers of thinking: hidden vs visible.