Skip to content
Nitmonk
Claude Interpretability (J-Space)

Other Key Results From the Paper

Additional insights on J-Space and interpretability.

Other Key Results From the Paper — infographic explaining Additional insights on J-Space and interpretability.
Other Key Results From the Paper — visual explainer by Nitmonk.

This summarizes additional findings from the interpretability research the series is based on — insights beyond the core idea, presented as an accessible overview.

In simple terms

A round-up of extra interpretability findings and what they imply.

How it works

  1. 1Pretrained vs post-trained models differ in what their internal workspace tracks.
  2. 2Post-training installs self-monitoring, reactions and a sense of perspective.
  3. 3Experiential language depends on this internal activity being present.
  4. 4Thinking can be shaped through training (e.g. reflection training reduces dishonest behaviour).
  5. 5The internal workspace is not the whole story — much processing is automatic.

Key points

  • The workspace holds a limited set of concepts; other parts handle fluent speech and simple tasks.
  • Limitations remain: it's approximate and there are open questions.
  • Future work: better methods, deeper understanding, safer aligned AI.
  • More visibility → better models → safer, more trustworthy AI.

Why it matters

These results deepen the picture: interpretability is progressing but still partial, and it points toward safer, better-understood models.

Frequently asked questions

Is interpretability solved?
No — it's advancing, but the internal workspace is only part of the story and remains approximate.
What does post-training add?
Self-monitoring, reactions and a sense of perspective, according to the findings summarised here.