The transparency paradox inside the black box
Anthropic’s J-lens can surface internal model patterns that never appear in the output. Governance now has to decide whether looking creates responsibility.
Anthropic’s global-workspace research describes J-space, a shared representation that can connect patterns across a model’s layers. J-lens offers a way to inspect that space. The result is not a transcript of thought, but evidence about internal features that may never be visible in a response.
The EU AI Act’s Article 13 sits on the provider of a high-risk system, and what it asks for is documentary: a system designed to be transparent enough to use, shipped with instructions that state its accuracy, its known limitations, its foreseeable risks and the human oversight it needs. J-lens works at a different level: it can surface internal patterns that appear in none of that. The instructions stay owed either way, which is the point. A new way of looking inside a model does not discharge the duty to say, on paper, what the model is for.
That creates a monitoring paradox. A team that does not look may remain unaware; a team that looks may identify a risk it can no longer ignore. Governance should not reward blindness. It should distinguish exploratory signals from confirmed findings, record proportionate assessment, and define who decides what remediation is warranted.
A safe harbour for good-faith detection would help. Subject to careful limits, organisations that probe models, document their methods and promptly address credible findings should not incur extra exposure merely because they sought to understand the system. Interpretability can then become a control, rather than a reason to keep the lens closed.