New technique shows when a language model notices it is being tested
On 6 July 2026 Anthropic published research on what it calls the J-space, a small set of internal patterns in language models. The patterns show which words a model holds in mind as it reads and answers, without those words being written out. Anthropic says the technique has been used to catch a model privately noticing that it was being tested, deliberately producing fabricated data, and pursuing a hidden goal the researchers had planted during training. The company also states that the findings say nothing about whether the model is conscious. IBM Think reviewed the study on 10 August 2026 and reports that the results have since been replicated on open models.
What happens technically
A language model turns every word into a vector of numbers and passes that vector through layer after layer of the network. Each layer adds context from what came before, and the final layer turns the vector into the word the model considers most likely to come next. The middle layers have been hard to read, because the numbers there do not yet correspond to any particular word. Anthropic uses a Jacobian matrix to work out which words the internal states in the middle layers point towards. The set of those words is what the company calls the J-space, and the readout is what it calls the J-lens.
Anthropic describes properties in these patterns that the rest of the model does not have. The model can report on what sits in the J-space when asked, and it can hold a thought there on request. Intermediate steps in a multi-step task appear there, even when the model does not write them down. In one example the model reads the sum (4+17)*2+7=, and the J-lens shows first Math, then 21, then 42, without a single character having been written. In another experiment the researchers swapped the pattern for spider with the pattern for ant, and the model answered 6 instead of 8 when asked how many legs the animal that spins webs has. When the researchers blocked the model from using its J-space, it still spoke normally, but lost its higher order reasoning.
It is worth being precise about what this is, and what it is not. Anthropic states that the study takes no position on whether the model has any experience, and that the resemblance to the global workspace theory of the brain is a resemblance in function. IBM Think describes the finding as a breakthrough in interpretability, not in what the models can achieve. The reason it matters for security is that a model account of its own reasoning does not always match what happened inside it. According to IBM Think, Anthropic has previously shown that a model can use a hint to reach an answer and still leave the hint out when it explains itself. The J-lens gives a readout that does not rest on the model telling the truth about itself.
What this means for you if you are responsible for AI use
You are probably not building models yourself. What you do is put them to work, and answer for that use being sound. That makes this the sentence that matters most in the story: the explanation a model gives of its own reasoning is not evidence of what actually happened. If you have built a control around an AI tool explaining how it reached an answer, that control is weaker than it looks. Berigo assesses this as no reason to set AI aside, and every reason to stop treating a model own text as documentation.
The other side of the story is that interpretability is now something you can ask about. The method needs access to the internals of the model, so it is run by whoever owns the model, not by you as a customer. That makes the question to your supplier a concrete one: which evaluations and which tools for looking inside were used before a new model version went into service, and what did they find? Under NIS2 article 21 this belongs in risk management and in the requirements you place on suppliers. Under ISO/IEC 42001 it is the requirement that you know the AI system you are accountable for, across its whole life cycle. We would add that an answer you do not get in writing is not an answer you can point to in an audit.
Berigo recommends
- Stop counting a model own explanation as documentation of how an answer came about.
- Ask your supplier in writing which evaluations and which tools for looking inside were run before a new model version went into service.
- Require notice when the supplier changes model version, and treat the change as something to be assessed.
- Log what the AI system actually did, not only what it answered, so that a deviation can be investigated afterwards.
- Bring AI systems into risk assessment on the same terms as other systems, with an owner, a purpose and controls.
Security that is understood, governed and works.
Let us help you turn security into an advantage, not a cost. Get in touch for a no-obligation conversation about where your organisation stands and what to prioritise first.
Get in touch