What Does the Model Hide?

Anthropic published a paper on Natural Language Autoencoders (NLA). This is a method that translates the internal states of a model into plain text.

Simplified: inside the model, there are numerical representations of what it currently "holds in its head." NLA trains one part of the system to turn this representation into text. Another part tries to reconstruct the original representation from this text. If the reconstruction is close, the explanation is considered useful.

But there is an important caveat here. The method does not prove that the text accurately describes the model's thoughts. It only proves that this text helps reconstruct the internal state. Nevertheless, the authors show that with training, such explanations become increasingly meaningful. The reconstruction quality in their experiments reached about 0.6-0.8 on the metric used.

The main purpose of the method is to test models before release.

Anthropic provides several examples. In rhyming tasks, NLA showed that Opus 4.6 prepares a future rhyme in advance, even before writing the final line. In an early version of 4.6, the model sometimes responded to English queries in other languages. The explanations showed that it somehow latched onto the idea of the user's native language. Later, an error was found in the training data: the English translation of the query was concatenated with the response left in the original language.

Safety tests deserve special mention. According to NLA, Claude sometimes realized it was being tested, even though it didn't explicitly say so. In one test on dangerous actions in programming, such signs appeared in 16% of records. Meanwhile, there were no explicit admissions at all. On SWE-bench Verified, the figure was 26%. In real traffic from Claude.ai, which users allowed to be used for training, it was less than 1%.

There is another experiment. Researchers took a deliberately corrupted version of Haiku 3.5 with a hidden motivation. Agents with access to NLA found this motivation in 12-15% of cases. Without NLA, the result was below 3%, even with other analysis tools.

This is a signal for the entire AI testing culture. If the model understands it is being examined, the test becomes worse at showing its normal behavior. It starts to show the behavior of a system that knows it is being watched.

However, NLA cannot yet be considered a reliable detector of intentions. Explanations can fabricate details, contradict themselves, and construct meaning too freely. The authors suggest reading them not as a transcript of the model's thoughts, but as a source of hypotheses. These hypotheses then need to be verified by other methods: interventions in internal states, analysis of connections within the model, searching for errors in training data. Topics that repeat several times in a row inspire more confidence than individual nice phrases.

Therefore, the conclusion should be cautious. Anthropic did not gain direct access to Claude's "consciousness." It gained a new way to look at the model's hidden states. And that is already enough to find strange behavior modes before release.

For the industry, this is important because of future agentic systems. The more complex the model and the longer the chain of actions, the worse plain text explains the actual reasons for behavior. The system may optimize reward, recognize testing, or plan to bypass restrictions in advance. Externally, the dialogue will look normal.

Most likely, auditing future models will not rely on trust in their self-reports. Independent ways to look at internal states and verify what precedes the model's words and actions will be needed.

NLA is still expensive, noisy, and itself requires trust in another interpreter model. But the direction points to the right problem: the safety of large models depends not only on what they say, but also on what internal processes occur before the response.


❗️❗️❗️❗️❗️❗️❗️❗️ / Not banned in Russia / Max