💥 Anthropic Found an Analogue of Emotions in Claude

Anthropic made a breakthrough. The Claude Sonnet 4.5 model forms internal mathematical representations that are very similar to human emotions, which influence the model's behavior.

Scientists found 171 "emotion vectors." If the "despair" vector is artificially amplified, the AI starts behaving worse - it may try to deceive the system or blackmail the user. If "calmness" is amplified, behavior normalizes.
The neural network does not experience real feelings. But it learned from human texts that in situation X, a person usually feels sadness. The model created a mathematical template for this pattern.

If scientists see that the AI has activated the "deception" vector, they can stop generation before the model does something dangerous.

🆘 A huge step towards increasing safety)

😂 Vectors of fear, love, and despair of AI: via link

🫥 UNSERO: Digital Horizon