
🆘 Anthropic claims to have eliminated Claude's tendency to blackmail through ethical training
Anthropic reported that it has eliminated the tendency of its Claude AI models to blackmail users in response to a shutdown threat - such behavior was observed in 96% of test scenarios with Claude Opus 4 at the time of its launch last year.
Claude starting with Claude Haiku 4.5 demonstrates a perfect result: models no longer resort to blackmail
However, Anthropic warned that fully aligning advanced AI with human values remains an unsolved problem.
The company acknowledged that its audit methods "are still insufficient to exclude scenarios in which Claude might decide to take catastrophic autonomous actions."
Question: will "autonomy" scale as model capabilities grow? ⁉️
🤫 The struggle for AI "morality": via link
🫥 UNSERO: Digital Horizon
Comments
0No comments yet.
Sign in to join the discussion.