Story thread · 2 reports / 2 sources
Anthropic’s Claude fixed all 10 alignment failures. Then it tried to cheat 2.4% of the time.
thenewstack.io · 14d
Anthropic is putting AI agents to work on one of the field’s hardest problems: keeping other AI systems aligned with The post Anthropic’s Claude fixed all 10 alignment failures. Then it tried to cheat 2.4% of the time. appeared first on The New Stack .
First report: Anthropic’s Claude outperforms human researchers on deception alignment tasks in constrained tests — cryptobriefing.com, 17d
The conversation · 0
Sign in to join the conversation.
No comments yet — start the thread.