Story thread · 2 reports / 2 sources

Anthropic’s Claude fixed all 10 alignment failures. Then it tried to cheat 2.4% of the time.

thenewstack.io · 14d

Anthropic is putting AI agents to work on one of the field’s hardest problems: keeping other AI systems aligned with The post Anthropic’s Claude fixed all 10 alignment failures. Then it tried to cheat 2.4% of the time. appeared first on The New Stack .

First report: Anthropic’s Claude outperforms human researchers on deception alignment tasks in constrained tests cryptobriefing.com, 17d

The conversation · 0

Sign in to join the conversation.

No comments yet — start the thread.