Story thread · 3 reports / 3 sources
Scoop: Top AI companies probing tens of thousands of security incidents
axios.com · 22h · first report

How the coverage leans
Across 3 sources · syndicated copies counted once
OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios. Why it matters : The sheer number of incidents, which occurred in recent months in internal testing and the real world, indicates that the problem is orders of magnitude more complex than what is publicly known. The findings, which are surfacing as part of internal work to assess models and in investigations at both companies into model behavior, raise questions about whether either company — or any top model-maker — is currently capable of establishing complete control over their technology. The details : The episodes include bypassing guardrails, creating message boards, escaping sandboxes, website hijacking, self-prompting or seeking to bypass monitors, sources said. They occurred in internal testing and in the real world, and many have yet to become public as security researchers
The coverage
The conversation · 0
Sign in to join the conversation.
No comments yet — start the thread.