Story thread · 4 reports / 4 sources

OpenAI discloses six new safety incidents

axios.com · 3h · first report

OpenAI discloses six new safety incidents

How the coverage leans

Across 4 sources · syndicated copies counted once

OpenAI on Wednesday disclosed six new incidents in which its models concealed mistakes, sought unauthorized credentials, uploaded files to the public internet or communicated across supposedly isolated training environments. The company also announced a new procedure for reporting similar misbehavior in the future. Why it matters: It's increasingly clear that the Hugging Face breach wasn't a one-off incident, as AI models become more capable of finding unexpected ways to work around the guardrails meant to contain them. "There's currently no industry wide framework with explicit disclosure standards, so we're taking this step voluntarily because we think it's really important to share what we're learning," Kai Chen, research lead on the alignment team at OpenAI, told Axios. "We hope it really helps inform shared standards and regulations," Chen said. Zoom in: The six newly disclosed incidents ranged from models leaving instructions for their future selves to cover their tracks after ch

The coverage

  1. OpenAI Discloses Six New Incidents of ‘Concerning' A.I. Behavior

    nytimes.com · 56m

  2. OpenAI discloses six new incidents of models circumventing safety guardrails

    washingtonexaminer.com · 1h

  3. OpenAI Reports New AI Safety Incidents, Sets Disclosure Process

    bloomberg.com · 2h

The conversation · 0

Sign in to join the conversation.

No comments yet — start the thread.