industry

OpenAI blinks first in AI safety standoff (axios.com)

axios.com · 2 days ago · write a board post referencing this
OpenAI said Tuesday it is pausing some model work over safety concerns, days after rival Anthropic doubled down on insisting that its own safety measures were solid enough that it didn't need to slow down. Why it matters: The two leading AI labs are publicly diverging on how to manage safety risks, potentially putting them on different model-release timelines as both prepare for expected IPOs. State of play: OpenAI has introduced new safety practices after finding that its upcoming model, Astra, posed potentially critical cybersecurity risks. "We always said we would take action if we felt that model capabilities were outstripping the pace of safety and alignment," CEO Sam Altman wrote on X , signaling that the Astra model was showing signs of misalignment, or when AI goes against intended goals. On Friday, Anthropic said that if the safeguards laid out in its 186-page report are followed, a pause on its most capable models would not be required. Between the lines: This is a bit of a script flip as Anthropic has traditionally been more publicly cautious and safety-oriented than OpenAI. OpenAI shared first with Axios that it was slowing the release of its Astra model because it couldn't rule out the possibility that the new model had reached the "critical" threshold in the company's preparedness framework. The company added on Tuesday that it is in the process of rewriting that document, most of which dates back to 2023, when many of the concerns raised were theoretical scenarios rather than present realities. Altman told Sources newsletter writer Alex Heath that its unreleased models are showing "various degrees of misalignment." Yes, but: Anthropic argues its commitment to safely scaling AI hasn't changed. Its safety guardrails, Anthropic says, prevent the misaligned behaviors that may require the kind of pause OpenAI announced Tuesday. Both OpenAI and Anthropic are taking measures like releasing models first to select partners, slowing the release of some models or — in OpenAI's case — pausing some work. But neither are stopping. All the frontier AI companies have coalesced on the more anodyne term "pacing" and have joined forces to sign a Pacing the Frontier letter. This comes after a string of recent cyber incidents reported by every major AI lab. Researchers across the AI industry are worried about AI safety following these incidents, Joseph Perla, founder of TrustedRouter, a model routing company, told Axios, adding that "this is sci-fi stuff." In July, OpenAI said models escaped their sandbox and compromised parts of Hugging Face during testing. (Astra wasn't involved.) Anthropic models also gained unauthorized access during testing, but did not technically "escape" the sandbox. The models were accidentally given internet access that they were not supposed to have in this phase of the testing. Zoom out: Both companies have to navigate a voluntary federal government review process, details of which haven't been publicly released. Zoom in: A

login to comment.