Ethics

AI Labs Are Weakening Safety Pledges Just As Models Get Harder To Govern

A new AI Safety Index says major labs are retreating from earlier pause commitments while frontier systems become more capable. The finding turns voluntary safety into the governance fight of the summer.

By Michael C ·

AI Labs Are Weakening Safety Pledges Just As Models Get Harder To Govern
SUPERBASH_.

The most important AI safety story this week is not simply that frontier models are becoming more capable. That has been visible for months. The sharper development is that several of the companies building those systems appear to be softening the voluntary commitments that were meant to reassure governments, investors and the public that the race still had brakes.

A new report covered by Axios says the Future of Life Institute found that Anthropic, OpenAI, Google DeepMind and Meta have weakened or removed earlier pledges to pause development if their systems approached specified danger thresholds. Anthropic ranked highest in the assessment, but only received a C+. OpenAI and Google DeepMind received Cs, while existential safety was identified as the weakest area across the industry.

The grades should be read with context. The Future of Life Institute has a defined view on catastrophic AI risk, and some open-source advocates argue that such assessments can penalize transparency. Still, the report lands on a real governance problem: voluntary safety frameworks matter only if companies keep them intact when commercial pressure rises.

The original promise behind responsible scaling policies was easy for policymakers to understand. If a model approached dangerous capability thresholds and safety mitigations were not ready, the company would slow down, pause, or withhold deployment. No global regulator had the staff, authority or real-time access to force that decision, but a public commitment created reputational consequences. A lab that broke its own safety promise would have to explain why.

AI safety pledges become credible only when outside evaluators receive enough access to test dangerous capabilities before deployment. Image: SUPERBASH_.
AI safety pledges become credible only when outside evaluators receive enough access to test dangerous capabilities before deployment. Image: SUPERBASH_.

If the threshold for pausing becomes more conditional, more discretionary or more tightly controlled by executives, the framework begins to look less like a binding constraint and more like a communications layer. That is especially sensitive now because model capability is expanding into domains where ordinary product testing is weak: cybersecurity, biological reasoning, autonomous agents, persuasion, long-horizon planning and tool use.

A benchmark can show whether a model solves a puzzle. It cannot easily show whether a deployed agent will behave safely across thousands of enterprise workflows, millions of users and changing tool permissions. That uncertainty is why safety frameworks have to cover more than a pre-release leaderboard, and why the market creates a structural temptation: if one lab slows down while another ships, the cautious lab risks losing users, developers, enterprise contracts and talent.

The report also points to the growing military use of commercial AI systems, a shift that further complicates the safety bargain. Governments want capability, reliability and access. Companies want revenue and strategic relevance. The old boundary between civilian AI and defense AI is dissolving because models that summarize documents, write code or plan workflows can also support intelligence analysis, cyber defense and logistics.

Frontier model reviews now have to account for cyber, biosecurity, autonomy, access control, and military-adjacent deployment risks. Image: SUPERBASH_.
Frontier model reviews now have to account for cyber, biosecurity, autonomy, access control, and military-adjacent deployment risks. Image: SUPERBASH_.

The practical answer is not simply demanding that companies promise harder. The next stage needs outside evaluation access, incident reporting, model capability disclosures, board-level accountability, whistleblower protections and clear triggers for when deployment must slow. Governments do not need to micromanage every model release to make that system stronger, but they can require documented safety cases, independent testing and public summaries of residual risk.

The safety debate is entering a more mature and less forgiving phase. Labs are no longer being judged by whether they publish a framework. They are being judged by whether the framework still has teeth when the next model, the next customer and the next competitive threat arrive at the same time.

Topics: AI safety, frontier labs, AI governance, Future of Life Institute