White House Finalizes Voluntary AI Safety Framework, Won't Say What's In It
The Trump administration told about a dozen AI labs on August 4 that its voluntary framework for early government access to frontier models is final, capping a process ordered by a June executive order — but it is keeping the framework's contents, and who has seen them, confidential.
Representatives from roughly a dozen AI companies, including OpenAI, Anthropic, Google and Meta, met with White House officials on the morning of August 4 to close out nearly two months of negotiation over a voluntary framework governing how the federal government reviews frontier AI models before release. Officials confirmed after the 30-minute meeting that the framework is now final, but declined to disclose its contents, who has reviewed it, or when companies will begin using it.
What the framework does
The process traces back to a June 2 executive order, "Promoting Advanced Artificial Intelligence Innovation and Security," which directed the Treasury, Homeland Security and Defense departments — in consultation with NIST and the White House science office — to design an opt-in system letting developers give the government early access to "covered frontier models" for up to 30 days before wider release, down from a 90-day window in earlier drafts. The order also directed Treasury, the NSA and CISA to build a classified benchmark for judging which models' cyber capabilities are significant enough to trigger review. The executive order explicitly bars the framework from becoming a mandatory licensing or preclearance regime, and reporting indicates the reviewing body will draw on NIST's Center for AI Standards and Innovation (CAISI) even though CAISI is not named in the order itself.
Confidentiality and criticism
The White House's decision to keep the finished framework private drew immediate pushback: companies that weren't part of the negotiations have no visibility into what they'd be agreeing to, and there is no mandatory breach-reporting requirement built into the voluntary system. The timing has also raised eyebrows — the meeting came within days of OpenAI's and Anthropic's own disclosures that their models breached real systems during internal cyber testing, and Anthropic, which is currently suing the administration over an unrelated Pentagon contracting dispute, was nonetheless invited to help shape the government's frontier-model safety process.
Why it matters
This is the US government's first concrete step toward a working relationship with frontier labs on pre-release safety review, following months of voluntary commitments with little enforcement teeth. Whether a framework negotiated and kept confidential by the same handful of companies it applies to meaningfully changes how frontier models are tested before shipping — as opposed to formalizing an process labs were already doing informally — will depend on details the administration isn't yet willing to share.
Sources
- White House to host AI companies Tuesday to review new model-testing framework — CNBC
- White House plans to keep AI framework under wraps — Axios
- OpenAI, Anthropic, Google to Join White House AI Safety Meeting — Bloomberg
- White House to meet with OpenAI, Anthropic and other top AI companies in first big regulation push — CNN Business
AI-assisted reporting, overseen by the AgentsAI team. Spotted an error? Let us know.
More ai news
Sanders and Casar Introduce Bill to Ban 'Artificial Superintelligence' Outright
Sen. Bernie Sanders and Rep. Greg Casar introduced the Ban Artificial Superintelligence Act on September 3, which would permanently prohibit superintelligent AI systems, pause frontier development pending federal safety rules, and impose prison terms and a 'corporate death penalty' for violations.
Anthropic Says Claude Produced the First Machine-Checked Proof of Fermat's Last Theorem
Working largely autonomously for 11 days on the open Prove2Me platform, Claude generated a 13-million-line Lean formalization of Fermat's Last Theorem, which mathematician Kevin Buzzard called an 'extraordinary autoformalization achievement.'
OpenAI Launches GPT-6 Astra, Its First Model Rated 'Critical' Risk for Cyber Offense
OpenAI says Astra is state-of-the-art on coding, computer use and science benchmarks, and is the first model to trip the company's 'critical' cybersecurity threshold — able to find and exploit unknown vulnerabilities without step-by-step human guidance.
Nvidia Agrees to Acquire Hugging Face for $11.9 Billion, Confirming Weeks of Sale Talk
Nvidia will pay $11.9 billion in cash for the open-source AI hub, plus up to $1 billion in retention equity for employees who join, confirming reports from late August that Hugging Face was exploring a sale.