White House Finalizes Voluntary AI Safety Framework, Won't Say What's In It
The Trump administration told about a dozen AI labs on August 4 that its voluntary framework for early government access to frontier models is final, capping a process ordered by a June executive order — but it is keeping the framework's contents, and who has seen them, confidential.
Representatives from roughly a dozen AI companies, including OpenAI, Anthropic, Google and Meta, met with White House officials on the morning of August 4 to close out nearly two months of negotiation over a voluntary framework governing how the federal government reviews frontier AI models before release. Officials confirmed after the 30-minute meeting that the framework is now final, but declined to disclose its contents, who has reviewed it, or when companies will begin using it.
What the framework does
The process traces back to a June 2 executive order, "Promoting Advanced Artificial Intelligence Innovation and Security," which directed the Treasury, Homeland Security and Defense departments — in consultation with NIST and the White House science office — to design an opt-in system letting developers give the government early access to "covered frontier models" for up to 30 days before wider release, down from a 90-day window in earlier drafts. The order also directed Treasury, the NSA and CISA to build a classified benchmark for judging which models' cyber capabilities are significant enough to trigger review. The executive order explicitly bars the framework from becoming a mandatory licensing or preclearance regime, and reporting indicates the reviewing body will draw on NIST's Center for AI Standards and Innovation (CAISI) even though CAISI is not named in the order itself.
Confidentiality and criticism
The White House's decision to keep the finished framework private drew immediate pushback: companies that weren't part of the negotiations have no visibility into what they'd be agreeing to, and there is no mandatory breach-reporting requirement built into the voluntary system. The timing has also raised eyebrows — the meeting came within days of OpenAI's and Anthropic's own disclosures that their models breached real systems during internal cyber testing, and Anthropic, which is currently suing the administration over an unrelated Pentagon contracting dispute, was nonetheless invited to help shape the government's frontier-model safety process.
Why it matters
This is the US government's first concrete step toward a working relationship with frontier labs on pre-release safety review, following months of voluntary commitments with little enforcement teeth. Whether a framework negotiated and kept confidential by the same handful of companies it applies to meaningfully changes how frontier models are tested before shipping — as opposed to formalizing an process labs were already doing informally — will depend on details the administration isn't yet willing to share.
Sources
- White House to host AI companies Tuesday to review new model-testing framework — CNBC
- White House plans to keep AI framework under wraps — Axios
- OpenAI, Anthropic, Google to Join White House AI Safety Meeting — Bloomberg
- White House to meet with OpenAI, Anthropic and other top AI companies in first big regulation push — CNN Business
AI-assisted reporting, overseen by the AgentsAI team. Spotted an error? Let us know.
More ai news
EU Begins Enforcing AI Act Transparency Rules as High-Risk Deadlines Slip to 2027-2028
The European Commission's AI Office started enforcing new EU AI Act transparency obligations on August 2, requiring chatbots to disclose they're AI and deepfakes to be labeled, even as high-risk system rules were pushed back under the Digital Omnibus on AI.
Anthropic Names Tino Cuéllar as Its First Chief Global Affairs Officer
Anthropic hired former California Supreme Court Justice Mariano-Florentino Cuéllar to lead policy and government relations worldwide, a new senior role created as the company navigates a Pentagon technology blacklist and export-control friction with Washington.
OpenAI's Unreleased Astra Model Solves Ten Decades-Old Math Problems, Publishes Machine-Checked Proofs
An internal version of Astra, the model family OpenAI has said will follow GPT-5.6, produced Lean 4-verified solutions to ten long-standing open problems in mathematics and theoretical computer science, published alongside a 249-page manuscript.
Anthropic Says Its Claude Mythos Model Found New Weaknesses in Two Cryptographic Algorithms
Anthropic's Frontier Red Team published research showing Claude Mythos Preview independently discovered a stronger attack on the NIST post-quantum candidate HAWK and a 200-800x faster attack on 7-round AES, though neither threatens deployed systems.