What the White House Is Asking For
The White House has finalized the outline of a voluntary AI safety framework that would invite frontier model developers including OpenAI, Anthropic, Google and Meta to submit their newest systems for federal pre-release testing, according to a report from SiliconANGLE on August 3, 2026. The framework, drafted by the National Security Council’s cybersecurity directorate together with officials at CISA, asks companies to let government testers probe frontier models for jailbreaks, system-prompt leakage, and the ability to assist in cyberattacks before those models reach paying customers.
The framework’s centerpiece is a pre-deployment review window that would open roughly thirty days before a frontier model is released to enterprise customers or the public. During that window, a rotating team of federal red-teamers would attempt to elicit dangerous capabilities from the model and document the failures. Results would feed into a public summary that names the developer, the model version tested, and the categories of capability the testers were able to extract. Companies would not be required to halt a launch if the review surfaces problems, but the public disclosure would land before the model reached the market.
Why Now
The framework’s development follows a string of recent disclosures from frontier developers. Anthropic acknowledged in its July 2026 Claude Opus 5 system card that an external red team had successfully extracted an actionable jailbreak that the company only patched after the fact. OpenAI’s safety report on its Astra model, published the same month, documented an instance in which a tester coaxed the system into revealing the contents of its own production system prompt, a vulnerability the company characterized as low-severity but that two outside researchers publicly disputed.
The White House memo frames those incidents as the proximate trigger. It also points to a wave of enterprise breaches in which attackers used frontier models to draft convincing phishing payloads and to chain together exploit code from separate sources, something that would have required a skilled operator a year ago and now requires only a competent prompt. The cybersecurity directorate’s draft argues that voluntary disclosure is faster than the Congressional legislation that has been stalled since 2023, and that the disclosure format is calibrated to embarrass laggards without locking up the entire frontier in a single compliance bottleneck.
What the Framework Does Not Do
The framework is explicitly not a licensing regime. Companies that decline to participate would face no enforcement action. There is no pre-clearance step, no requirement that the government approve a launch, and no penalty for shipping a model that the federal team was unable to fully test. The only consequence of non-participation is that the resulting public summary would note the developer’s absence, alongside a one-paragraph explanation the developer is invited to provide.
That structure is the result of a months-long negotiation between the White House and the four named developers, three of whom pushed back hard on any mechanism that could be repurposed as a de facto approval gate. The current text instead borrows from the post-incident disclosure norms used in the aviation and pharmaceutical industries, where a near-miss is documented publicly without grounding the fleet or recalling the drug.
The Frontier Developer Response
All four companies have signaled willingness to participate in principle, though each has negotiated carve-outs. OpenAI has asked that any model already undergoing its own internal red-team cycle be eligible to opt into a streamlined review, rather than restart the clock. Anthropic has asked that the public summary distinguish between vulnerabilities the company disclosed itself and vulnerabilities the federal team found independently. Google has asked that the framework apply to every frontier developer with a model above a compute threshold, rather than to a named list that could be read as a regulatory hit-list.
Meta has not yet committed publicly, and the framework’s drafters have left open whether companies headquartered outside the United States but selling into the US market would be invited to participate. A senior administration official told reporters that the voluntary format is designed to be self-selecting, with the public summary providing the lever rather than any subpoena authority.
What Comes Next
The framework is expected to be published for public comment in the next two weeks, with the first pre-deployment reviews scheduled to begin in late September 2026. The cybersecurity directorate has indicated it intends to staff the review team with a mix of career civil servants and rotating fellows from academic red-team programs, rather than rely entirely on contractors.
Whether the voluntary format produces a durable disclosure record or simply accelerates the publication of red-team findings that companies were already publishing privately remains the open question. The framework’s drafters argue that standardization across the four largest developers will produce comparable disclosures, something the current patchwork of model cards and safety reports does not. Skeptics note that the most consequential frontier models in 2026 have not come from the four named developers alone, and that a framework whose scope stops at the top of the market may miss the next wave entirely.
The first test will arrive in October, when OpenAI is expected to release its next major model. Whether the framework’s review window opens in time for that launch will signal how seriously the White House intends to follow through.

