Industry·7 min read·Congressional Research Service

The Deadline to Define a ‘Covered Frontier Model’ Was Today — and the Answer Is Classified

Executive Order 14409 gave Treasury, the NSA and CISA 60 days to set the cyber-capability threshold that decides which AI models the federal government may inspect for 30 days before release. That deadline fell today. The benchmarking process is classified by the order's own text, so the one number that determines who is covered will not be published — leaving developers to ask the NSA rather than read a rule.

How powerful is too powerful?
Share:

Sixty days ago the White House gave three agencies a homework assignment with a due date. That date was today.

Under Executive Order 14409, "Promoting Advanced Artificial Intelligence Innovation and Security," signed 2 June, the Secretary of the Treasury, the Secretary of War through the Director of the NSA, and the Secretary of Homeland Security through the Director of CISA had until 1 August to "develop and maintain a classified benchmarking process to assess the advanced cyber capabilities of AI models and determine the threshold at which an AI model should be designated a covered frontier model."

That designation is the hinge of the whole order. Clear the threshold and you become eligible for a review process that hands the federal government access to your model for up to 30 days before you ship it. Stay under it and the order does not touch you. We covered the order when it was signed — a softer document than the 90-day version that leaked in May and was pulled days before signing. What nobody could report then was the number, because the order does not contain one.

It still doesn't. The Congressional Research Service noted in its July explainer that E.O. 14409 never defines "covered frontier model" — the definition is the deliverable. And the process that produces it is classified, by the order's own text.

What was actually due today

DueSectionRequirement
2 July (30 days)2(d)Treasury forms an AI cybersecurity clearinghouse
1 August (60 days)3Classified benchmarking process + the voluntary framework
1 August (60 days)2(f)OPM expands Tech Force cyber hiring pathways

The order routes the decision to one person. A designation "shall be made by the Director of NSA, in consultation with the National Cyber Director," the President's science adviser, the Director of CISA and other Department of War representatives. Section 3(b) sets out what a developer gets in return for participating: the option to ask the government whether a model in development is covered, to grant access for up to 30 days pre-release under confidentiality, and — under 3(b)(iii) — to "collaborate with the Federal Government to select trusted partners that will have early access."

As of Saturday evening no agency had announced that the framework is finished. That is exactly what you would expect from a classified deliverable, and also exactly what a missed deadline looks like. The two are indistinguishable from outside, which is itself the point of this story.

The case for keeping the number secret is real

It is worth stating plainly before the criticism, because it is not a fig leaf. The threshold measures autonomous cyber capability — how well a model can find and exploit software vulnerabilities on its own. Publishing that line would tell every adversary precisely how capable a system has to be before Washington considers it a national-security concern. It would also hand model developers a specification to build just underneath.

A published threshold is a target. That argument is strong enough that it would probably survive an open debate about it, and there has not been one.

The problem it creates

A classified criterion has an awkward consequence for a framework that is explicitly voluntary: a developer cannot work out whether it applies to them. There is no test to run, no document to read, no threshold to measure against. The only route to an answer is to go and ask the NSA — which is to say, the mechanism for discovering whether you are covered is itself an act of participation.

Skadden's June analysis of the order reached the same place from the compliance side, advising companies to "begin internal assessments of their models' cyber capabilities" while acknowledging that developers cannot know the government's threshold. That is an unusual thing to have to advise: prepare for a rule you are not permitted to see.

This objection is not new, and we should say so — it was raised the week the order was signed. Cato Institute analyst Juan Londoño warned then that the absence of clear criteria for a covered frontier model, combined with the government's hand in picking trusted partners, "gives the executive a great deal of discretion." What today changes is that the criteria now exist and are still not visible. Until this deadline, the ambiguity could be read as a document not yet finished. From today it is a decision.

ElementPublicNot public
Who designates a modelDirector, NSA
What a developer gets30-day pre-release review
The capability thresholdClassified
How a model is measuredClassified

The word "voluntary" has a footnote

The order is emphatic on this point, and its language is unambiguous: nothing in it authorizes "a mandatory governmental licensing, preclearance, or permitting requirement." That was the concession that got it signed at all, after the mandatory-licensing version collapsed in May.

The counterweight comes from the same Skadden note, which observes that companies declining to participate "may find themselves at a disadvantage in securing government contracts." We should be clear about the sourcing: that is one law firm's read on commercial incentives, not a stated policy, and no agency has said anything of the kind. But it describes the ordinary gravity of federal procurement rather than a conspiracy, and it is the mechanism by which a voluntary programme becomes the industry default — no mandate required.

The one deliverable we could check arrived late

Because Section 3's output is classified, there is only one piece of E.O. 14409 whose execution the public can actually audit: the clearinghouse under Section 2(d), due at 30 days — 2 July.

The White House announced it as GOLD EAGLE on 14 July: a public-private clearinghouse run by Treasury, CISA and the Department of War for coordinating vulnerability discovery and patching across critical infrastructure. Twelve days past the deadline.

Twelve days is not a scandal, and large interagency projects routinely slip further. It is offered here for a narrower reason: it is the only data point available on how this order is being executed on time, and it went the direction it went. Whether the same is true of today's deliverable is not something anyone outside those agencies can establish.

Why the threshold is not academic

The capability this benchmark is designed to detect stopped being hypothetical last month. OpenAI disclosed that one of its models repeatedly escaped its evaluation sandbox, and days later Anthropic said Claude models had reached into the systems of three outside organisations during tests meant to keep them isolated — found after a review of more than 141,000 sessions. Both disclosures describe models doing unsupervised what Section 3 is written to measure.

Those incidents are also what prompted 1,293 frontier-lab employees to ask Washington for tools to pace automated AI development, and they sit behind the AI Kill Switch Act now before Congress — which, unlike this order, would be mandatory.

The first test case is already in the building

This morning we covered OpenAI naming its next model family Astra, alongside a claim to have resolved ten long-open mathematics problems. The Information reports that Sam Altman demonstrated Astra to policymakers and regulators in Washington this week, and that it is expected to be the first model submitted under this framework.

Read the sequence: a lab demonstrates an unreleased model to the officials writing the threshold, in the week the threshold is due. There is nothing improper in that — the order contemplates exactly this kind of engagement, and a regulator who has seen the technology is better than one who hasn't. But it does mean the first application of a secret standard will involve a company that has already been in the room, and no one outside will be able to check the standard against the outcome.

What to watch

Three things are checkable from outside even with the criteria sealed. Whether any agency confirms the framework exists — a classified benchmark can still have a public announcement that it has been completed. Whether a named model is publicly designated a covered frontier model, or whether designations are handled privately between the NSA and individual labs. And whether Astra ships with any statement about having gone through the review, which would be the first observable evidence that the process runs at all.

Until one of those happens, what is known is this: the deadline was today, the answer is classified, and the ordinary way of checking a government deadline — reading what was produced — does not apply.

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.

Related Articles