Industry·2 min read
By BitsMindsSource: HPCwire

Commerce's CAISI Signs Pre-Deployment Frontier AI Testing Pacts With Google DeepMind, Microsoft, and xAI

The Center for AI Standards and Innovation will get pre-release access to frontier models from Google DeepMind, Microsoft, and xAI — including versions with safeguards stripped — to evaluate cyber, bio, and chemical risk.

Commerce's CAISI Signs Pre-Deployment Frontier AI Testing Pacts With Google DeepMind, Microsoft, and xAI
Share:

The Center for AI Standards and Innovation (CAISI), the AI evaluation arm housed inside the Commerce Department's National Institute of Standards and Technology, announced new frontier AI testing agreements on May 5, 2026 with Google DeepMind, Microsoft, and xAI. Under the deals, all three labs will give the U.S. government pre-deployment access to their highest-capability models for security and capability evaluations, with additional assessments continuing after public release.

The agreements broaden the scope of what CAISI can probe. Evaluators will assess "demonstrable risks" tied to national security — explicitly including cybersecurity, biosecurity, and chemical-weapons capabilities — and will be allowed to test in classified environments. To make those tests meaningful, the developers have agreed to provide model variants with reduced or removed safety guardrails, the same approach that government red-teamers have argued is necessary to surface ceiling-level capabilities rather than the post-mitigation behavior end-users see.

"These expanded industry collaborations help us scale our work in the public interest at a critical moment," CAISI Director Chris Fall said in a statement. The center has now completed more than 40 evaluations, and its remit was reset earlier this year to align with Commerce Secretary Howard Lutnick's directives and the administration's AI Action Plan, which leans more heavily on national security framing than the framework CAISI's predecessor body operated under in 2024.

The new arrangement updates partnerships that the U.S. AI Safety Institute, CAISI's predecessor, signed with OpenAI and Anthropic in August 2024. Those original deals were focused on voluntary safety testing of frontier models. The renegotiated terms with Google DeepMind, Microsoft, and xAI fold in security and capability assessment work that CAISI now treats as a formal pre-deployment gate, even though compliance remains voluntary in the absence of federal AI legislation.

For the labs, the deals are a hedge. Pre-deployment evaluations let them get ahead of likely regulatory scrutiny without conceding to a binding licensing regime, while also giving CAISI a clearer picture of which capabilities — autonomous cyber operations, dual-use biology, agentic tool use — are crossing thresholds that warrant new policy. Anthropic, notably, has been pushing for stronger formal oversight, and OpenAI's existing arrangement remains in place; the May 5 announcement explicitly positions the new agreements as additions to, not replacements for, the broader frontier testing program.

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.

Related Articles

417–3 in the House. One objection in the Senate. Editorial illustration of the stalled Ratepayer Protection Act, H.R. 9340. A large 417–3 House vote meets a single red senator and a closed stop line. Below, AI data centres drawing 100 MW or more connect through a transmission tower to a home and an electricity bill. Amber conduits link the industrial load, grid upgrades and household costs. The proposed large-load standard would assign full incremental grid-upgrade costs to data centres, but a senator blocked unanimous consent, objecting that states were only asked to consider the standard. H.R. 9340 RATEPAYER PROTECTION ACT 417–3 HOUSE VOTE ONE SENATOR. BILL BLOCKED. The objection: states only asked to “consider”. AI DATA CENTRES 100 MW+ GRID UPGRADES WHO PAYS? BILL $ HOUSEHOLD BILLS BITSMINDS.COM
Industry

House Votes 417–3 to Make Data Centers Pay for the Grid

Hacktron: from an image upload to an internal repository An oversized HEIC photo file is connected by an electric green line to two glass security gates, labelled Forum and SSO, then to a dark repository cabinet bearing the OpenAI emblem and a pull-request symbol. A separate Claude Opus 5 plaque credits the model used by the researchers. The top caption says disclosed and patched. This is a conceptual illustration of Hacktron's reported security research and vulnerability chain. The two gates represent the two reported flaws, and their open padlocks are a visual metaphor. Original vector editorial illustration for BitsMinds, hacktron-claude-opus-5-openai-exploit-chain. 18 September 2026. Conceptual architecture based on the article, not a real interface or technical system diagram. HACKTRON / SECURITY RESEARCH DISCLOSED & PATCHED .HEIC IMAGE UPLOAD 01 FORUM IMAGE PROCESSING 02 SSO SIGN-IN TRUST INTERNAL REPO PULL REQUEST OPENED CLAUDE OPUS 5 RESEARCH ASSISTANT TWO FLAWS. ONE CHAIN. BITSMINDS.COM
Industry

Claude Opus 5 Wrote the Exploit That Broke Into OpenAI

Pew's global AI survey: uncertainty about the future of work An empty ochre office chair faces a computer in a quiet evening office. Beyond the window, buildings of unequal heights suggest an uneven economy. A large globe on the left is surrounded by 37 separate markers: 34 amber and three neutral. They represent surveyed publics, not individual respondents, countries mapped to locations, job counts or measured job losses. The caption states that in 34 of 37 publics, more people expect job losses than gains. This is an editorial illustration of public expectations reported in the Pew survey, not a prediction of what AI will actually do. Original vector illustration for BitsMinds, pew-global-ai-survey-2026-jobs-inequality. 17 September 2026. Survey finding, not observed labour-market outcomes. PEW / GLOBAL AI SURVEY 34 / 37 PUBLICS More expect job losses than gains AI BITSMINDS.COM
Industry

Most of the World Expects AI to Cost Jobs, Pew Finds