On May 5 2026 the US AI Safety Institute (now operating under the CAISI banner — Center for AI Standards and Innovation, the renamed and restructured successor to the original AISI under the Biden-era Executive Order 14110) signed pre-deployment evaluation agreements with three frontier AI labs: Google DeepMind, Microsoft Research, and xAI. The agreements grant CAISI evaluators access to new frontier models before public release, for the purpose of safety testing across a defined set of capability evaluations.

The signed parties matter. Google, Microsoft, and xAI are three of the five major US-headquartered frontier model developers. The two missing names from the May 5 list — OpenAI and Anthropic — had earlier partnership agreements with CAISI's predecessor (AISI) that are now in renegotiation under the Trump administration's revised AI Action Plan posture. The May 5 announcement is the first new pre-deployment eval signing since the administration transition, which makes the structure of the agreements informative about where federal AI oversight is heading.

This is a regulation-shaped change to how frontier models reach the market. The shape matters.

The Specific Terms

The May 5 agreements published in summary form (full text not released) cover four distinct elements:

Pre-deployment access: CAISI evaluators receive access to new frontier-tier model checkpoints at a defined window before public release. The exact window is unspecified in the public summary but is described as "sufficient for meaningful evaluation" — the original 2024 AISI agreement specified a 30-day window, and the May 5 agreement is rumored to use a similar timeframe.

Capability evaluation scope: Evaluations focus on dual-use risk categories defined in the National Security Memorandum on AI — CBRN (chemical, biological, radiological, nuclear) uplift, cybersecurity capability, autonomy, and persuasion. The scope is explicitly bounded to national-security-relevant capabilities and does not extend to general performance evaluations.

Information sharing: Findings are shared with the relevant lab privately first, with public disclosure of aggregate-level results on a delayed timeline. Specific vulnerabilities or capabilities are not publicly disclosed in detail. This is a meaningful change from the original AISI structure, which allowed for more detailed public reporting.

Voluntary structure: The agreements remain voluntary. Labs can withdraw with notice. No regulatory enforcement mechanism is attached to the agreements. This is also a change from the trajectory under the prior administration, which had been moving toward mandatory evaluations through forthcoming rulemaking.

The combination — pre-deployment access, narrowed scope, private-first findings, voluntary structure — describes a federal evaluation regime that is more cooperative and less adversarial than the trajectory under the prior administration. The May 5 agreements are best read as the Trump administration's signal of how it wants the AI evaluation layer to function.

Why Google, Microsoft, And xAI Signed First

The composition of the May 5 signing is informative. Google and Microsoft are the two large incumbent tech companies most exposed to federal regulation across multiple product lines beyond AI — antitrust posture, cloud contracts, defence-related work. Signing pre-deployment evaluation agreements is partially a relationship-management decision for these companies, even if the capability evaluations themselves are narrow.

xAI is the more interesting signing. The company has positioned itself politically as friendlier to the current administration than other frontier labs. Signing the agreement publicly demonstrates that friendliness while also extracting the legitimacy benefit of being included in the federal evaluation regime. The strategic logic is unusually clean.

OpenAI and Anthropic are the two notable absences. Both had existing relationships with the predecessor AISI under the prior administration and both are now in renegotiation. The renegotiation is rumored to focus on whether the agreements survive in similar form or whether the labs use the transition to revise terms. Neither lab has publicly indicated whether they will sign the May 5-style agreement or hold out for different terms.

The pattern: the labs with the strongest political alignment with the current administration signed first. The labs with stronger historical ties to the prior administration are negotiating. The labs that did not have prior agreements (some open-weight model developers) are not part of the discussion.

What This Means For Model Release Timelines

The practical question for builders is whether the May 5 agreements affect when new models reach the public. The answer is "marginally yes for the signed labs, less than the headlines suggest."

The 30-day pre-deployment window (assuming it matches the prior AISI structure) is short relative to the development timelines of frontier models. A model that finishes training in April typically reaches public release in May or June. A 30-day evaluation window inserted into that timeline pushes the release date by 30 days at most, and the lab can compress the gap between training completion and evaluation kickoff to minimise the net delay.

For Google specifically, the Gemini 3.1 Pro release on April 12 2026 happened before the May 5 agreement was signed. The next major Gemini release (likely Gemini 3.2 Pro in Q3) will be the first model that flows through the new evaluation process. The expected delay is on the order of two to four weeks compared to the unconstrained release timeline.

For Microsoft's Phi-4 line and the GPT-OSS forthcoming releases, the effect is similar. For xAI's Grok 4 line, the agreement may actually accelerate release readiness by providing a public legitimacy signal that the model is safety-evaluated.

The net effect across the three signed labs is a small marginal delay — measured in weeks, not months — applied to frontier model releases. This is unlikely to materially change the cadence of releases, but it does set a precedent that other labs may follow.

What This Does Not Cover

The narrowed scope is the part of the agreements that gets least attention and matters most. CAISI's evaluations cover national-security-relevant capabilities only. The agreements do not cover:

- Bias, fairness, or discrimination concerns - Misinformation or content-policy concerns - Economic displacement effects - Environmental impact of training compute - Concentration of capability in a small number of labs - Open-source model release decisions - API-level safety controls (jailbreaking, refusal behaviour)

These categories were part of the broader AI policy conversation under the prior administration and were partially within scope of the original AISI's mission. The May 5 agreements narrow the federal evaluation focus specifically to dual-use national-security risks. Other concerns are deferred to other agencies (FTC on bias, FCC on misinformation, etc.) or to industry self-governance.

For builders, this scope matters in two ways. First, the federal evaluation regime will not catch issues outside its defined scope, which means concerns like bias and misinformation are unlikely to delay model release. Second, the labs themselves remain responsible for managing the un-covered risk categories, which means lab-level governance and policy decisions become more, not less, important for the overall safety posture of frontier model releases.

The International Picture

The May 5 agreements are US-specific. The international AI evaluation landscape has continued to develop in parallel:

- UK AI Security Institute (AISI-UK, separate institution): Continues to operate pre-deployment evaluations under voluntary agreements with multiple labs. The UK AISI has historically been more aggressive in publishing specific capability findings. - EU AI Office: Operating under the EU AI Act framework, with mandatory rather than voluntary evaluation requirements for systemic-risk models. The first wave of mandatory evaluations is in progress through Q2 2026. - Japan AI Safety Institute: Established in early 2025, evaluations are voluntary and focused on Japanese-language and Japan-relevant capability concerns. - Singapore AI Verify Foundation: Operates a different model — auditable evaluation framework rather than direct pre-deployment access.

The US position under the May 5 agreements is now substantially more lab-friendly than the EU position and roughly comparable to the UK position. Labs operating in both jurisdictions face different evaluation regimes for the same models, which adds compliance complexity but is unlikely to change deployment decisions materially.

The international fragmentation is the part of the story that may compound. If the US, EU, UK, Japan, and other jurisdictions each maintain different evaluation regimes with different scopes, frontier labs eventually face a coordination cost that affects release timelines and possibly capability scope. The May 5 agreements do not address this fragmentation. Whether the Trump administration pursues international harmonisation is an open question.

Signals Worth Tracking

Three specific signals to watch through Q2 and Q3:

Whether OpenAI and Anthropic sign similar agreements. The renegotiation is the most-watched policy item in the category. If both labs sign before the next Opus or GPT release, the federal evaluation regime achieves full coverage of US-headquartered frontier labs. If either lab walks away from agreements entirely, that signals a more adversarial posture that would shape the next 12-24 months of federal-lab relations.

Whether the evaluation findings stay private. The May 5 structure puts findings into private-first communication with delayed public disclosure of aggregate results. If specific findings leak — through Congressional testimony, lab announcements, or media reporting — the regime's voluntary structure becomes harder to maintain.

Whether mandatory evaluation requirements develop. The Trump administration's AI Action Plan has signalled a preference for voluntary structures, but Congressional action or new executive orders could shift toward mandatory requirements. The trajectory through Q3 will tell.

The May 5 agreements are a concrete step in the federal AI governance picture. They are narrower in scope, more lab-friendly in structure, and more voluntary in enforcement than the trajectory under the prior administration. For frontier labs operating in the US market, the practical implication is a manageable evaluation overhead on new releases and a more cooperative relationship with the federal evaluation layer. For the broader AI policy conversation, the implication is that issues outside the dual-use national-security scope will be handled outside the federal evaluation regime, which leaves significant policy gaps to be filled by other mechanisms.