Connect with us

AI

White House Orders CAISI to Stop Publishing AI Safety Reports

Five weeks after the White House ordered CAISI to halt public AI model evaluations, the agency’s reports remain dark with no return date set.

Published

on

The Trump administration told the government’s main AI testing lab to stop publishing what it finds, and five weeks later, the silence hasn’t lifted. National Cyber Director Sean Cairncross ordered the halt at the Center for AI Standards and Innovation, or CAISI, in early June, according to The Wall Street Journal.

The directive arrived as a new executive order on AI security took effect, folding CAISI’s public-facing work into a wider, classified review process with no announced date for its reports to return.

CAISI Goes Quiet

CAISI is short for the Center for AI Standards and Innovation. It is the rebranded successor to what the Biden administration called the U.S. AI Safety Institute, a body created in 2023 to study risk in frontier AI systems.

The name changed in June 2025. Underneath, the job stayed close to the same: probe advanced models for cyber and biological risks before they reach the public, under the National Institute of Standards and Technology, or NIST, inside the Commerce Department.

Administration officials, Cairncross among them, told CAISI to halt publication of its model assessments while the new order takes hold. The move marked a win for Cairncross and Treasury Secretary Scott Bessent, who had pushed for security concerns to weigh more heavily in how models get evaluated.

The directive stops short of a full shutdown. Three things haven’t changed:

  • Testing continues. CAISI’s internal evaluation program keeps running.
  • Agreements hold. Its voluntary arrangements with AI developers remain in place.
  • No new mandate. The directive creates no additional legal requirement for companies to submit models.

What changed is where the findings go. Results that once became public write-ups now stay inside government channels. Before the halt, that public record included an April review of the Chinese open-weight model, when CAISI evaluated the open-weight model DeepSeek V4 Pro, one of the clearer examples of the agency measuring a rival country’s AI against its own government’s benchmarks.

Why Mythos Changed Washington’s Math

The trigger sits with a single model. Anthropic announced in April that it would limit distribution of its new Claude Mythos Preview model after discovering it could find and exploit software vulnerabilities most trained engineers would miss.

The disclosure rattled both the AI industry and federal officials. OpenAI’s GPT-5.5 drew similar scrutiny soon after, flagged for the same kind of vulnerability-hunting skill.

Cairncross himself has called Mythos “the model right now that everyone’s talking about,” acknowledging in April that the administration was already holding meetings on its cyber risks and benefits. He took over the Office of the National Cyber Director in August 2025, making him one of the newer voices now shaping frontier AI policy.

Trump had been close to signing a cybersecurity order built around a 90-day government review window for the most powerful models, then pulled back in May over worries it would slow American labs against Chinese competitors. He signed a narrower version on June 2, titled “Promoting Advanced Artificial Intelligence Innovation and Security,” cutting that review window to 30 days.

Anthropic was already sparring with the administration on a separate front. The Defense Department had slapped its tools with a “supply chain risk” designation after the company declined to loosen restrictions on domestic surveillance and fully autonomous weapons use, and the White House ordered federal agencies to phase out Anthropic tools entirely. Anthropic is fighting that order in court.

The same day Trump signed the June 2 order, Anthropic expanded Mythos access from roughly 50 organizations to about 200.

The Five Labs Now Working in the Dark

Five companies hold voluntary testing agreements with CAISI: OpenAI, Anthropic, Google DeepMind, Microsoft and xAI. All five submit frontier systems for government review before wide release.

On May 5, three of those companies signed a renegotiated deal that moved their testing into classified environments. Google DeepMind, Microsoft and xAI agreed to undergo safety testing in classified environments, a step up from the earlier, more open arrangement.

“Independent, rigorous measurement science is essential to understanding frontier AI and its national security implications,” CAISI Director Chris Fall said when the deal was announced.

Microsoft described its side of the arrangement as collaborative work to test Microsoft’s frontier models, framing government-level review as something no single company can fully replicate in house.

Lab CAISI Testing Status Recent Flashpoint
OpenAI Voluntary evaluation partner since the agency’s founding GPT-5.5 flagged for vulnerability-hunting skill
Anthropic Mythos submitted for review; access expanded to about 200 groups Fighting a Pentagon “supply chain risk” label in court
Google DeepMind Signed renegotiated classified pact on May 5, 2026 Models now tested inside classified government environments
Microsoft Signed the same May 5 pact Says in-house testing cannot replace government review
xAI Signed the same May 5 pact One of three labs moved to classified access this spring

OpenAI and Anthropic’s ties to the agency reach back further than the May deals, to research relationships signed when the office still operated as the Biden-era AI Safety Institute.

The Commerce Department Already Tried This Once

This isn’t the first time CAISI’s public-facing work has quietly disappeared.

Days after the May 5 agreements were announced, the Commerce Department page describing the testing framework vanished. CAISI staff were told to take down the page with no explanation, a source told Axios, and the old link redirected to CAISI’s general site, stripped of any detail about the arrangement.

The pattern traces back further than that. The Biden administration created the AI Safety Institute in 2023 under an executive order the Trump White House rescinded within hours of taking office in January 2025. Commerce Secretary Howard Lutnick rebranded the survivor as CAISI five months later.

“For far too long, censorship and regulations have been used under the guise of national security,” Lutnick said in the statement announcing the rebrand.

Industry groups largely welcomed it at the time. The Information Technology Industry Council welcomed CAISI’s focus on voluntary testing, calling continued reliance on voluntary evaluation good for U.S. AI leadership.

What Does the New Executive Order Actually Require?

The order does not create a licensing regime for AI. It gives Treasury, the National Security Agency (NSA), the Cybersecurity and Infrastructure Security Agency (CISA) and NIST 60 days to design a classified process for identifying which models count as “covered frontier models,” and it lets the government examine those models for up to 30 days before they reach wide release.

Nothing in the order forces a company to participate. Developers can still release on their own timeline, though doing so invites the kind of scrutiny Anthropic now faces.

Not everyone inside the administration is happy with how that authority got assigned. Atlantic Council analysts noted the order creates a parallel track for model evaluation that duplicates work CAISI was already doing.

The fact that the Department of the Treasury is the lead agency acting as a clearinghouse is deeply concerning

Nick Leiserson, senior vice president for policy at the Institute for Security and Technology, raised that objection shortly after the order was signed, adding that neither AI nor cybersecurity are core competencies of the department.

Lawfare, a national security policy publication, reported that some officials and industry figures worry Cairncross “lacks the expertise to lead on such a technically complex and emergent national security issue.”

The disagreements run in more than one direction.

  • Cairncross and Bessent pushed for the halt and the new classified framework, treating it as overdue given what Mythos showed.
  • Rival officials inside the administration believe the order simply hands Cairncross authority over evaluation work CAISI already performed.
  • Outside critics question why Treasury, rather than CISA or NIST, ended up running the new clearinghouse.

The Classified Benchmark Deadline Is Still Weeks Away

The order’s 60-day clock runs out August 1, 2026. Until then, the agencies designing the classified system have no legal obligation to show their work, and CAISI has no public deadline to resume publishing anything.

President Trump signaled on July 3 how far he’s willing to go with companies that resist. Asked about dangerous AI actors, he said if the situation turned even slightly dangerous, “we will move quickly and effectively to stop that actor,” in remarks widely read as aimed at Anthropic.

Two days later, Sriram Krishnan, a former White House AI policy adviser, told the Financial Times that “there will not be an AI equivalent of the U.S. Food and Drug Administration,” arguing centralized pre-approval would cost the U.S. its competitive edge.

Whichever agency runs the classified process once August arrives, CAISI’s public output stays exactly where it’s been since early June: at zero.

Frequently Asked Questions

What is the Center for AI Standards and Innovation?

CAISI is the Commerce Department’s primary AI testing body, housed inside NIST. It began in 2023 as the U.S. AI Safety Institute under a Biden-era executive order, then was renamed and refocused toward national security and global competitiveness by Commerce Secretary Howard Lutnick in June 2025. It also runs newer programs like the AI Agent Standards Initiative, launched in February 2026 to guide how autonomous AI agents get secured.

Does the halt mean CAISI stopped testing AI models?

No. CAISI’s internal evaluation work continues, and its voluntary agreements with developers remain active. What changed is where the results go: findings now stay inside government channels, reaching other federal agencies and presumably the company being tested, rather than becoming a public write-up outside researchers can read.

Which AI companies work with CAISI?

Five developers, OpenAI, Anthropic, Google DeepMind, Microsoft and xAI, hold voluntary testing agreements with CAISI covering their frontier models. That’s separate from the AI Safety Institute Consortium, a group of more than 200 companies, universities and civil-society organizations established in 2024 that advises the agency but doesn’t submit models for classified review.

What happened to evaluations CAISI finished before the halt?

They remain part of CAISI’s internal record. Its April review of the Chinese open-weight model DeepSeek V4 Pro, published while the agency still put out public write-ups, is one of the last examples of that work reaching outside readers.

When does the government’s classified AI review process take effect?

Treasury, the NSA, CISA and NIST have until August 1, 2026, 60 days after the executive order was signed, to finish designing the classified benchmarking process that decides which models qualify as “covered frontier models.” Hitting that deadline sets the criteria. It does not obligate CAISI to resume public reporting, since the order sets no date for that to happen.

Has CAISI published anything publicly since the halt?

Not that has surfaced. No new public CAISI evaluation has appeared since the directive took hold in early June, and the agency has not said when, or whether, that will change.

Logan Pierce is a writer and web publisher with over seven years of experience covering consumer technology. He has published work on independent tech blogs and freelance bylines covering Android devices, privacy focused software, and budget gadgets. Logan founded Oton Technology to publish clear, no nonsense tech news and reviews based on real hands on testing. He has personally tested and reviewed dozens of mid range and budget Android phones, written extensively about app privacy, and built and managed multiple WordPress publications over the past decade. Logan holds a bachelor's degree in English and studied digital marketing at a certificate level.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending