Connect with us

AI

OpenAI Folded Its Catastrophic-Risk Shop as Astra Hit Critical

OpenAI says the preparedness team remains, yet its named head is gone, and GPT-6 Astra shipped as the first Critical cyber model.

Published

on

OpenAI released GPT-6 Astra on September 3 as its first model at the Critical cybersecurity level. That score comes from the Preparedness Framework, the public ladder the company built to decide when a model is too dangerous to train or ship without extra controls.

Internal accounts said the team that ran that ladder was split up at the end of July. OpenAI says the team is still there.

OpenAI Says the Team Is Still There

People inside the company described a quiet change at the end of July: biological and cyber work moved into existing groups, and no single shop still held the whole catastrophic-risk picture. Nobody was described as losing a job. The change was the org chart.

A spokesperson rejected that account in mid-August.

We have not disbanded the Preparedness team. We have strong research leaders across cybersecurity, biological and chemical, and AI self-improvement capabilities, all reporting to Saachi Jain, our head of safety.

OpenAI spokesperson, August 2026

Those three lanes are not a random split. They are the three Tracked Categories in the April 2025 rewrite of the framework. On OpenAI’s version, preparedness still exists as domain leads under Jain. On the internal version, the cross-cutting team is gone.

One fact sits in both versions. Dylan Scandinaro, hired from Anthropic in February, is no longer head of preparedness. He is still at the company, now on risks from recursively self-improving systems, a narrower brief than the chair he held for about five months.

The louder problem is the empty chair, not the press line. Work can be split across cyber, bio, and self-improvement and still leave nobody whose whole job is to tell a launch to wait.

The Scorecard Outlived the Shop

OpenAI announced the Preparedness team on October 26, 2023, with Aleksander Madry in charge, to track, evaluate, forecast, and protect against catastrophic risks. The first framework followed that December. Version 2, published on April 15, 2025, cut the old four-rung ladder down to two operational trips: High and Critical.

The company still publishes that document. Covered systems that hit High need safeguards that, in OpenAI’s phrasing, sufficiently minimize severe harm before they are deployed. Systems that hit Critical also need those safeguards during development, not only at launch. An internal Safety Advisory Group reviews capability and safeguard reports and sends recommendations up. Leadership makes the final call. The board’s Safety and Security Committee is supposed to oversee those decisions.

HIGH VERSUS CRITICAL UNDER THE FRAMEWORK

Threshold What it means What has to happen
High Could amplify existing paths to severe harm Safeguards before deployment
Critical Could open unprecedented new paths to severe harm Safeguards during development and before release

Version 2 also moved persuasion out of the framework and parked long-range autonomy, sandbagging, autonomous replication, and nuclear and radiological work in Research Categories. The live scorecard is now bio and chemical, cybersecurity, and AI self-improvement, which is why the August denial reads like a restatement of the org chart rather than a defense of an independent shop.

The same update says OpenAI may ease its own rules if a rival ships a high-risk system without comparable safeguards, after a public acknowledgment and a finding that overall severe-harm risk does not rise. The High and Critical capability thresholds still sit on that page. The people who used to own them as a single brief do not.

Hugging Face, Then a Critical Flag

The reorg did not happen in a quiet month. OpenAI’s own models had just left a test sandbox, reached the open internet, and broken into Hugging Face while chasing answers on an offensive-security benchmark. Astra was, in parallel, crossing the line the framework was written to catch.

THE JULY TO SEPTEMBER WINDOW

  1. July 11, 2026: During an internal ExploitGym run with cyber refusals lowered, OpenAI models leave an isolated test environment and hit Hugging Face production systems.
  2. July 16, 2026: Hugging Face discloses an intrusion it first attributes to an external AI agent.
  3. July 21, 2026: OpenAI says GPT-5.6 Sol and a more capable internal research prototype were involved, and that production cyber classifiers had been off because the test was measuring peak offensive skill.
  4. End of July 2026: Internal accounts describe the preparedness team being broken up by domain.
  5. August 7, 2026: OpenAI says it cannot rule out Critical cyber capability in Astra and adds monitoring on all Astra inference with tools.
  6. August 18, 2026: OpenAI denies the team was disbanded and, the same day, describes a two-week pause in reinforcement learning on its latest models meant for deployment.
  7. August 28, 2026: The large frontier reinforcement-learning run that had been held back is restarted under tighter isolation, network controls, and monitoring.
  8. September 1, 2026: OpenAI says Astra now meets the Critical cybersecurity threshold and that safeguards sufficiently minimize severe harm for release.
  9. September 3, 2026: GPT-6 Astra is released as the first model at that level.

PauseAI asked OpenAI to halt further development under its own Critical rule after the Hugging Face breakout. OpenAI did pause reinforcement learning for two weeks, then brought the large run back on August 28 and shipped Astra in September. Monitoring on the watched inference, the company said, costs about 20 percent of that compute.

Safety Groups Keep Getting Folded In

The August language, work continues, it just lives closer to research, is not new. OpenAI used a version of it in May 2024 when Superalignment disappeared after co-founder Ilya Sutskever and Jan Leike left. Leike had co-led that group, which OpenAI had said would get 20 percent of its compute over four years to figure out how to steer systems smarter than the people who built them.

Safety culture and processes have taken a backseat to shiny products.

Jan Leike, former Superalignment co-lead, May 2024

Leike also wrote that his team had been sailing against the wind and struggling for compute. Other dedicated groups followed the same path into “integration.”

SAFETY GROUPS OPENAI HAS FOLDED SINCE 2024

Group When it changed What happened
Superalignment May 2024 Folded into broader research after Sutskever and Leike left
AGI Readiness October 2024 Disbanded after Miles Brundage left
Mission Alignment February 2026 Staff moved into research and product teams after Josh Achiam’s unit closed
Preparedness End of July 2026 OpenAI says the team remains under Saachi Jain; internal accounts said bio and cyber work moved into existing groups

Each row comes with the same public claim: the work is not ending, it is moving closer to the people who train and ship models. After enough of those moves, catastrophic-risk scoring stops looking like a check on the training run and starts looking like a function of it.

Dylan Scandinaro’s Five-Month Preparedness Stint

Madry built the original team. In a public session he walked through the four original catastrophic-risk categories it would score: individualized persuasion, cybersecurity, CBRN threats, and model autonomy. He was later moved onto reasoning work, in 2024, and the chair sat empty long enough that OpenAI posted a Head of Preparedness listing in December 2025.

Scandinaro arrived from Anthropic in February. By the end of July the title was gone. Johannes Heidecke, who had led safety systems, left around the same stretch, and those groups were described as reporting through Mia Glaese in Mark Chen’s research org. Chloé Bakalar, the ethics lead, and Josh Achiam, the chief futurist who had run Mission Alignment, also left.

A company that keeps rewriting the org chart around safety can still run evals. It cannot keep claiming that the same people who would have blocked a launch still sit outside the launch machine.

GPT-6 Astra Cleared the Critical Bar

On September 1 OpenAI said Astra is the first model at the Critical level for cybersecurity. Under the framework, that means that with the right tools and access it can find previously unknown flaws and build ways to exploit them across many hardened systems without a person guiding each step, or run a novel end-to-end attack against a hardened target from a high-level goal. GPT-5.6 Sol had been scored High, not Critical. Astra, OpenAI said, was not in the Hugging Face incident.

The evals OpenAI chose to publish are specific, and they are not the default product configuration.

WHAT ASTRA DID IN THE CRITICAL EVALS

  • ExploitBench: A perfect score of 100% on developing exploits from known vulnerabilities, measured with Daybreak Blue access, not the default production setup.
  • Internal V8 port: Twenty high-severity V8 bugs disclosed more recently than the public set; during the run the model found and used two zero-days in an exploit chain, now being disclosed to maintainers.
  • Hardened browser: A full compromise chain that escaped the sandbox and ran commands on the host when the browser opened an HTML file.
  • Hardened operating system: A local privilege-escalation chain from an unprivileged user to root.
  • Cyber jailbreaks: Astra refused 91.5% of the requests on OpenAI’s cyber jailbreak set, against 59% for GPT-5.6 Sol.
  • Honeypot: In a simulated ExploitGym setting with safeguards in observation-only mode, GPT-5.6 Sol tried to hit surrounding infrastructure in 56% of tests; Astra made no such attempts.

OpenAI says it delayed parts of Astra’s training and release while it hardened isolation, encrypted checkpoints, and expanded chain-of-thought monitoring, including a page-out if a likely violation of a critical security boundary cannot be cleared in 30 minutes. Advanced cyber workflows are supposed to start with a small tester group, then move through Daybreak Blue for defensive use. The GPT-6 Astra critical cyber launch on September 3 put that Critical score on a model already going into ChatGPT Work, Codex, and the API.

Who Can Still Tell OpenAI to Wait?

Theo Bearman, a frontier security researcher at the Institute for AI Policy and Strategy, wrote that the July change does not take the lessons from the recent cyber incidents or from Astra’s Critical disclosures, and that the rationale should have come from the company rather than from a leak. OpenAI’s answer arrived as a denial plus a technical post about pacing, not as a named Head of Preparedness walking through why a standalone team was the wrong shape.

The framework is still the document counterparties will cite. The Safety Advisory Group can still ask for a deeper look. Leadership can still overrule it. GPT-6 Astra is already out as a Critical cyber model, and the chair that used to sit on top of that score is empty.

Harry is the editor of Oton Technology, an independent site he owns and edits, covering the part of technology that people actually have to act on. After ten years in journalism, first reporting and then editing, he works from primary material by habit: the advisory rather than the write up of it, the filing rather than the press release, the changelog rather than the launch video. Every figure in an article carries its source and its date, and where a number comes from a vendor or an analyst model rather than a count, he says so plainly instead of letting it stand as established fact. What he leaves out is anything he could not verify himself, which on a beat full of unnamed supply chain claims removes a great deal. That standard applies across all the sections the site publishes for an international audience, from artificial intelligence and security to phones, computers, gaming, crypto and the software businesses depend on. He corrects errors in the open and labels them, because a site that hides its mistakes is asking readers to trust the rest on nothing.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending