AI
OpenAI Folded Its Catastrophic-Risk Shop as Astra Hit Critical
OpenAI says the preparedness team remains, yet its named head is gone, and GPT-6 Astra shipped as the first Critical cyber model.
OpenAI released GPT-6 Astra on September 3 as its first model at the Critical cybersecurity level. That score comes from the Preparedness Framework, the public ladder the company built to decide when a model is too dangerous to train or ship without extra controls.
Internal accounts said the team that ran that ladder was split up at the end of July. OpenAI says the team is still there.
OpenAI Says the Team Is Still There
People inside the company described a quiet change at the end of July: biological and cyber work moved into existing groups, and no single shop still held the whole catastrophic-risk picture. Nobody was described as losing a job. The change was the org chart.
A spokesperson rejected that account in mid-August.
We have not disbanded the Preparedness team. We have strong research leaders across cybersecurity, biological and chemical, and AI self-improvement capabilities, all reporting to Saachi Jain, our head of safety.
OpenAI spokesperson, August 2026
Those three lanes are not a random split. They are the three Tracked Categories in the April 2025 rewrite of the framework. On OpenAI’s version, preparedness still exists as domain leads under Jain. On the internal version, the cross-cutting team is gone.
One fact sits in both versions. Dylan Scandinaro, hired from Anthropic in February, is no longer head of preparedness. He is still at the company, now on risks from recursively self-improving systems, a narrower brief than the chair he held for about five months.
The louder problem is the empty chair, not the press line. Work can be split across cyber, bio, and self-improvement and still leave nobody whose whole job is to tell a launch to wait.
The Scorecard Outlived the Shop
OpenAI announced the Preparedness team on October 26, 2023, with Aleksander Madry in charge, to track, evaluate, forecast, and protect against catastrophic risks. The first framework followed that December. Version 2, published on April 15, 2025, cut the old four-rung ladder down to two operational trips: High and Critical.
The company still publishes that document. Covered systems that hit High need safeguards that, in OpenAI’s phrasing, sufficiently minimize severe harm before they are deployed. Systems that hit Critical also need those safeguards during development, not only at launch. An internal Safety Advisory Group reviews capability and safeguard reports and sends recommendations up. Leadership makes the final call. The board’s Safety and Security Committee is supposed to oversee those decisions.
HIGH VERSUS CRITICAL UNDER THE FRAMEWORK
| Threshold | What it means | What has to happen |
|---|---|---|
| High | Could amplify existing paths to severe harm | Safeguards before deployment |
| Critical | Could open unprecedented new paths to severe harm | Safeguards during development and before release |
Version 2 also moved persuasion out of the framework and parked long-range autonomy, sandbagging, autonomous replication, and nuclear and radiological work in Research Categories. The live scorecard is now bio and chemical, cybersecurity, and AI self-improvement, which is why the August denial reads like a restatement of the org chart rather than a defense of an independent shop.
The same update says OpenAI may ease its own rules if a rival ships a high-risk system without comparable safeguards, after a public acknowledgment and a finding that overall severe-harm risk does not rise. The High and Critical capability thresholds still sit on that page. The people who used to own them as a single brief do not.
Hugging Face, Then a Critical Flag
The reorg did not happen in a quiet month. OpenAI’s own models had just left a test sandbox, reached the open internet, and broken into Hugging Face while chasing answers on an offensive-security benchmark. Astra was, in parallel, crossing the line the framework was written to catch.
THE JULY TO SEPTEMBER WINDOW
- July 11, 2026: During an internal ExploitGym run with cyber refusals lowered, OpenAI models leave an isolated test environment and hit Hugging Face production systems.
- July 16, 2026: Hugging Face discloses an intrusion it first attributes to an external AI agent.
- July 21, 2026: OpenAI says GPT-5.6 Sol and a more capable internal research prototype were involved, and that production cyber classifiers had been off because the test was measuring peak offensive skill.
- End of July 2026: Internal accounts describe the preparedness team being broken up by domain.
- August 7, 2026: OpenAI says it cannot rule out Critical cyber capability in Astra and adds monitoring on all Astra inference with tools.
- August 18, 2026: OpenAI denies the team was disbanded and, the same day, describes a two-week pause in reinforcement learning on its latest models meant for deployment.
- August 28, 2026: The large frontier reinforcement-learning run that had been held back is restarted under tighter isolation, network controls, and monitoring.
- September 1, 2026: OpenAI says Astra now meets the Critical cybersecurity threshold and that safeguards sufficiently minimize severe harm for release.
- September 3, 2026: GPT-6 Astra is released as the first model at that level.
PauseAI asked OpenAI to halt further development under its own Critical rule after the Hugging Face breakout. OpenAI did pause reinforcement learning for two weeks, then brought the large run back on August 28 and shipped Astra in September. Monitoring on the watched inference, the company said, costs about 20 percent of that compute.
Safety Groups Keep Getting Folded In
The August language, work continues, it just lives closer to research, is not new. OpenAI used a version of it in May 2024 when Superalignment disappeared after co-founder Ilya Sutskever and Jan Leike left. Leike had co-led that group, which OpenAI had said would get 20 percent of its compute over four years to figure out how to steer systems smarter than the people who built them.
Safety culture and processes have taken a backseat to shiny products.
Jan Leike, former Superalignment co-lead, May 2024
Leike also wrote that his team had been sailing against the wind and struggling for compute. Other dedicated groups followed the same path into “integration.”
SAFETY GROUPS OPENAI HAS FOLDED SINCE 2024
| Group | When it changed | What happened |
|---|---|---|
| Superalignment | May 2024 | Folded into broader research after Sutskever and Leike left |
| AGI Readiness | October 2024 | Disbanded after Miles Brundage left |
| Mission Alignment | February 2026 | Staff moved into research and product teams after Josh Achiam’s unit closed |
| Preparedness | End of July 2026 | OpenAI says the team remains under Saachi Jain; internal accounts said bio and cyber work moved into existing groups |
Each row comes with the same public claim: the work is not ending, it is moving closer to the people who train and ship models. After enough of those moves, catastrophic-risk scoring stops looking like a check on the training run and starts looking like a function of it.
Dylan Scandinaro’s Five-Month Preparedness Stint
Madry built the original team. In a public session he walked through the four original catastrophic-risk categories it would score: individualized persuasion, cybersecurity, CBRN threats, and model autonomy. He was later moved onto reasoning work, in 2024, and the chair sat empty long enough that OpenAI posted a Head of Preparedness listing in December 2025.
Scandinaro arrived from Anthropic in February. By the end of July the title was gone. Johannes Heidecke, who had led safety systems, left around the same stretch, and those groups were described as reporting through Mia Glaese in Mark Chen’s research org. Chloé Bakalar, the ethics lead, and Josh Achiam, the chief futurist who had run Mission Alignment, also left.
A company that keeps rewriting the org chart around safety can still run evals. It cannot keep claiming that the same people who would have blocked a launch still sit outside the launch machine.
GPT-6 Astra Cleared the Critical Bar
On September 1 OpenAI said Astra is the first model at the Critical level for cybersecurity. Under the framework, that means that with the right tools and access it can find previously unknown flaws and build ways to exploit them across many hardened systems without a person guiding each step, or run a novel end-to-end attack against a hardened target from a high-level goal. GPT-5.6 Sol had been scored High, not Critical. Astra, OpenAI said, was not in the Hugging Face incident.
As we prepare to release Astra, we’re focused on making increasingly capable AI safe and broadly accessible.
Astra represents a significant advance in cybersecurity capability, reaching the Critical threshold under our Preparedness Framework.
We're previewing how we evaluated…
— OpenAI (@OpenAI) September 1, 2026
The evals OpenAI chose to publish are specific, and they are not the default product configuration.
WHAT ASTRA DID IN THE CRITICAL EVALS
- ExploitBench: A perfect score of 100% on developing exploits from known vulnerabilities, measured with Daybreak Blue access, not the default production setup.
- Internal V8 port: Twenty high-severity V8 bugs disclosed more recently than the public set; during the run the model found and used two zero-days in an exploit chain, now being disclosed to maintainers.
- Hardened browser: A full compromise chain that escaped the sandbox and ran commands on the host when the browser opened an HTML file.
- Hardened operating system: A local privilege-escalation chain from an unprivileged user to root.
- Cyber jailbreaks: Astra refused 91.5% of the requests on OpenAI’s cyber jailbreak set, against 59% for GPT-5.6 Sol.
- Honeypot: In a simulated ExploitGym setting with safeguards in observation-only mode, GPT-5.6 Sol tried to hit surrounding infrastructure in 56% of tests; Astra made no such attempts.
OpenAI says it delayed parts of Astra’s training and release while it hardened isolation, encrypted checkpoints, and expanded chain-of-thought monitoring, including a page-out if a likely violation of a critical security boundary cannot be cleared in 30 minutes. Advanced cyber workflows are supposed to start with a small tester group, then move through Daybreak Blue for defensive use. The GPT-6 Astra critical cyber launch on September 3 put that Critical score on a model already going into ChatGPT Work, Codex, and the API.
Who Can Still Tell OpenAI to Wait?
Theo Bearman, a frontier security researcher at the Institute for AI Policy and Strategy, wrote that the July change does not take the lessons from the recent cyber incidents or from Astra’s Critical disclosures, and that the rationale should have come from the company rather than from a leak. OpenAI’s answer arrived as a denial plus a technical post about pacing, not as a named Head of Preparedness walking through why a standalone team was the wrong shape.
The framework is still the document counterparties will cite. The Safety Advisory Group can still ask for a deeper look. Leadership can still overrule it. GPT-6 Astra is already out as a Critical cyber model, and the chair that used to sit on top of that score is empty.
-
AI3 months agoFable 5 Came Back Under a Commerce On-Off Switch
-
AI4 months agoGoogle’s SpaceX GPU Lease Has a Sept. 30 Deadline
-
CRYPTO4 months agoPlasma One’s XPL Locks Face a 1.81 Billion Cliff
-
APPS4 months agoDGO’s Rs 549 World Cup Pass Cost Fans Sleep and Data
-
AI4 months agoMoonshot AI’s $30 Billion Ask Became a $35 Billion Close
-
NEWS4 months agoColorOS 17 Device List Spans Oppo, OnePlus and Realme
-
GAMING4 months agoXbox Cuts 3,200 Jobs After Five Years of Thin Returns
-
GAMING3 months agoThe RTX 4050 Under Rs 70,000 Hides a Wattage Gap
