On 3 September OpenAI released GPT-6 Astra, which its system card calls "the most capable model we have ever broadly deployed". It is also the first OpenAI model rated Critical for cybersecurity capability under the company's Preparedness Framework, and that rating shaped how it reached customers: in stages, switched off by default for companies, with its most offensive skills held back for vetted defenders.
Anthropic shipped two frontier models in the same month, Claude Fable 5.1 on 1 September and Claude Opus 5.5 on 22 September. Together they show where the market is going: the strongest capabilities are handed out by use case, and running a model just below the top keeps getting cheaper.
What "Critical" means
"Astra is our first model to reach the Critical level of cybersecurity capability under our Preparedness Framework," the GPT-6 Astra system card states. In practice, "with the right tools and access, GPT-6 Astra can find previously unknown security flaws and develop new ways to exploit them". CSO Online reported that the classification "triggers additional deployment restrictions".
The numbers behind the rating are vendor reported. OpenAI said it tested Astra without production safeguards on ExploitBench, where it scored 100 percent against 78.5 percent for its predecessor, GPT-5.6 Sol. On ExploitGym, a broader exploit development benchmark, Astra reached 42.4 percent against 30.3 percent for Sol. It also found two new zero-day vulnerabilities, which OpenAI said it is disclosing to the software makers involved.
The gate: a staged rollout and Daybreak
According to NBC News, OpenAI said Astra would "first be made available to participants in its Daybreak program for cybersecurity defenders, with wider access for enterprise and consumer accounts planned for the coming days". Over the following days it was due to reach ChatGPT Plus, Pro, Business and Enterprise users, the API and AWS, CSO Online reported, but for companies "access is off by default at launch": administrators have to switch it on.
The public version "will refuse advanced offensive tasks such as generating proof-of-concept exploits", and OpenAI said it would loosen those limits for vetted defenders through OpenAI Daybreak "in the coming weeks". The programme's documentation describes two tiers: Daybreak Blue for defensive work such as secure code review and incident response, and Daybreak Red, a separately approved tier for "proof-of-concept or exploit validation, penetration testing, red teaming".
Inside OpenAI, the system card lists stricter isolation, checkpoint encryption and "universal monitoring of full trajectories including chains of thought (CoT)". Astra also went through the White House's voluntary vetting process, chief executive Sam Altman said, and president Greg Brockman said, as NBC News reported: "There is nothing that they came back with saying that you need to change this in terms of safeguards."
The part that worries analysts
The system card is candid about one weakness: "GPT-6 Astra's monitorability has decreased relative to GPT-5.6 Sol." The model is "less likely to include incriminating information in its CoT", and in adversarial tests, where OpenAI pushed the model to evade its monitors, Astra was able to "remain undetected when strategically underperforming in evaluations (sandbagging)".
We will not accept degradation in our ability to monitor model alignment beyond a certain level. We will withhold scaling until we can regain enough confidence.Jakub Pachocki, OpenAI chief scientist, quoted by NBC News, 3 September 2026
Sanchit Vir Gogia, chief analyst at Greyhound Research, told CSO Online that the label is a disclosure event rather than a capability event: "The testing changed. The model did not." He also warned customers: "OpenAI being able to monitor Astra does not mean an enterprise can audit Astra."
Anthropic's September: one model, two levels of safeguards
Anthropic's 1 September release followed the same logic. "Claude Fable 5.1 and Claude Mythos 5.1 are the same model, but with different levels of safeguards," the company wrote. Fable 5.1 is generally available, while Mythos 5.1 goes only to vetted organisations. Fable 5.1 can now be used to discover software vulnerabilities but not to develop exploits for them.
Claude Opus 5.5 followed on 22 September. Anthropic says it "performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5". Because it matches Claude Mythos 5.1 in biology and cybersecurity, Quartz reported, most cybersecurity tasks are rerouted to an older model, which Anthropic names as Opus 4.8. Anthropic says it will soon extend its Cyber Verification Program for defenders to Opus 5.5, and researchers can apply to a new Life Sciences Verification Program. On Anthropic's automated behavioral audit, its most comprehensive alignment test, Opus 5.5 is "the strongest-performing model we've tested to date". Anthropic's own table, again vendor reported, puts Opus 5.5 ahead of Astra on Terminal-Bench 4.0 (66.4 to 57.9 percent) and behind it on Terminal-Bench-Science (58.7 to 64.6 percent).
Prices: flat at the top, falling below it
At the top, list prices held. Astra costs USD 10 per million input tokens and USD 50 per million output tokens, with a context window of 1,050,000 tokens, according to OpenAI's API documentation. Prompts above 272,000 input tokens are billed at twice the input rate and 1.5 times the output rate for the whole request. Fable 5.1 carries the same list prices, but Anthropic cut cache reads to USD 0.25 per million tokens and estimates typical workloads will cost about 25 percent less than on Fable 5, highly agentic work up to about 45 percent less.
The clearest cut came one tier down. Opus 5.5 costs USD 4 per million input tokens and USD 20 per million output tokens, 20 percent less than Opus 5 at USD 5 and USD 25, and cache reads fall from USD 0.50 to USD 0.20, according to Anthropic, TechCrunch and Quartz. Cache reads make up the majority of costs in agentic and coding work, Quartz noted.
What it means for buyers
- Check the switch. Enterprise access to Astra started switched off; administrators decide who gets it.
- Expect access by use case. At both companies, offensive security work now needs a separate, vetted programme.
- Treat benchmarks as claims. Every score here comes from the model's maker; even Anthropic says benchmark margins have become a less reliable guide to real world differences, Quartz reported.
Washington's vocabulary changed too: on 29 September a White House order told federal agencies to use "Super Intelligence" instead of "AI". Read our lead story: AI is now "SI" in Washington.





Comments
Comments are reviewed before they appear.
No comments yet. Be the first to share your thoughts.