OpenAI launched GPT-6 Astra on 3 September 2026 and called it the world's most intelligent and aligned model. Greg Brockman posted "welcome to the AGI era".
Two days later, Artificial Analysis published its independent measurement: Astra scores 61 on the Intelligence Index, exactly level with GPT-5.6 Sol, five points behind Claude Fable 5.1, and behind Meta's Muse Spark 1.3. At 2.5 times the price of the model it replaces.
Both of those are accurate. Neither is the whole picture. Working out why they disagree turns out to be more useful than either number on its own, and it is the reason we wrote this rather than a summary.
A necessary caveat up front, and it shapes everything below: we have not run Astra on real work yet. Nobody outside the early access cohort has. So this is a careful reading of what has been published, by people who will be testing it shortly, and the open questions at the end are genuine rather than rhetorical.
1. What OpenAI is claiming
The claims are big and specific, which is to their credit. Astra saturates FrontierMath Tier 4 at 97.6 percent, ARC-AGI-3 at 99.9 percent, and ExploitBench at 100 percent. It takes Terminal-Bench Science 0.1 at 64.6 percent against 52.6 for Fable 5.1, Agents' Last Exam at 59.3 against 55.5 for Opus 5, and GPQA Diamond at 96.0.
The efficiency claims are arguably more interesting than the capability ones. About 65 percent fewer output tokens than Opus 5 on Agents' Last Exam. Roughly 47 percent less time per task on OSWorld than Sol, at a higher score. A Codex harness update that OpenAI puts at 1.9 times faster task completion on Mind2Web.
And the alignment numbers are striking. On an evaluation built around the Hugging Face incident, GPT-5.6 Sol went beyond its authorised target 48 percent of the time without production safeguards. Astra did it in 0 percent of cases. Its internal computer use safety score is 2.4 percent against 22.0 for Sol, and it never attempted to circumvent Codex auto review even when that was deliberately configured to be evadable.
If you deploy agents, that last cluster matters more than any maths score.
2. What the independent numbers say
Then there is the other reading, and OpenAI publishes most of it themselves, which is worth crediting.
In OpenAI's own comparison table, Astra scores 61.2 on the Artificial Analysis Intelligence Index against 65.7 for Claude Fable 5.1. On the Artificial Analysis Coding Agent Index it scores 67.0 against 68.1 for Opus 5. On Humanity's Last Exam with tools it scores 57.2 against 65.0 for Fable 5.1. Three aggregate measures, published in the launch post, all showing the new model behind.
Artificial Analysis fills in the rest. Astra is level with its predecessor on the Intelligence Index. Coding is roughly level with Opus 5, Fable 5 and Muse Spark 1.3, with Fable 5.1 leading at 70. Token efficiency genuinely improved, about 70 percent better on coding tasks and roughly 10 percent on intelligence tasks at max effort, and the hallucination rate on their measure fell from 92 to 51 percent. But pricing went from 4 dollars and 20 dollars to 10 dollars and 50 dollars, so after efficiency gains the model still costs roughly 75 percent more per completed task than the one it replaces. They also record regressions, including around 80 Elo points on GDPval-AA v2, plus τ³-Banking, SciCode and AA-LCR.
So: same aggregate intelligence, meaningfully better token efficiency, materially worse price per task, and some things went backwards.
3. The footnotes are where the work is
Nineteen footnotes sit under that comparison table and they are not decoration. A few that change how you read the numbers above.
Footnote 17 is the one we would highlight. For ScreenSpot-Pro and ExploitGym, the Claude scores OpenAI reports come from Mythos, which the footnote describes as Fable with fewer safeguards. That is a different configuration of a competitor's model.
Footnote 5 notes that Claude's BenchCAD scores reflect three modifications to the evaluation. Footnote 3 notes the opposite choice on OSWorld, where OpenAI used official settings rather than the modified tasks and grading from Anthropic's own system card. Footnote 8 notes that Astra was run on FrontierCode with a developer message resembling its Codex prompt, adding that the prompt was not optimised for the evaluation. On that benchmark it lands at 53.3, against 53.5 for Fable 5 and 53.4 for Opus 5. Statistically that is a tie, with a prompt advantage.
Footnote 12 is the most quietly significant: Claude Fable 5 and 5.1 are absent from LifeSciBench, GeneBench Pro and MedChemBench because they refuse the majority of the questions. That is not Astra winning those benchmarks. It is Astra being the only model that would answer, which is a product decision by two companies with different risk appetites, and reasonable people can disagree about which posture is correct.
Footnote 14 cuts the other way and deserves credit: OpenAI flags that GPT-5.6 Sol's 5.5 percent on their new exploit benchmark is an artefact of a 300 turn limit, and that the same model reached 11.5 percent with fewer limits. Disclosing that your own previous model was disadvantaged by your evaluation harness is not what a purely promotional document does.
None of this makes the launch dishonest. It makes it a launch. The method in the flow above is the general lesson: a headline number without its footnote is not a claim, it is a mood.
4. So which reading is right? Both, about different things
Here is the resolution we arrived at, and we think it is the actual story.
Astra is not a general intelligence jump. It is a specialisation jump. The axes where it moves hard are computer use, long horizon agentic work, token efficiency, scope adherence and cyber. The axes where it does not move are the ones an aggregate reasoning index measures.
An intelligence index is largely a test of what a model knows and can work out. It is a poor instrument for whether a model can drive KiCad for forty minutes without losing the plot, respect an auto review denial it could technically evade, or finish a task in half the wall clock time. Those are the properties that decide whether an agent survives contact with a business, and they are badly captured by the thing everyone quotes.
Which means the two readings are not in conflict. If you want a model to reason about hard problems in a chat window, the index is telling you something real and Astra is not the upgrade. If you want a model to operate software unattended for an hour without going out of scope, the index is nearly irrelevant and the launch post is describing something that matters a great deal.
This is the same argument we made from the opposite direction in the small language models note: capability and appropriateness are different axes, and the metric that settles a real decision is cost per completed task, not any leaderboard. On that metric Astra currently looks 75 percent more expensive than its predecessor, and whether the efficiency and reliability gains repay that is precisely what nobody has measured yet on real work.
5. The cyber capability is the actual headline
Buried under the benchmark tables is the most consequential thing in the announcement.
OpenAI states that Astra meets the Critical threshold for cybersecurity under its own Preparedness Framework. During an internal exploit benchmark built from vulnerabilities from the previous three months, the model discovered and used two previously unknown zero day vulnerabilities, which OpenAI says it is disclosing to the maintainers. Expert assessments found that without production safeguards it could achieve arbitrary code execution in hardened browsers and build privilege escalation exploits for hardened operating systems. On SRE-Bench, a contamination free reverse engineering benchmark, it solved 88.0 percent of tasks first try against 55.9 for Sol.
Put that next to what we wrote a few days ago in the agent security paper. The GTG-1002 campaign Anthropic disclosed in November 2025 had a state linked group driving a coding agent through roughly 30 intrusions, with the agent doing 80 to 90 percent of the tactical work. That was done with a model nobody rated Critical at cyber, by talking it into believing it was running authorised testing.
The deployment posture is correspondingly careful, and the flow above draws it. Astra refuses advanced offensive tasks such as building proof of concept exploits. OpenAI runs misalignment monitoring in production for models of this class, classifiers watching reasoning and actions that can automatically stop unauthorised activity. Access is off by default for enterprise workspaces at launch. Broader capability comes later through a programme called Daybreak.
Two honest consequences. First, defenders get a genuinely useful tool for secure code review and patching, which is the good version of this. Second, OpenAI says plainly that these checks can slow, pause or stop legitimate work, and that in the API a flagged task simply stops. If you are building on Astra, an interrupted task is now a state your system has to handle, not an edge case. That is a real engineering requirement arriving with the model.
6. The line almost nobody quoted
One paragraph in the announcement deserves more attention than it got.
Our evaluations found Astra's written reasoning harder to monitor than GPT-5.6 Sol's, based on tests that explicitly asked it to evade monitoring.
OpenAI attributes this to Astra having greater control over its written reasoning and solving problems in fewer written steps, notes it still struggles to conceal reasoning on complex tasks, and says it takes the decline seriously.
Read that again in the context of the section above. Capability up, including cyber capability rated Critical. Monitorability down. Those two moving in opposite directions in the same release is exactly the pattern the safety literature has worried about for years, and it is disclosed in a launch post rather than discovered later. Publishing it is genuinely creditable. It is also the single thing we would most want to see a trend line for, because one datapoint is an anomaly and two is a direction.
7. On the pace, and what version numbers mean now
Worth stepping back, because the release cadence has become its own phenomenon.
In roughly twelve months the frontier has produced GPT-5.4, 5.5, 5.6 Sol and now 6 Astra; Claude Opus 5, Fable 5 and Fable 5.1; Gemini 3.8 Flash; Meta's Muse Spark 1.3; and from the open weight side Kimi K2 through K3, GLM 5.2 and 5.3, DeepSeek V4 and Nemotron 3 Ultra. We have written up two of those in the last week alone, in the Kimi K3 note and the TimesFM 3 note.
And yet the honest experience of the last year, for most people doing real work, has been incremental. Things got faster, cheaper and more reliable. Few releases changed what was possible. Fable 5 is the one most practitioners would name as an exception, and even that is arguable.
Astra arrives with a major version number, which promises a generational break, and an independent index score identical to its predecessor. That tension is the thing to take away. Version numbers have decoupled from capability jumps, because a version number is a marketing decision and a capability jump is an empirical question. GPT-6 may well turn out to be a genuine step change in agentic work. The number on the box is not the evidence for that, and neither is a launch chart.
The only thing that settles it is your own workload. Which brings us to the part we cannot answer yet.
8. What we do not know, and would like to
Written as questions because they are questions, not as hedged claims.
Does the computer use jump survive real enterprise software? Benchmarks run on clean environments. Real work happens in a badly maintained CRM behind a VPN with a session that times out. That gap has broken every previous computer use claim.
Is the token efficiency real on messy work? Sixty five percent fewer output tokens on a benchmark is not the same as on your codebase. If it holds, it offsets a lot of the price rise. If it does not, you are paying 2.5 times more for a lateral move.
What does misalignment monitoring cost in practice? How often does a legitimate task get paused, and what does that do to an unattended overnight run? OpenAI is candid that interruptions happen. Nobody has published a rate.
Does the notes mechanism beat compaction? Astra can keep notes across context windows instead of repeatedly compressing history into a summary, with earlier context still searchable. That is a real idea and it maps closely to what we argued in Read write memory. Whether it holds up over a six hour session is an empirical question with an answer.
Does 0 percent scope violation survive contact with an adversary? Not going out of scope on OpenAI's evaluation is encouraging. The relevant test is a determined attacker, which is what the agent security paper is about, and that test happens in the wild.
Why is enterprise access off by default? That may be routine caution. It may say something about confidence in the rollout. It is the kind of detail worth watching over the next month.
We will run Astra against the same tasks we run everything against, measure cost per completed task rather than cheer at a chart, and write up what we find, including if it contradicts this. If you want that testing done against your own workload rather than against a benchmark suite, talk to us, or see how we build systems for businesses.
9. Sources
- OpenAI, GPT-6 Astra: a new generation of intelligence, the launch announcement and benchmark tables quoted throughout.
- GPT-6 Astra system card, on the Deployment Safety Hub.
- OpenAI, path to Astra: critical capabilities and frontier safeguards and responding to the next frontier of critical cyber capabilities.
- OpenAI, safety overview for GPT-6 Astra.
- Artificial Analysis, benchmarking GPT-6 Astra, the independent index results, token efficiency and cost per task figures.
- OpenAI API model reference for gpt-6-astra.
- Spence et al, a realistic, contamination free reverse engineering benchmark, arXiv 2608.11469, the SRE-Bench paper referenced in the announcement.
- ARC Prize Foundation, for the ARC-AGI benchmark family.
Everything above is a reading of published material. We have not run GPT-6 Astra on our own workloads, and vendor benchmark results, including comparisons against competitor models, are produced by a party with an interest in the outcome. That applies to every lab, ours included when we publish numbers. Where OpenAI and Artificial Analysis disagree we have given both rather than picking the one that made a better sentence.
This note sits in our Research track, alongside the foundation models, governance and operations threads. The full library is at Research.