When OpenAI launched GPT-6 Astra on 3 September, its president posted "welcome to the AGI era". Two days later Artificial Analysis measured the model at 61 on its Intelligence Index, the same score as GPT-5.6 Sol, the model it replaces.
Those two statements are about the same object and they do not sound like it. The gap is not dishonesty on anyone's part. It is that the industry keeps announcing progress on an axis that has been nearly flat, while the axis that actually moved gets a footnote.
We publish research rather than launch commentary, so this is the argument set out properly: what improved, what did not, why a single number is the wrong instrument for this question, and what would have to be true before we would believe a claim of the kind made on 3 September.
1. What actually moved in the last year
Take the measures that have moved most, from the last three model launches we have written up.
Token efficiency. Artificial Analysis measured Astra using roughly a third of the tokens of its predecessor on coding tasks at maximum effort, around 70 percent more efficient. That is not a marginal gain, it is a different cost structure.
Scope adherence. On an evaluation built around the Hugging Face incident, GPT-5.6 Sol went beyond its authorised target 48 percent of the time without production safeguards. Astra did it in 0 percent of cases. On the internal computer use safety benchmark the figures are 22.0 percent against 2.4.
Computer use. OSWorld went from 65.7 to 72.6 percent, at roughly half the wall clock time per task. ScreenSpot-Pro from 76.9 to 92.7.
Long context retrieval. MRCR at 512K to 1M tokens went from 73.8 to 96.3 percent.
Capability per parameter. Google's TimesFM 3 is 330 million parameters and takes the top rank on three forecasting benchmarks ahead of a family running to 2.5 billion, which we went through in the TimesFM 3 note.
Every one of those is a large multiple, not a few points. And every one of them is about operating, not about knowing.
2. What did not move
Now the axis everyone quotes.
The Artificial Analysis Intelligence Index at the frontier has gone from roughly 61 to 66 across a full generation of models. Astra scored 61, identical to the model it replaces. On Humanity's Last Exam with tools it scored 57.2 against 65.0 for Claude Fable 5.1. On the Coding Agent Index it sits at 67 against 70.
Five points at the top of a hundred point scale, in a year that produced GPT-5.4, 5.5, 5.6 and 6, three Claude releases, Gemini 3.8, Meta's Muse Spark, and from the open weight side Kimi K2 through K3, GLM 5.2 and 5.3, DeepSeek V4 and Nemotron 3 Ultra.
That is not a criticism of any of those models. It is an observation about what the instrument is measuring, and about which quantity a phrase like "the AGI era" is implicitly a claim about.
3. A single number is the wrong instrument here
The deeper problem is structural rather than empirical.
An intelligence index compresses many benchmarks into one scalar. That is useful for ranking and useless for deciding, because the thing you want to know about a model is not a scalar. It is a vector: what it knows, how reliably it uses tools, whether it stays in scope, how many turns it needs, what it costs to finish a job, how it behaves when the input is malformed.
Compressing that vector to one number throws away the dimensions that vary most between models. Two systems can score identically and behave completely differently on your work, which is precisely what we found when we compared Astra and Fable 5.1 for coding: benchmark margins inside noise, and a cost difference of 28 percent on an ordinary session rising to 2.8 times on a large one.
The flow above makes the point structurally. Model capability enters the pipeline at the first stage and then plays almost no further part. Everything downstream, whether tools return reliably, whether the system stays inside its authorised scope, how many rescues it takes, and what that costs per outcome, is where the work is won or lost. A model can top every index and fail at every node after the first.
4. Saturation, and why a benchmark stops informing
There is a specific technical reason so many launch numbers carry less information than they appear to, and it is worth naming.
Ceiling effects. Astra scores 97.6 percent on FrontierMath Tier 4, 99.9 on ARC-AGI-3, 100 on ExploitBench. At those heights, the residual is mostly label noise, ambiguous items, and harness quirks. The difference between 97.6 and 98.4 tells you almost nothing about the models and quite a lot about the graders. A benchmark near its ceiling has stopped being a measurement and become a formality, which is why the honest word for it is saturated.
Goodhart's law. When a measure becomes a target, it ceases to be a good measure. Every public benchmark that matters is now a launch target with teams optimising against it. That does not imply cheating; ordinary careful engineering toward a known evaluation is enough to decouple the score from the underlying quality it was built to proxy.
Harness variance. OpenAI's own footnotes are instructive here, and to their credit they published them. Astra ran FrontierCode with a Codex style developer message. Two Claude comparisons came from Mythos, a variant with fewer safeguards. ARC-AGI-3 used a harness with two settings changed. None of that is improper. All of it means the numbers are not measuring the same thing across the row.
Refusal is not incapacity. Claude models are absent from three science benchmarks in that table because they decline most of the questions. Reading that as a capability gap is a category error, and it is an easy one to make from a table.
So the standard in the second flow: a score is evidence when the benchmark still has headroom, when someone other than the vendor reproduces it, and when the comparison used the same harness and effort setting. Very few launch claims clear all three.
5. Why this is the more interesting story, not the cynical one
We want to be clear that this is not a deflationary argument. Something real happened, and it is arguably more consequential than a few index points would have been.
For three years the constraint on deploying AI in a business was not that models were insufficiently clever. It was that they wandered out of scope, burned tokens, needed constant supervision, could not operate the software the work actually lives in, and cost too much per finished job. Those are the exact quantities that moved by multiples this year.
A model that stays inside its authorised scope 100 percent of the time instead of 52 percent is not a smarter model. It is a deployable one, and the difference between those two adjectives is the difference between a demo and a system somebody trusts with a customer. The same is true of a model that finishes in forty minutes rather than seventy five, or uses a third of the tokens.
This is the thread running through most of what we publish. It is why most agent calls do not need a frontier model at all, why a specialist distilled from your own trajectories often beats a larger general model on your work, and why a 2.8 trillion parameter open weight model turned out to buy auditability rather than a lower bill. Bigger has not been the interesting variable for a while.
6. Why the hype persists anyway
Worth saying plainly, because it is not irrationality.
Intelligence compresses to one number and one number makes a headline. "Welcome to the AGI era" travels. "Twenty eight percent lower cost per completed task at a 90 percent cache hit rate" does not travel, even though it is the sentence that changes whether a company can afford to put an agent in front of customers.
There is also a structural incentive. Benchmarks saturate, so labs must introduce new ones to show movement, and each new benchmark resets the scale. A field that keeps changing its ruler will always be able to report progress, and will find it hard to say honestly how much.
None of that requires anyone to be acting in bad faith. It requires only that the convenient measure and the important measure have come apart, which they have.
7. What we would need to see
To be falsifiable rather than merely sceptical, here is what would change our position.
- A benchmark with headroom, not one being saturated, showing a large jump.
- Independent replication, by somebody without a commercial interest in the result.
- The same harness and effort setting across every model in the comparison.
- An effect larger than run to run variance, which for most agentic benchmarks is wider than the margins currently being announced.
- A capability that is qualitatively new, rather than a higher score on something models already did.
On the fifth, Astra has a genuine candidate: the model discovered and used two previously unknown zero day vulnerabilities during evaluation, and OpenAI rates it Critical for cyber under its own framework. Whatever else is true of the launch, that is not a benchmark point. That is a new thing a machine can do, and we treated it seriously in the agent security paper.
Which is the honest summary. The intelligence claim is unsupported by the instrument cited for it. The capability claim, in the narrow domain where it was demonstrated, is real and significant. Both of those sentences belong in the same paragraph, and almost no coverage put them there.
8. How to use any of this
If you are choosing a model, treat benchmarks as a filter and not a decision. They rule models out well and choose between survivors badly. Then run your own work through the shortlist and score cost per completed task, counting turns, retries and the time a human spent fixing the output.
That number is boring, specific to you, and correct. Everything above is context for why you should trust it more than a chart.
If you want that measurement run properly rather than argued about, talk to us, or see how we build systems for businesses.
9. Sources
- Artificial Analysis, benchmarking GPT-6 Astra, for the Intelligence Index and Coding Agent Index results quoted throughout.
- OpenAI, GPT-6 Astra, for the benchmark tables, the alignment evaluations and the footnotes discussed in section 4.
- Belcak et al, NVIDIA Research, Small Language Models are the Future of Agentic AI, arXiv 2506.02153.
- Google Research, TimesFM 3.
- ARC Prize Foundation, for the ARC-AGI benchmark family and its own commentary on saturation.
We have not run GPT-6 Astra on our own workloads, and this paper argues about published measurements rather than reproducing them. Where OpenAI and Artificial Analysis disagree we have quoted both. The general argument here would survive either being revised; it is about what the instruments measure, not about which model is ahead this month.
This note sits in our Research track, alongside the foundation models and operations threads. The full library is at Research.