GPT-6 Astra shipped in the first week of September 2026 and the reaction split immediately into two camps that were not looking at the same evidence. One camp read a 99% score on a reasoning benchmark and called it the end of the argument. The other read the footnotes.
This piece is the second reading. Not because scepticism is a personality, but because a production studio has to make purchasing and pipeline decisions from these numbers, and the gap between the two readings is the difference between rebuilding a workflow and leaving it alone.
What it is, before what it scores
Astra is a text and image input model that returns text. It does not generate video. It does not generate images. For a video production business this is the first fact that matters, because a large share of the coverage implied a general leap forward in creative capability, and the model's own interface rules that out.
The concrete parameters: roughly 1,050,000 tokens of context, $10 per million input tokens and $50 per million output, and a reasoning_effort control that now extends past high into xhigh and max. That last control is not a detail. It is where the benchmark story starts.
The 99% came from the harness
ARC Prize evaluated Astra on ARC-AGI-3 under two different test harnesses and published both. Under the standard harness at maximum reasoning, Astra scored 62.71%, at a cost of $26,098. Under a new Provider Adapter harness, it scored 99.95% for $18,817. Like for like at max reasoning, the comparison is 62.7% at $26,098 against 98.55% at $17,332.
Two numbers, same model, same week, differing by more than thirty-six points. The variable is the scaffolding around the model rather than the model itself. Anyone quoting the 99% without naming the harness is quoting an engineering result as if it were a cognitive one.
The 99% came from the harness
GPT-6 Astra on ARC-AGI-3, score and cost by test harness
| Harness | Score | Cost |
|---|---|---|
| Standard harness, maximum reasoning | 62.71% | $26,098 |
| Provider Adapter harness | 99.95% | $18,817 |
For scale on the cost side: humans solve these environments at 100% for roughly $12.78 per game.
ARC Prize's own summary contains the sentence most of the coverage left out: we are not claiming that it is AGI.
The claim that got distorted
The line that travelled furthest was that Astra beats humans on 96% of levels. ARC's actual statement is about action efficiency, that the model surpassed human performance on that share of levels in terms of taking fewer actions. Fewer actions is not higher accuracy. The sentence was compressed into something it does not say, and then the compressed version became the fact everybody repeated.
This is worth naming as a general pattern rather than a single error. Benchmark statements are conditional, scoped and usually hedged in the original. Every hop they make through social and secondary coverage strips a condition. By the fourth hop you have a headline the original authors would not sign.
The independent read is far less dramatic
Artificial Analysis, which runs its own harness rather than reprinting vendor numbers, places Astra at 53 on its Intelligence Index. That is tied, not ahead. It places third on Humanity's Last Exam at 54.7%.
Two results in that evaluation matter more than the aggregate for anyone doing commercial creative work. Astra regressed on GDPval-AA v2, the measure built around real professional deliverables, by roughly 45 Elo. And on AA-Briefcase, which scores presentation quality of finished output, GPT-5.6 Sol still leads.
So on the two axes closest to producing client-ready work, the newer, more expensive model is level at best and behind at worst. That is an uncomfortable finding and it is the one with the most direct bearing on a production budget.
The measurement environment itself got less trustworthy
Fortune documented OpenAI revising published benchmark numbers after launch, including a competitor's score on a third-party evaluation. Not a correction issued and flagged, a quiet edit. Stanford researchers have started using the word benchmaxxing for the broader behaviour of optimising against the test rather than the capability the test proxies.
The practical consequence for a buyer is that launch-week numbers now carry an implicit asterisk, and the rational response is a delay. Give any new frontier model two to four weeks before you act on its scores, and weight independent harnesses over vendor charts.
One thing that is unambiguously new
Astra is the first model rated Critical for cyber capability under OpenAI's own preparedness framework. That is a vendor's own risk classification, not a marketing claim, and it is the single least disputed statement in the whole launch. If you run agentic tooling with shell access, that rating is more relevant to your week than any reasoning score.
What this means if you produce video
Nothing in the model's capabilities touches generation. What it touches is the layer above generation: scripting, shot list construction, continuity tracking across a long project, prompt engineering at scale, brief interpretation. That work is text work, and a million-token context does change what is possible there. You can hold an entire production bible, a character sheet set and a full shot list in one context and ask questions across all of it.
Whether that is worth $50 per million output tokens depends entirely on whether the reasoning uplift shows up in your specific task, and the GDPval regression is a genuine warning that it may not. The honest position is that this needs testing on your own work, not inference from a leaderboard.
How to test it in a week
- Pick three tasks you already do. Not benchmark tasks. A script pass, a shot list build, a continuity audit across an existing project.
- Run each on your current model and on Astra at high, not max. Max reasoning is where the cost curve goes vertical and the ARC results show the return on it is harness-dependent.
- Score blind. Strip the labels before you or a colleague judge the outputs. Model preference is heavily contaminated by knowing which is which.
- Log tokens and dollars per task, not per million. Per-million pricing is meaningless until you know your own consumption shape.
- Decide per task, not per model. The likely outcome is that it wins on long-context work and loses on short generative work, in which case you route rather than switch.
The short version
Astra is a real step forward in long-context reasoning and a genuinely strong model. It is also the subject of a benchmark claim that does not survive reading the source, and it regressed on the measure closest to professional deliverable quality. Both statements are true at once and the second one is the one that should shape your spending.
The transferable lesson is bigger than this model. When a number is quoted without its harness, its date and its conditions, it is not yet evidence. Ask for those three things before you let a benchmark move your budget.
Sources
- ARC Prize, OpenAI's GPT-6 Astra on ARC-AGI-3
- ARC Prize, GPT-6 Astra results page
- Artificial Analysis, Benchmarking GPT-6 Astra
- Artificial Analysis, GPT-6 Astra (max) model page
- Fortune, OpenAI quietly boosts some of Astra's evaluation metrics (4 September 2026)
- The New Stack, Astra's 98.6% and the fine print
- TNW on the Astra harness discrepancy
- Stanford HAI, 2026 AI Index, technical performance