ChatGPT
ChatGPT
Text & Chatfreemium
4.5

GPT-6 Astra Review: Frontier Intelligence, Half the Cost

Astra launched to a brutal independent read: intelligence level with the model it replaces, at 2.5x the price. That framing measured cost per token. Measured per finished task — the number you actually pay — Astra completes the Artificial Analysis Intelligence Index for $4,064 where Claude Fable 5.1 needs $10,816, and scores two points behind it. Our own testing found the same speed advantage, and one repeated weakness no benchmark catches.

Pros

  • Finishes the Artificial Analysis Intelligence Index for 37% of what Claude Fable 5.1 spends
  • Uses 49M output tokens where Fable 5.1 needs 160M and Opus 5 needs 120M
  • Third overall on Intelligence Index v4.2 at 55 — a point above Claude Opus 5
  • Cheapest route to a top-three intelligence score available today
  • 25-45% faster than the Anthropic flagship on wall-clock in our own build tests
  • OSWorld 2.0 tasks in roughly 40 minutes against about 75 for GPT-5.6 Sol
  • Category lead on computer use: ScreenSpot-Pro 92.7% and SRE-Bench 88.0%
  • 1.05M-token context window with 128K maximum output

Cons

  • Not the smartest available — Claude Fable 5.1 scores 57 to Astra’s 55
  • Costs more per finished task than the GPT-5.6 Sol it replaces
  • 322 seconds to first token at max effort — slowest of the frontier group
  • Prompts above 272K tokens bill at 2x input and 1.5x output
  • Our own testing found polished surfaces over repeated mechanism failures
  • Muse Spark 1.3 is a third of the price and 3.7x faster if a score of 53 suffices
  • A Critical cyber rating means staged access and restricted capabilities
  • The 99.9% ARC-AGI-3 headline depends on an expensive stateful harness

Three days after GPT-6 Astra launched, the independent verdict looked brutal: intelligence level with the model it replaces, at two and a half times the price. We wrote that ourselves. It was the correct reading of the numbers available at the time, and it measured the wrong thing.

Per-token list price is not what you pay. What you pay is that price multiplied by the number of tokens the model burns reaching an answer — and on that number Astra is not expensive at all. It is the cheapest way to buy a top-three intelligence score that currently exists.

This review combines our own hands-on testing with the published benchmark record. Where the two disagree, we say so. On one point they disagree sharply.

The number that changes the verdict

Artificial Analysis runs an identical Intelligence Index across every frontier model and publishes what each run cost to complete. That figure is the closest thing the industry has to a like-for-like price on finished work, and it is the only pricing number in this review that really matters.

ModelIntelligenceOutput tokensCost to finishList price
Claude Fable 5.157160M$10,815.70$10 / $50
GPT-6 Astra5549M$4,063.77$10 / $50
Claude Opus 554120M$5,545.57$5 / $25
Meta Muse Spark 1.353130M$1,340.75$1.25 / $4.25
GPT-5.6 Sol5176M$2,463.56$4 / $20
What each model spent to finish the same benchmarkUSD to complete Intelligence Index v4.2 at max effort, ordered by score · lower is betterCost to complete the full index run02000400060008000100001200010816Claude Fable 5.14064GPT-6 Astra5546Claude Opus 51341Muse Spark 1.32464GPT-5.6 Sol
Scores run 57, 55, 54, 53, 51 from left to right. The bills do not. Astra and Claude Fable 5.1 carry the identical $10/$50 list price, and Astra still finishes for 37% of the money — because it needs 49M output tokens where Fable 5.1 burns 160M. Data: Artificial Analysis.

Astra and Claude Fable 5.1 charge exactly the same list price — $10 per million input tokens, $50 per million output. Fable 5.1 scores two points higher. It also spends 2.7 times as much money getting there, because it emits 160 million output tokens where Astra emits 49 million. Against Claude Opus 5 the comparison is starker still: Astra scores a point higher and spends 27% less, on a list price exactly double Opus 5’s.

One caveat before anyone holds these scores against the launch-week coverage. Artificial Analysis has moved to Intelligence Index v4.2, and the figures across the entire field are lower than the ones that circulated in early September. Compare models within the table above, not against last week’s write-ups.

Where the value case does not hold

Two places — and both matter enough that anyone quoting the paragraph above should quote these as well.

Against its own predecessor, Astra costs more. GPT-5.6 Sol finishes the same index for $2,464; Astra needs $4,064 to add four points. Artificial Analysis’s own summary of the model is that its token savings are “outweighed by higher prices,” and measured against Sol that is simply correct. The value case here is specifically against the two models that outscore Astra — not against the field.

It is nowhere near the cheapest. Meta’s Muse Spark 1.3 scores 53 to Astra’s 55, finishes for $1,341 — a third of the money — and runs at 230 output tokens per second against Astra’s 62.5. If a score of 53 does your job, Astra is the wrong purchase by a wide margin. The per-token side of that comparison is in our pricing calculator.

There is a third irritant no index captures. At max effort Astra takes 322 seconds to produce its first token — over five minutes of nothing, and the slowest first-token time in the frontier group; Fable 5.1 waits 277 seconds, Opus 5 only 63. Astra is quick to finish and slow to start, which is an awkward shape for interactive work. Prompts above 272,000 tokens also bill at 2x input and 1.5x output, so the long-context bargain is narrower than the 1,050,000-token window suggests.

What we found running it ourselves

Alongside the published record we put Astra through a set of from-scratch build tasks against the current Anthropic flagship: identical briefs, one attempt each, maximum reasoning effort, and no ability to open a browser and check its own work. Two findings came out, and they point in opposite directions.

The speed advantage is real and it is large. Astra finished every task between 25% and 45% faster in wall-clock time. That is the token-efficiency number showing up as a stopwatch reading rather than an invoice, and it is the most underrated thing about this model. A five-minute wait for the first token still nets out ahead when the whole job lands a third sooner.

The weakness is consistent enough to plan around. In every task where something had to move, change state, or run through a sequence of stages, Astra produced work that looked right and did not behave right — while its surface detail was the richest of anything we tested. The decoration never failed. The mechanisms did, repeatedly. A reviewer glancing at a screenshot would have approved all of it.

That is a specific and manageable failure mode rather than a disqualifying one. The practical translation: Astra is an outstanding first-draft engine that needs a reviewer checking behaviour rather than appearance. If your verification step is a human looking at output, this model will slip things past you. If your verification step is a test that runs, it will not.

What OpenAI is actually selling

The AGI language around this launch is marketing. The computer-use capability underneath it is not.

Computer use and cyberSelf-reported by OpenAI · higher is better · versus the model Astra replacesGPT-6 AstraGPT-5.6 Sol020406080100100.078.5ExploitBench92.776.9ScreenSpot-Pro88.055.9SRE-Bench72.665.7OSWorld 2.0
The rows the AGI language rests on. Worth an asterisk: on a contamination-controlled ExploitBench limited to vulnerabilities disclosed between June and August 2026, the perfect score falls to 39.0% — still far ahead of Sol at 5.5%. Data: OpenAI.

OSWorld 2.0 is the row to sit with: Astra completes those tasks in roughly 40 minutes where GPT-5.6 Sol needed about 75. That is the same efficiency story again, measured in a different unit. It is also why Astra became the first model OpenAI has rated Critical for cyber capability, why access is staged, and why cyber tasks are restricted. The capability that makes it valuable as an agent is the capability that makes it dangerous as one — they are the same capability.

Two asterisks belong on the headline numbers. The marquee ARC-AGI-3 result of 99.9% depends on a stateful and expensive evaluation harness; through ordinary stateless API calls it reportedly falls to somewhere between 17% and 63%. And on a contamination-controlled ExploitBench restricted to vulnerabilities disclosed between June and August 2026, the perfect 100% becomes 39.0% — still far ahead of Sol’s 5.5%, but a useful measure of how much of that score is memorisation.

Who should be running it

If you are…Do this
Running Claude Fable 5.1 or Opus 5 in productionRe-run your evals against Astra. The same tier of intelligence for 37–73% of the bill is the largest cost lever on the table right now
Still on GPT-5.6 SolUpgrade only if you need the four extra points. Sol finishes the same benchmark for 60% of the money
Doing computer use or browser automationThis is the category leader, and it is not close
Shipping against a deadlineThe 25–45% wall-clock advantage we measured compounds across a working day
Building anything with moving partsUse it — but assert behaviour in tests rather than reviewing output by eye
Happy with a score of 53Muse Spark 1.3 is a third of the price and 3.7x faster. Astra is the wrong purchase
Serving interactive usersBudget for a first token that can take five minutes at max effort, or drop the effort level

Verdict

The claim doing the rounds — that Astra is the best model available today once you weigh performance, cost and speed — is not quite right, and it is a great deal closer to right than the launch-week coverage suggested. It is not the best on performance; Fable 5.1 is. It is not the cheapest; Muse Spark 1.3 is, by a distance. What it is, precisely, is the cheapest way to buy intelligence at the very top of the table and the fastest model up there at finishing work.

Against the only two models that outscore it, Astra costs 37% and 73% of what they do to complete the same benchmark. That is not a rounding error, and it is a larger effect on a real budget than the two points of intelligence separating it from first place. For most production workloads it is the trade you want.

What keeps it off the top of our scale is the thing benchmarks cannot see and our own testing kept finding: beautiful surfaces over mechanisms that do not work. Set that beside a five-minute wait for the first token and a price that only looks reasonable next to Anthropic’s, and this becomes a model to deploy deliberately rather than by default — behind tests, on work where finishing quickly matters, and never as a casual drop-in for a cheaper model that was already good enough.

Related Reviews

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.