Models·8 min read·Axios

GPT-6 Astra Lands and OpenAI Declares the AGI Era

OpenAI's largest training run ever produced a model that drives a computer at superhuman speed and is the first rated Critical on cyber. Independent scoring puts its intelligence level with the model it replaces — at 2.5x the price.

GPT-6 ASTRA OPENAI · "WELCOME TO THE AGI ERA" ARC-AGI-3 99.9 COMPUTER USE BITSMINDS.COM
Share:

OpenAI released GPT-6 Astra on Thursday and did not hedge about it. President Greg Brockman called the model a generational leap and told reporters "it's not unreasonable to feel that we are now in the AGI era, and I think that if you want to say this [model is] the first one, I think it's reasonable." His post announcing it was shorter: "Welcome to the AGI era!"

Astra came out of the largest training run OpenAI has ever done by a wide margin — more than 100,000 GPUs at the Stargate site in Texas, the company's first pretraining run above that threshold. It ships with a 1,050,000-token context window, a 128,000-token maximum output, a knowledge cutoff of 30 April 2026, and five reasoning-effort levels: low (the API default), medium, high, xhigh and max.

The rollout is staged and unusually cautious. Access went first to enterprises in OpenAI's Daybreak program for cybersecurity defenders; ChatGPT Plus, Pro, Business and Enterprise users, the API and AWS follow "in the coming days," with cybersecurity tasks restricted. That caution is not decoration — Astra is the model that crossed the Critical cybersecurity tier of OpenAI's Preparedness Framework earlier this week, the first ever to do so, and its release triggered internal safeguards no previous model has.

The capability OpenAI is actually selling is computer use. Brockman says the model can "zip through spreadsheets, fill out forms, and navigate across web pages often at superhuman speed," and the company's own framing is blunter still: "Astra can really do anything a human can do with a computer." Mia Glaese of OpenAI framed it as a threshold crossing, saying it shows how far the field has come from aspirationally training for computer use to delivering everyday value. That is the claim on which the AGI language rests.

OpenAI did not just publish its own scores; it published a full head-to-head against GPT-5.6 Sol, Claude Opus 5 and Claude Fable 5.1, and the table is more revealing than the press quotes.

Reasoning and knowledgeSelf-reported by OpenAI, not independently reproduced · higher is better · blank = not publishedGPT-6 AstraGPT-5.6 SolClaude Opus 5Claude Fable 5.102040608010097.683.073.287.8FrontierMath T496.094.693.793.7GPQA Diamond57.2n/r63.665.0Humanity's LastExam
FrontierMath Tier 4 is a rout. GPQA Diamond is a 2.3-point edge on a benchmark the whole field has nearly saturated. And on Humanity's Last Exam, Astra is beaten by both Claude models OpenAI chose to compare itself against. Data: OpenAI.

Some of it is decisive. FrontierMath Tier 4 at 97.6% against Fable 5.1's 87.8% and Opus 5's 73.2% is not a close call, and Terminal-Bench-Science at 64.6% against 52.6% and 30.0% is a rout. Terminal-Bench 4.0 at 57.7% leads a tight pack. But GPQA Diamond at 96.0% sits 2.3 points above GPT-5.6 Sol on a benchmark everyone has nearly saturated, and on Humanity's Last Exam — with tools — Astra scores 57.2% against 63.6% for Opus 5 and 65.0% for Fable 5.1. OpenAI published that row anyway, which is to its credit, and it is the row least compatible with the AGI framing.

Coding and agentsSelf-reported by OpenAI · higher is betterGPT-6 AstraGPT-5.6 SolClaude Opus 5Claude Fable 5.102040608010057.737.352.355.8Terminal-Bench4.059.353.655.548.7Agents' LastExam
Terminal-Bench 4.0 leads a tight pack — 1.9 points over Claude Fable 5.1 — while the jump over its own predecessor is 20 points. Agents' Last Exam is the wider win, and OpenAI reports reaching it with about 65% fewer output tokens than Claude Opus 5. Data: OpenAI.

One headline number deserves an asterisk. The marquee ARC-AGI-3 result of 99.9% depends on a stateful, expensive evaluation harness; through ordinary stateless API calls the score reportedly falls to somewhere between 17% and 63% depending on tier. That is the same methodological argument OpenAI and ARC Prize had in July over GPT-5.6 Sol, and it has not been settled — it has just been restaged at a higher number.

The gains that look most solid are the ones tied to operating a computer.

Computer use and cyber — where the jump is realSelf-reported by OpenAI · higher is better · ScreenSpot-Pro figure for Claude is Fable 5GPT-6 AstraGPT-5.6 SolClaude Opus 5Claude Fable 5.1020406080100100.078.570.0n/rExploitBench92.776.9n/r87.3ScreenSpot-Pro88.055.9n/rn/rSRE-Bench72.665.770.2n/rOSWorld 2.0
The rows that carry the AGI framing. A perfect ExploitBench score and a 32-point SRE-Bench lead over its predecessor are the reason this model shipped behind restrictions, and the computer-use gains come with a speed claim too: OSWorld tasks in about 40 minutes against 75 for GPT-5.6 Sol. Data: OpenAI.

ExploitBench at a perfect 100% against 78.5% for GPT-5.6 Sol and 70.0% for Opus 5, SRE-Bench at 88.0% against 55.9%, ScreenSpot-Pro at 92.7% against 76.9%, and OSWorld 2.0 at 72.6% in roughly 40 minutes per task where its predecessor needed about 75. OpenAI also reports Astra finishing Agents' Last Exam with about 65% fewer output tokens than Opus 5. Worth noting on the cyber row: under a contamination-controlled ExploitBench restricted to vulnerabilities from June to August 2026, Astra scores 39.0% rather than 100% — still far ahead of Sol's 5.5%, but a useful measure of how much of the perfect score is memorisation.

Then there is the independent read, and it complicates the announcement considerably.

The independent scorecardArtificial Analysis, GPT-6 Astra at max effort · higher is better · not every model is scored on both indicesGPT-6 AstraGPT-5.6 SolClaude Opus 5Claude Fable 5.10204060806161n/r66IntelligenceIndex67n/r6770Coding AgentIndex
Artificial Analysis puts Astra's Intelligence Index at 61 — level with GPT-5.6 Sol, the model it replaces, and behind Claude Fable 5.1 at 66. On the Coding Agent Index it ties Claude Opus 5 at 67. Data: Artificial Analysis.

Artificial Analysis scores Astra at max effort at Intelligence Index 61. That is the same score it gives GPT-5.6 Sol, and below Claude Fable 5.1 at 66. On its Coding Agent Index Astra reaches 67, matching Claude Opus 5 and Fable 5 rather than clearing them. More awkwardly, the firm records a roughly 80-Elo regression on GDPval-AA v2, the benchmark that scores models on economically valuable professional work — the exact territory where an AGI claim would want to be strongest.

Astra's real efficiency story is tokens, not scores. Artificial Analysis measures it using about a third of the tokens GPT-5.6 Sol needs at max, and a fifth of what Claude Opus 5 uses at xhigh — roughly 70% more token-efficient than its predecessor. That is why it lands the same coding score as Claude Fable 5 at less than half the cost, even though the list price went up sharply: $10 per million input tokens and $50 per million output, 2.5x GPT-5.6 Sol's $4/$20, with cached input at $1 and cache writes at $12.50. Prompts over 272,000 tokens are billed at 2x input and 1.5x output.

The safety picture is where the discomfort concentrates. OpenAI says it has deployed heightened cybersecurity protocols and monitoring designed to "rapidly detect and contain potentially misaligned actions," and chief scientist Jakub Pachocki was candid about the direction of travel: "As these models become more capable, understanding exactly what they can do gets harder." He added a line that reads as a commitment and a warning at once — "we will not accept degradation in our ability to monitor model alignment beyond a certain level." Outside researchers were unconvinced that monitoring is sufficient for a model at the Critical threshold, and The Information reported that a training technique used on Astra could meaningfully reduce human insight into how the system processes instructions, which drew immediate criticism.

So the honest summary is narrower than "AGI is here" and more interesting than "nothing happened." On aggregate intelligence Astra is level with the model it replaces and behind Anthropic's current flagship, at two and a half times the price. What it appears to have genuinely moved is autonomous computer use and offensive cyber capability — the two things that make a model useful as an agent, and the two things that make it dangerous as one. Those are the same capability. That is the part worth sitting with.

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.

Related Articles