Products·6 min read
By BitsMindsSource: OpenAI

GPT-6.1 Sol: Near-Astra Scores for a Fifth of the Price

OpenAI released GPT-6.1 Sol at DevDay on 29 September, a week after GPT-6 Sol, at the same $2/$10 per million tokens with cached input halved to $0.10. On OpenAI's own charts it beats Astra on coding for a seventh of the cost per task, but still trails Astra and Claude Opus 5.5 on business automation and science.

GPT-6.1 SOLBITSMINDS.COM
Share:

GPT-6.1 Sol is OpenAI’s new mid-priced model, released on 29 September 2026 at DevDay in San Francisco, exactly one week after GPT-6 Sol. OpenAI’s announcement pitches it as a model that “nearly matches GPT‑6 Astra’s intelligence on agentic coding, computer use, and professional work at one-fifth of Astra’s standard input and output token prices.” In the API it is gpt-6.1-sol. It arrived while the top of the range stood still: a day earlier the company said it had scrapped GPT-6.1 Astra after the model failed its safety standards.

Price: the same sticker, cheaper caching

The headline rates have not moved from GPT-6 Sol. On OpenAI’s pricing page, GPT-6.1 Sol costs $2 per million input tokens and $10 per million output tokens, with cache writes at $2.50, for prompts up to 272K tokens. Longer prompts cost $4 in and $15 out. The change is cached input, which drops to $0.10 per million, half of GPT-6 Sol’s $0.20 and 95% below the standard input rate. That matters most for agents, which resend the same long context on every turn. GPT-6 Astra stays at $10 and $50, so the “fifth of the price” is a straight per-token comparison.

GPT-6.1 Sol is available from today in ChatGPT Work and Codex on the Plus, Pro, Business, Enterprise and Edu plans, and in the API. It is not yet in the ordinary ChatGPT chat window. OpenAI says a GPT-6.1 Sol version of Ultrafast, its new premium speed tier, is coming “in the coming days” with up to eight times faster token generation in Codex. Ultrafast launched today for GPT-6 Astra only, at $60 per million input tokens and $300 per million output tokens in the API, six times Astra’s standard rate, and in ChatGPT on the new $500-a-month Pro 500 plan. The Codex documentation says Ultrafast draws down a Pro 500 allowance at eight times the standard rate.

What OpenAI’s charts show

OpenAI published scores for each model at five reasoning-effort settings, from low to max, together with the cost per task at each one. That is more useful than a single headline number, and it also shows where the announcement picks its comparisons. The chart below takes each model’s best score at any setting.

GPT-6.1 Sol against its family and Opus 5.5Best score at any tested effort, % · higher is better · OpenAI, 29 Sept 2026GPT-6.1 SolGPT-6 SolGPT-6 AstraClaude Opus 5.5 (with fallbacks)02040608075.268.874.1n/rDeepSWE 1.132.028.032.228.8GDP.pdf36.133.241.442.5AutomationBench71.464.473.5n/rOSWorld 2.0offline57.027.668.163.3Terminal-BenchScience
Best score at any of the five effort settings OpenAI tested. Opus 5.5 was run with fallbacks and was not tested on DeepSWE or OSWorld. Data: OpenAI.

Coding is the clearest win. On DeepSWE 1.1, a test of long software-engineering tasks in real codebases, GPT-6.1 Sol reaches 75.2% at high effort, ahead of Astra’s best of 74.1% and 6.4 points above GPT-6 Sol’s best. It gets there for $0.65 a task, against $4.43 for Astra. Oddly, its score falls to 71.9% at the two higher settings, xhigh and max, so more thinking does not help it here. On computer use, OSWorld 2.0, GPT-6.1 Sol scores 71.4% at max effort, 2.1 points short of Astra, for $1.27 a task against $9.44. On GDP.pdf, which asks professional questions about complex PDF documents, it is effectively level with Astra and ahead of Claude Opus 5.5 at every setting.

The comparison with Anthropic’s model is where the wording needs care. OpenAI says that on AutomationBench, which tests multi-step business workflows, GPT-6.1 Sol scores “2.2 percentage points above Opus 5.5 at medium reasoning effort”. That is accurate: 31.7% against 29.5%. But the same chart shows Opus 5.5 at max effort reaching 42.5%, the highest score on the chart and above Astra’s 41.4%, while GPT-6.1 Sol peaks at 36.1%. On Terminal-Bench Science, which tests research work such as data analysis and simulations, GPT-6.1 Sol more than doubles GPT-6 Sol’s score to 57.0%, but Astra (68.1%) and Opus 5.5 (63.3%) are still clearly ahead. Both of them cost more than four times as much per task, so the choice there comes down to budget.

The cost column is what separates GPT-6.1 Sol from both rivals. Here is what each model spent per task at the setting where it scored best:

BenchmarkGPT-6.1 SolGPT-6 SolGPT-6 AstraOpus 5.5
DeepSWE 1.1$0.65 (high)$2.74 (max)$4.43 (xhigh)n/r
GDP.pdf$0.35 (high)$0.35 (high)$1.91 (xhigh)$0.83 (high)
AutomationBench$0.30 (max)$0.27 (xhigh)$1.73 (max)$1.44 (max)
OSWorld 2.0 offline$1.27 (max)$3.37 (max)$9.44 (max)n/r
Terminal-Bench Science$5.47 (max)$12.18 (max)$23.80 (max)$23.21 (max)

The table also shows that a lower per-token price does not always mean a cheaper task. GPT-6.1 Sol costs the same per token as GPT-6 Sol, yet it is cheaper per task on DeepSWE, OSWorld and Terminal-Bench Science, largely because it needs fewer tokens to reach its best result. All of these figures are OpenAI’s own, from its research environment. OpenAI took competitors’ results from public reports, and no independent benchmark had published GPT-6.1 Sol results by the time of writing.

Safety, a day after the Astra decision

Given the timing, OpenAI spent more of the announcement on alignment than it usually does. It says GPT-6.1 Sol is “more transparent about its limitations and more reliable at respecting user intent and safety constraints,” and its stress tests support that against GPT-6 Sol. In a test designed to provoke the behaviour, GPT-6 Sol found ways around a warning in 64.4% of cases, and GPT-6.1 Sol does so in 23.5%. On a computer-use safety stress test the failure rate falls from 17.4% to 4.3%. Neither model tried to bypass the automated safety reviewer. The new model still fails more often than Astra on all three tests charted below, and the tasks are chosen to produce failures, so the rates say nothing about everyday use.

OpenAI’s stress tests: GPT-6.1 Sol closes most of the gap to Astra% of adversarial cases failed · lower is better · max effort, computer use at xhigh · OpenAIGPT-6.1 SolGPT-6 SolGPT-6 AstraGPT-6 Luna0102030405060702.14.91.528.7Hides a brokensearch tool23.564.417.442.4Works around awarning4.317.42.413.7Computer-usesafety
Adversarial tasks chosen to elicit failures; they do not measure typical use. Data: OpenAI.

OpenAI also reports fewer factual errors on hard prompts. At low effort, the share of answers with at least one error falls from 11.4% to 7.7%. The test uses de-identified ChatGPT conversations in which users had flagged a mistake by an earlier model, so it is deliberately difficult.

Where it fits in the day

GPT-6.1 Sol was one of more than 20 announcements in OpenAI’s DevDay recap, which led with dots, always-on agents that run on GPT-6 Astra rather than on Sol. Before the keynote we listed the rumours and called a new model the weakest bet, because GPT-6 Sol had shipped only a week earlier. That bet was wrong. Our GPT-6 Sol build-off found only a small step up from GPT-5.6 Sol. On OpenAI’s numbers, 6.1 is a much larger jump for the same money, and the next thing to watch is whether independent benchmarks agree.

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.

Related Articles