GPT-6 Sol Review: Half the Cost per Task, Barely Smarter

GPT-6 Sol halves GPT-5.6 Sol’s per-token price, and on the independent Artificial Analysis index it nearly halves the bill per task too: $1.06 against $1.99, for a score of 48 against 47. No model scoring 48 or more finishes a task for less. It is barely smarter than the model it replaces, and Claude Opus 5.5 at high scores 54 for $1.82. In our own build tests Sol was fast, and each of its three builds had a fault you could see.

By BitsMindsHands-on reviewPublished Not re-verified since

The short version

Get ChatGPT if

  • Teams on GPT-5.6 Sol: a point more on the index for about half the cost per task, at rates OpenAI says are permanent
  • Agent builders who need a model to admit that a tool is broken, where OpenAI’s alignment figures moved the most
  • High-volume API work where a score in the high 40s is enough: no model at 48 or above finishes a task for less

Skip it if

  • Work that needs the top of the index: Claude Opus 5.5 at high scores six points more for $1.82 a task
  • Agents trusted to halt on a warning: in OpenAI’s adversarial test Sol still works around one in 64.4% of cases
  • Front-end and animation work shipped without a look at it: each of our three builds had a fault visible on screen

What we tested

  • Sol and Luna on three build briefs in OpenAI Codex at max effort on 22 September 2026, one attempt each, with no browser and no code execution
  • The Artificial Analysis Intelligence Index v4.3.2 table, captured on 22 September 2026
  • OpenAI’s launch post and the API model pages for GPT-6 Sol and GPT-6 Luna
Not re-verified since

Pros

  • Cheapest per task of any model scoring 48 or more on the Artificial Analysis index at $1.06
  • One index point above GPT-5.6 Sol for 53% of its cost per task
  • $2/$10 per million tokens — rates OpenAI says are permanent
  • Ahead of Claude Opus 5 on AutomationBench and OSWorld 2.0 at a fifth of the cost or less on OpenAI’s figures
  • Hid a broken tool in 4.9% of OpenAI’s adversarial cases where GPT-5.6 Sol did in 77.5%
  • Finished each of our three build briefs in under seven minutes — about a quarter of GPT-6 Astra’s time
  • A none setting for replies that need no reasoning with effort running up to max
  • Batch and every hosted tool from computer use to MCP at launch — plus GitHub Copilot

Cons

  • One index point above GPT-5.6 Sol — the step is in price rather than capability
  • Claude Opus 5.5 at high scores 54 for $1.82 a task — six points more
  • OpenAI’s head-to-heads pit it against last-generation Claude and GPT-6 Astra at low effort
  • Trails Claude Fable 5 on DeepSWE 1.1 at 68.8% against 69.9%
  • Still works around a warning in 64.4% of OpenAI’s adversarial cases against 17.4% for Astra
  • Each of our three builds had a fault you could see on screen
  • Not yet in regular ChatGPT chat — only in ChatGPT Work and Codex
  • Below a score of 48 it is no bargain — MiMo-V2.6-Pro scores 46 for $0.13 a task

GPT-6 Sol is the middle of OpenAI’s new range. It arrived on 22 September alongside the budget GPT-6 Luna and 19 days after GPT-6 Astra took the top slot, and it replaces GPT-5.6 Sol, which led OpenAI’s range until Astra arrived. OpenAI says Sol was trained with methods similar to Astra’s, and its pitch is about the bill: $2 per million input tokens and $10 per million output, a fifth of what Astra charges.

That case can be checked more thoroughly than most launch claims. Artificial Analysis measured Sol on launch day, and we ran Sol and Luna through the three build briefs that Astra answered in early September. The saving holds up, and it is much larger than the gain in capability, which is small. For anyone already on GPT-5.6 Sol, the move is easy. For anyone choosing from scratch, it comes down to whether a score in the high 40s is enough, because the obvious step up from Sol is not Astra. It is Claude Opus 5.5 at high effort.

The number that matters: $1.06 a task

Artificial Analysis runs the same evaluation suite on every model and publishes what one task cost to finish, which is the list price multiplied by the tokens the model actually spent. This site judges value on that figure rather than on the per-token sticker, and on that figure Sol has a strong case:

Model and effortIntelligenceCost per taskAgainst Sol’s billList price
Claude Opus 5.5, max58$5.985.6x$4 / $20
Claude Opus 5.5, high54$1.821.7x$4 / $20
GPT-6 Astra, max53$3.263.1x$10 / $50
GPT-6 Sol, max48$1.06$2 / $10
Muse Spark 1.3, max48$1.601.5x$1.25 / $4.25
GPT-5.6 Sol, max47$1.991.9x$4 / $20 (promotion)
Grok 4.7, xhigh46$3.743.5x$2 / $6
MiMo-V2.6-Pro46$0.130.12x
Intelligence Index score against cost per taskIndex v4.3.2 · USD per task · score axis starts at 44 · up and left is better · Artificial AnalysisIndex score44464850525456$0$1$2$3$4Cost of one index task (USD)GPT-6 Sol (max)48 · $1.06 a taskMuse Spark 1.3 (max)48 · $1.60 a taskGPT-5.6 Sol (max)47 · $1.99 a taskClaude Opus 5.5 (high)54 · $1.82 a taskGPT-6 Astra (max)53 · $3.26 a taskGrok 4.7 (xhigh)46 · $3.74 a task
No point sits above and to the left of GPT-6 Sol: nothing here scores as high for less. The dashed arrow is the step from GPT-5.6 Sol: one point higher and 47% cheaper per task. Claude Opus 5.5 at high sits well above both for $1.82, which is the trade a new buyer actually faces. Data: Artificial Analysis.

Against GPT-5.6 Sol, the gain is in the bill. Sol scores 48 to its predecessor’s 47 and finishes a task for $1.06 against $1.99, 53% of the cost. One point is the smallest step this index can show. The accurate summary is the same tier of result for about half the money, not a smarter model.

At 48 and above, nothing is cheaper. Of the models that score 48 or more, Sol has the lowest cost per task. Muse Spark 1.3 matches its score and costs half as much again. GPT-6 Astra scores five points more for three times the bill, and Claude Opus 5.5 at max leads the index at 58 for $5.98.

For a buyer starting fresh, the comparison that matters is Claude Opus 5.5 at high. It scores 54 for $1.82, six points above Sol for 72% more per task. That is a bigger lead over Sol than Astra has, at 44% less than Astra costs. Our Opus 5.5 review calls that configuration the value pick at the top of the market, and nothing about Sol changes that. What Sol changes is the price of the tier below it.

Below 48, Sol stops being the bargain. Xiaomi’s MiMo-V2.6-Pro scores 46 for $0.13 a task, an eighth of Sol’s cost for two points less. OpenAI’s own Luna scores 37 for $0.07. A workload that does not need a score of 48 is overpaying on Sol. Grok 4.7, which sold itself on value, scores 46 at xhigh and costs more than three times as much as Sol per task.

All of these figures are for Sol at max, the top of its effort dial. The API defaults to medium, so a request that leaves effort unset is not running the configuration measured here. It will cost less, and probably score lower, but neither figure has been published yet.

OpenAI’s benchmarks, and who it chose to beat

Unless marked otherwise, every figure in this section is OpenAI’s own, from its launch post, run at effort levels OpenAI chose. None has been reproduced independently. Read closely, the table does not say Sol is better than the models beside it. It says Sol is level with them and much cheaper, and on those terms it holds up.

OpenAI’s head-to-heads for GPT-6 SolScore, % · higher is better · Sol at xhigh, at max on DeepSWE · rival effort as labelled · OpenAIGPT-6 SolThe rival OpenAI put beside it02040608033.230.3AutomationBenchvs Astra low33.231.4AutomationBenchvs Fable 5.1 max33.226.9AutomationBenchvs Opus 5 max60.560.3OSWorld 2.0vs Opus 5 medium68.869.9DeepSWE 1.1vs Fable 5 xhigh
Every pair is close to level. Sol’s widest lead is 6.3 points over Claude Opus 5 at max, and on DeepSWE it trails Claude Fable 5 by 1.1. Fable 5.1’s AutomationBench run fell back to Opus 5 on about 40% of tasks. Claude Opus 5.5, released earlier the same day, is not in OpenAI’s table. All figures are OpenAI’s own.

The score margins are narrow. On AutomationBench 1.0.6, Zapier’s test of business workflows across 47 tools, Sol at xhigh scores 33.2%. That is 2.9 points above Astra at low effort, 1.8 above Claude Fable 5.1 with an Opus 5 fallback and 6.3 above Claude Opus 5 at max. On OSWorld 2.0 offline it leads Opus 5 at medium by 0.2 points. On DeepSWE 1.1, at max, it trails Claude Fable 5 at xhigh by 1.1. The cost margins, by contrast, are wide:

BenchmarkSolRival OpenAI choseRival’s scoreCost gap, in OpenAI’s terms
AutomationBench 1.0.633.2% (xhigh)Claude Opus 5, max26.9%Opus 5 costs 11.1x Sol
AutomationBench 1.0.633.2% (xhigh)Claude Fable 5.1, max, Opus 5 fallback31.4%Over 8.9x Sol, fallbacks not counted
AutomationBench 1.0.633.2% (xhigh)GPT-6 Astra, low30.3%Astra costs 3.9x Sol
OSWorld 2.0 offline60.5% (xhigh)Claude Opus 5, medium60.3%Sol about 80% cheaper
DeepSWE 1.168.8% (max)Claude Fable 5, xhigh69.9%Sol about 80% cheaper
Agents’ Last Exam56.4% (max)Claude Opus 5, its bestlower, not publishedSol 60% cheaper
FrontierCode 1.1not publishedClaude Fable 5.1, xhighlevel with Sol, per OpenAISol much cheaper, no figure given

The table has two blind spots. The first is which rivals it uses. Every Claude model in it had been superseded by the time Sol shipped, because Claude Opus 5.5 launched earlier the same day. That is timing rather than selection, but it leaves out the comparison buyers now face. Anthropic’s own AutomationBench figure for Opus 5.5 is 40.0%, from a different company’s run, so it is not a head-to-head; it is still well above Sol’s 33.2%. The second is the effort settings. The Astra that Sol beats on AutomationBench is running at low, the bottom of Astra’s dial, while Sol at xhigh sits one notch below the top of its own.

Factuality gets a ratio, not a percentage. OpenAI replayed real ChatGPT exchanges that users had reported for a factual error, and counts roughly half as many errors from Sol as from GPT-5.6 Sol. Every item in that set is one an earlier model got wrong, so it measures the hard cases rather than a typical day.

OpenAI’s alignment tests: where Sol still trails Astra% of adversarial cases, run at max effort · lower is better · OpenAI’s labelled barsGPT-5.6 SolGPT-6 SolGPT-6 LunaGPT-6 Astra02040608010.41.32.80.5Coding deception77.54.928.71.5Hides a brokensearch tool68.264.442.417.4Works around awarning
Sol’s rate of hiding a broken tool fell from 77.5% to 4.9%, and coding deception from 10.4% to 1.3%. Working around a warning barely moved, from 68.2% to 64.4%, where Luna improved to 42.4% and Astra sits at 17.4%. The tests are built to provoke these failures. Data: OpenAI.

For agent work, the alignment results may matter more than any score above. In the two tests closest to an agent reporting on its own work, Sol changed dramatically. Its coding-deception rate, on tasks designed to lure a model into overstating what it did, fell from 10.4% to 1.3%. Given a broken search tool, it pressed on as though the tool were fine in 4.9% of cases, down from 77.5%. Bypassing a code reviewer went from 7.3% to zero, and acting without authorisation from 51.9% to 11.3%. Each of those is a small fraction of GPT-5.6 Sol’s rate.

Plan around the row that barely moved. Faced with a warning that ought to stop the job, Sol worked around it in 64.4% of OpenAI’s cases, against 68.2% for GPT-5.6 Sol and 17.4% for Astra. Luna, at 42.4%, does better than Sol on this test. The tests are adversarial by design, so the figure describes a worst case. Even so, an agent that can spend money, deploy code or delete data should have its stop conditions enforced by the harness around Sol, not left to the model.

What we found running it ourselves

On 22 September we ran Sol and Luna in OpenAI’s Codex at max reasoning effort on the three briefs GPT-6 Astra answered in its September build-off: a busy motorway interchange and a seaside fairground at dusk, each as one animated SVG, and a cinematic rocket launch as one web page. Each model had one attempt per brief, with no browser, no preview and no way to run code, so neither could look at what it had built. Every run stayed inside those rules; the Codex logs show nothing but file writes.

The six builds, scored side by side with Astra’s September files, are in GPT-6 Astra vs Sol vs Luna. Astra won every round and finished with 10 points; Sol took 5 and Luna took 3.

Time to a finished file on our three build briefsWall-clock minutes, one decimal · lower is faster · Sol and Luna at max · Astra at xhigh · BitsMindsGPT-6 SolGPT-6 LunaGPT-6 Astra (September)01020304.88.116.3Motorwayinterchange6.39.126.6Seasidefairground6.64.528.0Rocket launch
Sol finished every brief in under seven minutes, about a quarter of Astra’s total time, at a higher effort setting. Wall-clock time also depends on server load, and Astra’s runs were on other days, so read the ratio rather than the seconds. Runs by BitsMinds in OpenAI Codex.

The first thing the runs showed was speed. Sol took 4 minutes 45 seconds on the interchange, 6 minutes 18 seconds on the fairground and 6 minutes 36 seconds on the launch. Astra, at xhigh, which was then the top of Codex’s dial, took 16 minutes 15 seconds, 26 minutes 33 seconds and 28 minutes. Sol ran one setting higher and still used about a quarter of Astra’s total time.

Sol’s interchange draws its flyover ramps over the lower road with real layering and a cast shadow, which the brief makes a hard requirement. The trouble is at both ends of every ramp. At the lower road each one starts as a rounded stub in the middle of the carriageway rather than peeling off a lane, and the cars that should join the north–south overpass are drawn beneath its deck, so they disappear under the bridge instead of driving onto it. The ramps begin in the middle of one road and lead nowhere on the other.

Its fairground has both rides moving as the brief describes. The Ferris wheel’s rim, spokes and gondolas turn together as one assembly, with the gondolas held level. The carousel turns too: eight horses ride an elliptical track under a rotating canopy, upright and bobbing out of phase. But they never turn to face the way they travel, so every horse points the same way all the way round and rides backwards on the far side. The reflection in the water that the brief requires never appears. Its clipping rectangle sits in the flipped coordinate space of the reflection and cuts the reflected rides away.

Its launch is the cleanest-looking of Sol’s three builds: a countdown, ignition, liftoff and a sustained climb under a mission-control HUD, with the camera holding the rocket. The flight never gets past ascent. There is no engine cutoff and no staging, the engines stay at 100% for as long as the page runs, and during the climb the rocket’s nose sits cut off at the top of the frame.

Luna took 8 minutes 6 seconds, 9 minutes 6 seconds and 4 minutes 29 seconds. Its interchange has the same ramp fault as Sol’s from the other direction: each ramp climbs out from under the middle of the overpass, and the cars on it are hidden beneath the deck. Its fairground reflection works, but the carousel canopy and platform spin flat in the picture plane like a propeller, on 14- and 18-second cycles, so the horses loop vertically. Its launch has the fuller flight plan of the two models, with stage separation and orbital insertion written into the HUD, but the code adds the camera offset to the rocket twice. By about T+11 the rocket has sunk out of the bottom of the frame, and staging and orbit play out against an empty sky.

Across Sol’s three builds the pattern is consistent. The parts that need planning, such as ramps layered over a road, a wheel that turns as one assembly and a camera that follows the climb, are handled well. Each build then has one thing wrong that a single look at the output would have shown. At under seven minutes a brief, a look and a second pass together still take far less time than one Astra run. That is the practical way to use Sol for visual work: generate a quick draft, check it by eye, and send it back. These are three visual briefs, and they say nothing about the business-workflow and agent tasks where OpenAI’s table is strongest.

What changes for developers

The API ids are gpt-6-sol and gpt-6-luna. Both take text and images and return text, with a 1.05M-token context window, up to 922K tokens of input and 128K of output, on the Responses, Chat Completions and Batch APIs. The Sol model page and its Luna counterpart list the details that decide the bill:

  • Effort now starts at none. The dial runs none, low, medium, high, xhigh and max, and the default is medium. Astra has no none setting. The independent figures above are for max, so measure medium and high on your own jobs before paying for the top. None is there for replies that need no reasoning at all.
  • Cache reads are cheap, and cache writes cost a little extra. Cached input is $0.20 per million tokens, a tenth of the normal input rate, and a cache write is $2.50, a quarter above it. A prefix that is read back even once has already paid for its write: two uses cost $2.70 per million prefix tokens, against $4.00 uncached. OpenAI says changing effort or switching tools mid-conversation no longer discards the cached prefix, which helps agents that do both.
  • Every hosted tool is supported: computer use, web search, file search, code interpreter, a hosted shell, apply_patch, MCP and tool search.
  • The knowledge cutoffs differ. Sol’s training data runs to 20 April 2026 and Luna’s to 18 May, so the cheaper model knows about more recent events.

The 50% cut is real, but it is measured from a discounted price. OpenAI compares Sol’s $2/$10 with GPT-5.6 Sol’s $4/$20, and that rate is itself a promotion that began on 21 August. GPT-5.6 Sol listed at $5/$30 at launch, and measured from that, Sol takes 60% off input and two thirds off output. OpenAI has said the GPT-6 prices are not introductory: a spokesperson described them to VentureBeat as permanent. For a team weighing the move, the discount on the old model is temporary and the new price is not.

Inside ChatGPT, subscribers on Plus, Pro, Business, Enterprise and Edu get both models in ChatGPT Work and in Codex, while the Free and Go plans get Luna through the desktop app. Neither is in regular Chat yet, so someone using ChatGPT the ordinary way will not see either model for now. GitHub Copilot carries Sol on Pro+, Max, Business and Enterprise, and Luna on those plans and on Pro. The pricing calculator has both at their new rates, but it multiplies list prices, and this launch is a reminder that the list price is only one of the two numbers that set what a task costs.

And Luna?

GPT-6 Luna is the budget tier, at $0.10 per million input tokens and $0.50 output, and it is the model that free ChatGPT users now meet in the desktop app. On the independent index it matches its predecessor’s score for 39% of the cost:

Model and effortIntelligenceCost per taskList price
GPT-6 Sol, max48$1.06$2 / $10
MiMo-V2.6-Pro46$0.13
GPT-6 Luna, max37$0.07$0.10 / $0.50
GPT-5.6 Luna, max37$0.18$0.20 / $1.20

At $0.07 a task, Luna has the lowest cost per task on our leaderboard. Its score of 37 is eleven points below Sol’s. OpenAI’s own figures are kinder. On DeepSWE 1.1 at max it scores 66.6%, which OpenAI puts level with Claude Opus 5 and Fable 5 run at medium effort, at 93% and 96% lower cost per task respectively. At higher effort, OpenAI says, it matches GPT-5.6 Sol on the factuality test for about a hundredth of the cost.

Its alignment record is mixed. Luna never acted without authorisation in OpenAI’s test, and at 42.4% it works around a warning less often than Sol does. It also carried on with a broken search tool in 28.7% of cases, nearly six times Sol’s rate, and each of its three builds for us had a fault you could see. Luna is for high-volume classification, extraction, routing and short drafting, where cost is the constraint and a person or a test checks the output. It is not a model to leave running an agent on its own. For teams not committed to OpenAI, MiMo-V2.6-Pro scores nine points higher for six cents more per task.

Who should be running it

If you are…Do this
On GPT-5.6 SolSwitch. Sol scores a point more for 53% of the cost per task, and its rates are permanent where GPT-5.6 Sol’s $4/$20 is a promotion
On GPT-6 AstraKeep Astra for computer use and the hardest agent runs, and test Sol on everything else. Astra scores five points more for three times the cost per task
Choosing between Sol and Claude Opus 5.5Decide what score the work needs first. At high, Opus 5.5 scores 54 for $1.82 a task; if 48 is enough, Sol gets there for $1.06
Running agents that can spend, deploy or deleteEnforce the stop conditions outside the model. In OpenAI’s own adversarial test Sol worked around a warning in 64.4% of cases
Building front-ends or animationsUse the speed, then look at the result. Our three builds took under seven minutes each, and each had a fault one look would catch
Serving high volume on a tight budgetTry Luna first at $0.07 a task. If you are not tied to OpenAI, MiMo-V2.6-Pro scores 46 for $0.13
Picking an effort levelStart at medium, the default, or high. The independent figures are for max, and none is there for replies that need no reasoning

Verdict

GPT-6 Sol earns a 4.2 out of 5, below the 4.5 we gave GPT-6 Astra and above the 4.0 we gave Grok 4.7. It delivers the price the launch promised, and the independent numbers confirm it. GPT-5.6 Sol customers get the same tier of result for about half the cost per task, and no model scoring 48 or more on Artificial Analysis finishes a task for less. It was quick in our hands. On OpenAI’s tests it misreports its work and covers for broken tools far less often than the model it replaces.

The missing points come from four places. The capability step is one index point. The head-to-heads OpenAI published are against Claude models that Opus 5.5 had superseded earlier the same day, and Claude Opus 5.5 at high scores six points more for $1.82. The warning row in OpenAI’s alignment chart barely moved. And in our own build-off it took half of Astra’s points, with a fault in each of its three builds that one look at the output would have caught. None of this stops Sol from being the sensible default in OpenAI’s range. It does stop it from being the default for everything.

Anyone on GPT-5.6 Sol should move now and start at medium or high rather than max. Anyone choosing fresh should first decide whether the work needs a score in the mid-50s. If it does, Opus 5.5 at high is the better purchase. If it does not, Sol is the cheapest way to get a score of 48, and the leaderboard shows what is available below that line for less.

About the score. 4.2/ 5 is BitsMinds' editorial verdict from our own testing and research — not an average of user ratings, which we do not collect. Prices and plan tiers are as published by the vendor on the fact-check date shown above. How we rate →

Related Reviews

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.