Claude Opus 5.5 vs GPT-6 Sol and Luna: Benchmarks and Cost
Claude Opus 5.5 vs GPT-6 Sol and Luna across 12 benchmarks, with cost per task at every effort level: which model wins, what it costs, and when to use each.
OpenAI says GPT-6.1 Sol nearly matches GPT-6 Astra for a fifth of the token price. I checked that claim against every chart in the launch and every independent leaderboard that had already run it, then put Claude Opus 5.5 next to both. Here's where the claim holds, where it breaks, and which model to use for what.
Mostly true. GPT-6.1 Sol lands within 3.5 points of GPT-6 Astra, or beats it, on 8 of the 10 charts that include both, at 8% to 23% of Astra's cost per task. It breaks on science: 11.1 points behind. Claude Opus 5.5 still wins on coding and documents, but Sol now matches it on terminal work for about a third of the cost.
Your default for well-specified work at volume: automations, reviews, computer use, terminal jobs with tests. Run it at medium, and go to max for shell and desktop work.
Still the pick when quality decides: real codebases, front ends, documents people read, and anything long and ambiguous. Run it at medium.
Keep it for the jobs 6.1 Sol measurably loses: scientific computing, the hardest multi-app automations, and runs that already failed once on Sol.
Same token price as 6.1 Sol and worse on almost every chart. Switch the model ID.
OpenAI's headline is near-Astra performance at one-fifth the price. The price part is exact. The performance part depends on the job.
The token price is a clean one-fifth: $2 and $10 per million input and output tokens against Astra's $10 and $50. Cost per task isn't a fixed ratio, because the two models use different amounts of tokens and tools. On the charts below, Sol costs 8% to 23% of what Astra costs for the same task.
The headline number comes from Artificial Analysis, which runs every model the same way. Sol at max scores 51.8 on its Intelligence Index for $0.72 a task. Astra at max scores 52.7 for $3.26. That's 98% of the score for 22% of the cost. It's an index, not a percentage of intelligence, but it's the fairest single comparison on launch day. For Astra's own launch numbers, see GPT-6 Astra release date, price and benchmarks.
Computer use is the strongest case. On OpenAI's OSWorld 2.0 chart, Sol at max scores 71.4% for $1.27 and Astra at max scores 73.5% for $9.44. On DeepSWE, OpenAI's long engineering benchmark, Sol at high (75.2%) beats Astra's best setting (74.1% at xhigh) for about a seventh of the cost. On GDPval-AA, the one independent test of real work products that has both, Sol at max outscores Astra at max: 1575 Elo against 1542 Elo.
Scientific computing. On OpenAI's own Terminal-Bench-Science chart, Sol at max scores 57.0% and Astra at max scores 68.1%. Even at its best setting, Sol trails by 11.1 points. Business automations are the other soft spot: Sol tops out at 36.1% on AutomationBench against Astra's 41.4%. And on FrontierCode, Astra at max still beats Sol's best by 3.1 points, though it costs $4.59 a task to Sol's $0.36.
Intelligence Index
Hover any dot to read its score and cost. Click it to pin it and see the cost per solved task, plus the cheapest setting where each other model catches up. Use the legend to toggle models on and off.
Index v4.3.2. GPT-6.1 Sol captured Sept 29 at every effort. Rival efforts come from the Sept 22 capture; every rival's max point was re-captured Sept 29 and is unchanged. Source: Artificial Analysis.
| Model | Effort | Index | Cost per index task |
|---|---|---|---|
| GPT-6.1 Sol | low | 42.1 | $0.13 |
| GPT-6.1 Sol | medium | 47.8 | $0.21 |
| GPT-6.1 Sol | high | 50.2 | $0.32 |
| GPT-6.1 Sol | xhigh | 51.0 | $0.39 |
| GPT-6.1 Sol | max | 51.8 | $0.72 |
| GPT-6 Astra | low | 45.8 | $0.82 |
| GPT-6 Astra | medium | 49.6 | $1.54 |
| GPT-6 Astra | high | 50.9 | $1.73 |
| GPT-6 Astra | xhigh | 52.4 | $2.31 |
| GPT-6 Astra | max | 52.7 | $3.26 |
| Claude Opus 5.5 | low | 42.3 | $0.55 |
| Claude Opus 5.5 | medium | 51.2 | $1.34 |
| Claude Opus 5.5 | high | 53.6 | $1.82 |
| Claude Opus 5.5 | xhigh | 56.0 | $3.46 |
| Claude Opus 5.5 | max | 57.6 | $5.98 |
| Claude Fable 5.1 | low | 46.8 | $2.37 |
| Claude Fable 5.1 | medium | 48.9 | $2.98 |
| Claude Fable 5.1 | high | 51.2 | $3.91 |
| Claude Fable 5.1 | xhigh | 53.2 | $5.98 |
| Claude Fable 5.1 | max | 53.4 | $7.63 |
Charts give you averages. These are real builds people posted on September 29, most of them the same task on two or more models. Each is the creator's own test, with what they did and didn't disclose.
jjcm rebuilt the same Star Trek LCARS-style landing page design with GPT-6.1 Sol and with Claude Opus 5.5 and posted both. They prefer Opus for alignment, SVG animation and polish, and call Sol quick and cheap.
Caveat. Prompt, effort, retries and cost are not disclosed.
BridgeBench gave GPT-6.1 Sol, GPT-6 Sol, GPT-6 Astra and Claude Opus 5.5 the same ocean scene. It reports 97 seconds and $0.07 for GPT-6.1 Sol, 44 seconds and $0.07 for GPT-6 Sol, 3 minutes 54 seconds and $0.59 for Astra, and 6 minutes 45 seconds and $0.90 for Opus 5.5, and it rates Sol's ocean the weakest of the four.
Caveat. One scene, BridgeBench's own harness, effort settings not disclosed.
Lumina ran the same prompt on four models and says Astra and the new Sol feel very similar.
Caveat. The video labels the Astra run as GPT-6 Astra Pro, and the prompt, effort and cost are not published.
A split-screen NYC scene built by GPT-6.1 Sol and Claude Sonnet 5.5 at max. The detailed prompt is in the author's reply under the post.
Caveat. The comparison model is Sonnet 5.5, not Opus 5.5, and Sol's effort and cost aren't stated.
An animated ship in a glass bottle with a moving ocean and a lantern-lit table, built with GPT-6.1 Sol. Conor reports 55 minutes and an estimated API cost of 3.47.
Caveat. The currency isn't stated, the idea was his, and the prompt isn't published.
Theo had early access. He likes GPT-6.1 Sol for code reviews, audits and investigation, and keeps Claude Opus 5.5 for implementation and long unattended work, where Sol stalled on a TypeScript-to-Rust port.
Caveat. Early access without payment from OpenAI, per Theo. The video is sponsored by Depot. His own benchmark runs were incomplete, so this page doesn't use them.
Simon ran "Generate an SVG of a pelican riding a bicycle" at every effort. Output grew from 2,100 tokens and 32 seconds at low to 13,480 tokens and 252 seconds at max. He saw no notable difference from the GPT-6 family's pelicans.
Caveat. A playful visual probe, not a coding test. OpenAI gave Simon a free DevDay ticket, which he disclosed.
Demos load from their original posts and pages as you scroll to them, and each has a link to the original. Each one is the creator's own test, not a controlled benchmark.
Opus 5.5 still has the higher ceiling. What changed this week is how close Sol gets for the money, especially on terminal work.
Cognition's FrontierCode grades real pull requests on whether a maintainer would merge them. Opus 5.5 at its default medium effort scores 54.6% for $0.80 a task, still the best result on the board. Sol's best is 50.2%, also at medium, for $0.36. That beats Opus 5.5 at low (47.3% for $0.40) for less money, so for small, well-scoped tickets Sol is the cheaper good-enough option.
Last week this was Opus 5.5's biggest win. Now, on Artificial Analysis's Terminal-Bench 4.0 run, Sol at max scores 56.1% for $1.82 a task, and Opus 5.5 at high scores 56.6% for $5.12. Opus 5.5 still has the top score (59.6%), but it pays $13.11 a task for it.
On GDPval-AA, where judges compare real work products head to head, Opus 5.5 is far ahead: 1846 Elo against Sol's 1575 Elo. If a person has to read the output, start with Opus. On OpenAI's AutomationBench chart, which reuses Zapier's published Opus 5.5 results with fallbacks, Sol at xhigh (35.5% for $0.25) matches Opus 5.5 at xhigh (35.8% for $0.89). Opus 5.5 at max tops that chart at 42.5%. On OpenAI's office-document test, GDP.pdf, Sol at high scores 32.0% against Opus 5.5's best of 28.8%, on OpenAI's run of it. On science, Opus 5.5 at max is 63.3% to Sol's 57.0%.
Terminal-Bench 4.0
Hover any dot to read its score and cost. Click it to pin it and see the cost per solved task, plus the cheapest setting where each other model catches up. Use the legend to toggle models on and off.
AA's own Terminal-Bench 4.0 run, not the tbench.ai leaderboard (which has no GPT-6.1 Sol yet). Cost is AA's real cost per Terminal-Bench task. Rival efforts from Sept 22; max points re-checked Sept 29. Source: Artificial Analysis.
| Model | Effort | Score | Cost per task | Output tokens |
|---|---|---|---|---|
| GPT-6.1 Sol | low | 30.8% | $0.38 | 17.4k |
| GPT-6.1 Sol | medium | 48.0% | $0.61 | 32.0k |
| GPT-6.1 Sol | high | 51.5% | $0.83 | 43.6k |
| GPT-6.1 Sol | xhigh | 54.0% | $1.03 | 56.9k |
| GPT-6.1 Sol | max | 56.1% | $1.82 | 108k |
| GPT-6 Astra | max | 59.1% | $8.50 | |
| Claude Opus 5.5 | low | 31.3% | $2.08 | |
| Claude Opus 5.5 | medium | 52.5% | $4.04 | |
| Claude Opus 5.5 | high | 56.6% | $5.12 | |
| Claude Opus 5.5 | xhigh | 59.6% | $8.78 | |
| Claude Opus 5.5 | max | 59.6% | $13.11 | |
| Claude Fable 5.1 | max | 52.0% | $19.22 |
I compared Opus 5.5 with last week's GPT-6 Sol in Claude Opus 5.5 vs GPT-6 Sol and Luna, and tested Opus 5.5 hands-on in my Claude Opus 5.5 review. If you want to run both side by side in one editor, Claude Code can point at any provider: how to change the model in Claude Code.
GPT-6 Sol launched on September 22. Seven days later, 6.1 Sol has the same input and output price and better numbers on almost every chart.
The biggest jumps are in long tool loops: Terminal-Bench at max goes from 43.9% to 56.1% while getting cheaper per task, and OpenAI's science chart goes from 27.6% to 57.0%. FrontierCode at medium improves from 45.9% to 50.2% at less than half the cost.
Two things didn't improve. FrontierCode at max slips from 49.3% to 47.6%. And the factual-error rate at max is flat (4.6% to 4.6%). The error-rate improvement is real at low effort, from 11.4% to 7.7%, so don't read it as a blanket drop in hallucinations. OpenAI's launch post doesn't say GPT-6 Sol is being retired; if you pinned gpt-6-sol, switch the model ID to gpt-6.1-sol yourself.
Terminal-Bench 4.0
Hover any dot to read its score and cost. Click it to pin it and see the cost per solved task, plus the cheapest setting where each other model catches up. Use the legend to toggle models on and off.
AA's own Terminal-Bench 4.0 run, not the tbench.ai leaderboard (which has no GPT-6.1 Sol yet). Cost is AA's real cost per Terminal-Bench task. Rival efforts from Sept 22; max points re-checked Sept 29. Source: Artificial Analysis.
| Model | Effort | Score | Cost per task | Output tokens |
|---|---|---|---|---|
| GPT-6.1 Sol | low | 30.8% | $0.38 | 17.4k |
| GPT-6.1 Sol | medium | 48.0% | $0.61 | 32.0k |
| GPT-6.1 Sol | high | 51.5% | $0.83 | 43.6k |
| GPT-6.1 Sol | xhigh | 54.0% | $1.03 | 56.9k |
| GPT-6.1 Sol | max | 56.1% | $1.82 | 108k |
| GPT-6 Astra | max | 59.1% | $8.50 | |
| GPT-6 Sol | low | 9.1% | $0.49 | |
| GPT-6 Sol | medium | 18.7% | $1.12 | |
| GPT-6 Sol | high | 26.3% | $1.60 | |
| GPT-6 Sol | xhigh | 30.3% | $1.91 | |
| GPT-6 Sol | max | 43.9% | $4.00 |
More thinking isn't always better with 6.1 Sol. On two coding charts, the top score comes before max.
On FrontierCode, Sol scores 45.5%, 50.2%, 48.0%, 49.3%, 47.6% from low to max. Medium is the peak, and max costs more than twice as much for a lower score. On OpenAI's DeepSWE chart, high (75.2%) beats both xhigh and max (71.9%). Readers on r/codex noticed the same shape on OpenAI's chart on launch day.
Terminal and computer-use work is the exception. Sol keeps climbing to max on Terminal-Bench, OSWorld and the science chart. My rule: start at medium, the default. Go to high for long engineering tasks, and max for shell and desktop agents. Drag the slider to see what each model gets you at a fixed budget per task.
Official prices captured on launch day, September 29, 2026. Check OpenAI's pricing page before you budget; rollout and plan limits change.
Ultrafast isn't out yet. OpenAI announced a GPT-6.1 Sol Ultrafast mode at DevDay for the coming days, with up to 8x faster token generation in Codex. Its own pages disagree on the API figure (up to 6x in the DevDay recap, up to 8x in the API guide), and there is no Sol-specific price yet. Today's Astra Ultrafast is limited to the $500 Pro tier and eligible Enterprise and Edu plans. Faster tokens also don't guarantee faster tasks: Artificial Analysis measured 6.1 Sol at max generating slower than GPT-6 Sol at max.
If you want one API key across OpenAI and Anthropic models, OpenRouter listed 6.1 Sol on launch day: what OpenRouter is and how it works.
Pick the job and what matters most. Each answer shows the numbers behind it, one source per row.
The pattern: 6.1 Sol for well-specified work you run a lot, Opus 5.5 when quality decides, Astra only where a chart shows it winning. One split worth trying is the one Theo describes in his review: let Opus 5.5 build and let Sol review and audit the result.
Launch-day reactions from people who actually used it, weighed by how many independent sources back each theme and whether the benchmarks agree.
| Signal | Strength | What it rests on | What to do with it |
|---|---|---|---|
| Close to Astra on most tests, for a fraction of the cost | Strong | Artificial Analysis, Cognition and seven of OpenAI's own charts. Lumina, who ran both on the same prompt, says they feel very similar. | Move Astra workloads to 6.1 Sol and keep Astra for the jobs your own tests say it wins. |
| Science and long terminal work still favor Astra | Solid | OpenAI's own Terminal-Bench-Science chart shows an 11-point gap. Artificial Analysis's separate run shows a smaller one. | Don't route research computing or long shell jobs to Sol without testing first. |
| Opus 5.5 still makes the more polished front end | Early | jjcm's side-by-side LCARS pages and BridgeBench's ocean scene both favor Opus. GDPval-AA, a judged deliverables test, has Opus far ahead. | For UI and anything a person will read, start with Opus 5.5. |
| Sol for reviews, Opus for building | Thin | One early-access reviewer (Theo) with a detailed video. No controlled test yet. | A good split to try: let Sol audit and review what Opus writes. |
| More effort doesn't always score higher | Solid | FrontierCode peaks at medium and DeepSWE at high. Readers on r/codex spotted the same shape on OpenAI's chart. | Start at medium, the default, and only go up where a chart shows it pays. |
| Faster tokens don't mean faster tasks | Early | BridgeBench timed 6.1 Sol at 97 seconds on a scene GPT-6 Sol did in 44. Ultrafast isn't out yet. | Judge speed on your own tasks, not on tokens per second. |
Treat all of this as early. Most launch-day posts talked about the claims and older models rather than new tests, and nobody has published a controlled head-to-head yet. The gaps matter too: I couldn't read the AP and WIRED originals, and I didn't count search-preview comments that weren't in the threads I opened.
All ten benchmarks, each model's best or max setting. Switch to independent sources only to hide OpenAI's own charts.
5 cases where the same model at the same effort gets different numbers from two sources. This is why every chart above sticks to one source at a time.
GPT-6 Astra at max: 68.1% on OpenAI's chart and the official board, 63.3% in Artificial Analysis's run
OpenAI's number matches the official board. AA's separate run lands 4.8 points lower. The cause isn't published, so compare within one source.
| Astramax OpenAI launch chart | 68.1% | $23.80 |
| Astramax Terminal-Bench-Science board | 68.1% | no cost |
| Astramax Artificial Analysis | 63.3% | no cost |
Claude Opus 5.5 at max: 63.3% on OpenAI's chart and the official board, 59.0% in Artificial Analysis's run
OpenAI's number matches the official board. AA's separate run lands 4.3 points lower. The cause isn't published, so compare within one source.
| Opus 5.5max OpenAI launch chart | 63.3% | $23.21 |
| Opus 5.5max Terminal-Bench-Science board | 63.3% | no cost |
| Opus 5.5max Artificial Analysis | 59.0% | no cost |
GPT-6.1 Sol at max: 57.0% on OpenAI's chart, 58.1% in Artificial Analysis's run
Here AA scores Sol slightly higher than OpenAI does, while AA scores Astra and Opus lower. In AA's run the gap to Astra shrinks from about 11 points to about 5.
| 6.1 Solmax OpenAI launch chart | 57.0% | $5.47 |
| 6.1 Solmax Artificial Analysis | 58.1% | no cost |
Claude Opus 5.5 at max on GDP.pdf: 26.2% on OpenAI's chart, 30.6% on Surge's board
Surge's own board scores it 4.4 points higher than OpenAI's chart. Surge hasn't added GPT-6.1 Sol, so there is no independent GDP.pdf number for it yet.
| Opus 5.5max OpenAI launch chart | 26.2% | $1.55 |
| Opus 5.5max Surge GDP.pdf board | 30.6% | no cost |
GPT-6 Astra at max on GDP.pdf: 31.0% on OpenAI's chart, 34.2% on Surge's board
Surge's own board scores it 3.2 points higher than OpenAI's chart. Surge hasn't added GPT-6.1 Sol, so there is no independent GDP.pdf number for it yet.
| Astramax OpenAI launch chart | 31.0% | $2.08 |
| Astramax Surge GDP.pdf board | 34.2% | no cost |
Close on most tests, not all. On the Artificial Analysis Intelligence Index, GPT-6.1 Sol at max scores 51.8 against Astra's 52.7, for 22% of the cost per task. It lands within 3.5 points of Astra, or ahead, on 8 of the 10 charts that include both. The big exception is scientific terminal work, where it trails by 11.1 points on OpenAI's own chart.
Not overall. Opus 5.5 scores higher on the Intelligence Index (57.6 vs 51.8), on FrontierCode coding (54.6% vs 50.2%) and by a wide margin on GDPval-AA work products. Sol is the value pick: on Terminal-Bench it gets within a point of Opus 5.5 at high for 36% of the cost per task, and on business automations it matches Opus 5.5 at xhigh for under a third of the price.
Same token price, better results, one week apart. The Intelligence Index at max rises from 47.5 to 51.8 while the cost per task drops from $1.06 to $0.72. Terminal-Bench at max jumps from 43.9% to 56.1%, and OpenAI's Terminal-Bench-Science chart roughly doubles, from 27.6% to 57.0%. Cached input got cheaper too: $0.10 per million tokens instead of $0.20.
In the API, $2 per million input tokens, $0.10 per million cached input tokens and $10 per million output tokens, one fifth of GPT-6 Astra's $10 and $50. Requests above 272K input tokens are billed at 2x input and 1.5x output for the whole request. Batch and Flex are half price, and the Fast tier is double. In ChatGPT it comes with paid plans, starting with Plus at $20 a month.
No. At launch OpenAI made it available in the API and to eligible paid ChatGPT plans in ChatGPT Work and Codex, rolling out by account and workspace. There is no free-tier access in the launch materials, and it is not in the regular ChatGPT chat model picker.
A faster serving mode OpenAI announced at DevDay 2026 as coming soon, with up to 8x faster token generation in Codex. It was not available at launch, and OpenAI had not published its price. Today's Astra Ultrafast bills 8x included usage or 6x purchased credits, but OpenAI hasn't said the Sol version works the same way.
Medium, the default, for most coding: it is Sol's best setting on FrontierCode (50.2% for $0.36), and every higher setting scores lower there. Use high for long engineering tasks, where OpenAI's DeepSWE chart peaks. Use max for terminal work and computer use, where the curves keep climbing.
OpenAI's launch post doesn't mention one. TechCrunch reported that OpenAI dropped a planned GPT-6.1 Astra release over safety concerns raised during internal testing. That is press reporting, not an OpenAI statement I could read directly.
Partly. On launch day, Artificial Analysis had it at every effort level and Cognition's FrontierCode had it at every effort. Zapier AutomationBench, Cursor's CursorBench, Datacurve's DeepSWE, tbench.ai, the OSWorld board, ARC Prize, LMArena and Surge's GDP.pdf board had not added it yet, so several numbers on this page come from OpenAI's own charts, and each chart says so.
Every number on this page comes from a public source captured on launch day, September 29, 2026. Nothing is estimated.
OpenAI's launch charts were read from the chart data embedded in the page, and every point was checked against the rendered chart labels. Artificial Analysis and Cognition numbers come from their published data, re-parsed from the saved pages. Rival effort curves on Artificial Analysis come from my September 22 capture; I re-checked every rival's max point on September 29 and none had changed. Terminal-Bench costs are Artificial Analysis's real cost per task. Every chart is labeled with its one source; none combines two. You can download all 193 data points as a CSV.
Community builds and reactions are attributed to their creators and linked. I didn't rerun them, and I only describe what each creator disclosed. If a leaderboard adds 6.1 Sol later, I'll update this page and its date.
Reuse the data: the compiled dataset is licensed CC BY 4.0. Use it anywhere, including commercially, if you credit Moe Lueker, link back here and to the license, and say what you changed. The underlying benchmark results still belong to the organizations that published them.
Claude Opus 5.5 vs GPT-6 Sol and Luna across 12 benchmarks, with cost per task at every effort level: which model wins, what it costs, and when to use each.
Claude Opus 5.5 vs GPT-6 Astra, Sol and Luna on the same two game prompts. Opus built the best games both times, but the runner cost $15.25 to Astra's $7.09.
OpenAI released GPT-6 Astra on September 3, 2026. The release date, the API price, the benchmarks OpenAI buried, and the line in the system card nobody is quoting.