Back to blog
AI Tools471 data points · 12 benchmarks · 7 models

Claude Opus 5.5 vs GPT-6 Sol and Luna: Benchmarks and cost per task, at every effort level

Anthropic and OpenAI shipped three models on September 22. I pulled every chart from both launches and five independent leaderboards, then plotted score against cost at every effort level. Here's where each model wins, what it costs, and how to use them together.

Intelligence per dollar, every effort levelArtificial Analysis Intelligence Index
2030405060$0.01$0.10$1$10cost per task, log scaleGPT-5.6 SolGPT-6 AstraLuna37.3 · $0.068 maxSol47.5 · $1.06 maxOpus 5.557.6 · $5.98 max

The short answer

Opus 5.5 is the best model of the three and, per task, often the cheapest way to get a good result. Sol is GPT-5.6 Sol at half the price. Luna is the pennies tier. Most teams should run all three.

  • Claude Opus 5.5Best overall · new

    Your default for hard work: real codebases, terminal and infrastructure, documents, anything ambiguous. Run it at medium.

    • Tops the Artificial Analysis Intelligence Index: 57.6 at max
    • Best FrontierCode score on the board: 54.6% at medium for $0.80
    • Terminal-Bench: 52.5% at medium, the same $0.40 Sol pays at max for 43.9%
    $4 / $20in / out
    per 1M tokens
  • GPT-6 SolBest value · new

    Well-specified work at volume: business automations, tickets with tests, anything you ran on GPT-5.6 Sol. Run it at high or xhigh.

    • Same Intelligence Index score as GPT-5.6 Sol at every effort, for 45% to 53% of the cost
    • AutomationBench at high: 31.2% for $0.24, vs Opus 5.5 at 32.0% for $0.65
    • Weak spot: Terminal-Bench, 43.9% at max
    $2 / $10in / out
    per 1M tokens
  • GPT-6 LunaCheapest · new

    Triage, extraction, summaries and sub-agent grunt work. Always run it at max: it's still pennies.

    • Cheapest model on the Intelligence Index board: 37.3 for $0.068
    • FrontierCode 42.4% for $0.10, above Sol at low
    • Falls apart in long tool loops: 12.6% on Terminal-Bench
    $0.10 / $0.50in / out
    per 1M tokens
  • GPT-6 AstraOpenAI's ceiling · flagship

    The escalation model for long workflows across many apps and for computer use.

    • Top AutomationBench score: 41.4% for $1.73
    • Ties Opus 5.5 on Terminal-Bench: 59.1% vs 59.6%
    • Opus 5.5 at high beats Astra at max on the Intelligence Index for 56% of the cost
    $10 / $50in / out
    per 1M tokens
  • Claude Fable 5.1Hard to justify now · flagship

    Opus 5.5 beat it on every benchmark both were run on. Keep it only where your own tests say otherwise.

    • Costs 2.5x Opus 5.5 per token
    • CursorBench: Opus 5.5 at medium, 52.5% for $2.91, beats Fable 5.1 at max, 51.8% for $17.28
    • Stronger at low effort, but Opus 5.5 at medium beats that for less
    $10 / $50in / out
    per 1M tokens

Intelligence vs cost, at every effort level

Each line is one model. Each dot is one effort setting, from low to max. Up is smarter, left is cheaper, and the x-axis is logarithmic, so every gridline is a big jump in price.

Both labs launched with charts like these, and neither could plot the other's new model. So I rebuilt them from the sources that run every model the same way: Cognition's FrontierCode leaderboard, Zapier's AutomationBench, Cursor's CursorBench, Datacurve's DeepSWE board and Artificial Analysis. Where only a vendor chart exists, the badge on the chart says so. Click a legend entry to hide a model, hover a dot for its numbers, and click it to see the cheapest way each rival matches it.

Same method for every modelAgentic coding

FrontierCode 1.1

Score vs cost per task
2030405060$0.02$0.05$0.10$0.20$0.50$1$2$5$10Cost per task (USD, log scale)Score (%)lowmedhighmaxlowmedhighxhighmax

Hover a point for its score and cost. Click one to see what it costs per solved task and what each rival has to spend to match it. Click a model in the legend to hide it.

Cognition ran every model and priced every run the same way. Claude models ran in Claude Code, GPT models in Codex CLI. Costs are rounded to the cent by the leaderboard. Source: Cognition FrontierCode leaderboard.

View this chart as a table
ModelEffortScoreCost per taskOutput tokens
Claude Opus 5.5low47.3%$0.409.1k
Claude Opus 5.5medium54.6%$0.8018.5k
Claude Opus 5.5high54.0%$1.0926.1k
Claude Opus 5.5xhigh51.4%$2.2560.2k
Claude Opus 5.5max54.4%$6.19166k
GPT-6 Sollow37.3%$0.436.5k
GPT-6 Solmedium45.9%$0.7711.5k
GPT-6 Solhigh47.7%$1.0416.5k
GPT-6 Solxhigh48.4%$1.3222.8k
GPT-6 Solmax49.3%$2.0741.3k
GPT-6 Lunalow25.7%$0.027.3k
GPT-6 Lunamedium35.5%$0.0521.2k
GPT-6 Lunahigh37.3%$0.0629.0k
GPT-6 Lunaxhigh37.1%$0.0733.2k
GPT-6 Lunamax42.4%$0.1056.6k
Claude Fable 5.1low49.8%$2.3819.1k
Claude Fable 5.1medium50.9%$3.2826.1k
Claude Fable 5.1high50.3%$5.2739.9k
Claude Fable 5.1xhigh48.7%$9.2767.7k
Claude Fable 5.1max50.3%$12.8391.7k
GPT-6 Astralow45.3%$1.706.8k
GPT-6 Astramedium48.8%$2.4310.6k
GPT-6 Astrahigh50.9%$3.0114.3k
GPT-6 Astraxhigh50.6%$3.2817.0k
GPT-6 Astramax53.3%$4.5930.1k

What your budget buys

Curves are good for shape. For a decision, flip the question: if you can spend a fixed amount per task, which model gets the most done? Drag the slider. On most boards the leader changes as the budget grows, usually from Luna to Sol to Opus 5.5.

Pick a budget per task

Best published score at or under your cost per task
Benchmark
$1.00per task
  • Opus 5.554.6%med · $0.80
  • Sol45.9%med · $0.77
  • Luna42.4%max · $0.10
  • Astraneeds $1.70+
  • 5.6 Solneeds $1.89+
  • Fable 5.1needs $2.38+
  • Opus 5needs $2.68+

At $1.00 a task, Claude Opus 5.5 at medium effort gets the most done on FrontierCode 1.1. Source: Cognition FrontierCode leaderboard. Same method for every model

Coding: Opus 5.5 wins, and it isn't close on terminal work

Four coding benchmarks, four different angles: mergeable pull requests, command-line tasks, long engineering projects and real Cursor sessions.

FrontierCode: the best score is also cheap

FrontierCode grades real pull requests on whether a maintainer would merge them, so sprawling changes lose points. Opus 5.5 peaks at its default medium effort with 54.6% for $0.80 a task. That's the top score on the whole board, above GPT-6 Astra at max (53.3% for $4.59). Sol's best is 49.3% at max for $2.07.

FrontierCode 1.1 score and cost per taskCognition FrontierCode leaderboard
Opus 5.5med54.6%$0.80
Astramax53.3%$4.59
Solmax49.3%$2.07
Lunamax42.4%$0.10

More thinking doesn't help Opus here: xhigh and max both score below medium. My guess is the grader, which penalizes sprawling changes, and longer runs tend to produce bigger diffs. Luna is the surprise: 42.4% at max for $0.10 beats Sol at low (37.3% for $0.43) and GPT-5.6 Sol at medium (39.9% for $2.69).

Terminal-Bench 4.0: the biggest gap on the page

Artificial Analysis ran all seven models on Terminal-Bench 4.0 in one harness and priced the launch models at every effort. At the same $0.40 per task, Opus 5.5 at medium scores 52.5% and Sol at max scores 43.9%. Opus 5.5 at xhigh (59.6%) ties GPT-6 Astra at max (59.1%) for about the same money, and Luna tops out at 12.6%.

Terminal-Bench 4.0 score and cost per taskArtificial Analysis
Opus 5.5xhigh59.6%$0.88
Astramax59.1%$0.85
Opus 5.5med52.5%$0.40
Solmax43.9%$0.40
Lunamax12.6%$0.024

Anthropic's own chart shows higher scores and much higher costs for the same model, because it uses a different harness: 66.4% at xhigh for $7.35. Both sources agree on the shape. Opus 5.5 gains nothing from xhigh to max, and Sol is the one model where max is worth it on this test: it climbs from 30.3% at xhigh to 43.9%.

DeepSWE: only OpenAI's chart has Sol and Luna

DeepSWE is 113 long engineering tasks written from scratch. There is no version 1.3. The latest is v1.1, and Datacurve's public board hadn't run any of the September 22 models when I checked. The only Sol and Luna numbers come from OpenAI's launch chart: Sol at max scores 68.8% for $2.74, and Luna at max scores 66.6% for just $0.22. Anthropic's system card lists Opus 5.5 at max at 74.2% with no cost, in line with GPT-6 Astra's best of 74.1%.

Treat Luna's DeepSWE number with care. OpenAI ran its own models itself and copied Claude numbers from Datacurve, and on GPT-5.6 Luna the two sources don't even agree on the score. Until Datacurve runs Luna, it's one vendor's claim.

CursorBench: Cursor hasn't run GPT-6 yet

Cursor's own board had Opus 5.5 at every effort and no GPT-6 model of any size. Opus 5.5 at medium (52.5% for $2.91) beats Claude Fable 5.1 at max (51.8% for $17.28). Even Opus 5.5 at low (43.7% for $1.17) beats GPT-5.6 Sol at max (41.7% for $8.23). Unlike FrontierCode, effort keeps paying inside Cursor, up to 57.8% at max.

If you want to try these side by side in your own editor, Claude Code can point at any provider. I wrote up the two-line switch in how to change the model in Claude Code.

Agents and knowledge work: Sol earns its keep here

Business automations, office deliverables, computer use and long professional tasks. The picture is closer, and it's where Sol and Astra look best.

Same method for every modelBusiness workflows across apps

AutomationBench 1.0.6

Score vs cost per task
01020304050$0.01$0.02$0.05$0.10$0.20$0.50$1$2Cost per task (USD, log scale)Score (%)offlowmedhighxhighlowmedhighxhighmax

Hover a point for its score and cost. Click one to see what it costs per solved task and what each rival has to spend to match it. Click a model in the legend to hide it.

Zapier ran and priced every model the same way. The unconnected Fable 5.1 point used Opus 5 as a fallback on about 40% of tasks, and its cost excludes the fallback tokens. Source: Zapier AutomationBench leaderboard.

View this chart as a table
ModelEffortScoreCost per task
Claude Opus 5.5low23.3%$0.46
Claude Opus 5.5medium28.6%$0.60
Claude Opus 5.5high32.0%$0.65
Claude Opus 5.5xhigh34.4%$0.80
Claude Opus 5.5max40.0%$1.28
GPT-6 Solnone9.1%$0.25
GPT-6 Sollow21.2%$0.19
GPT-6 Solmedium26.9%$0.21
GPT-6 Solhigh31.2%$0.24
GPT-6 Solxhigh33.2%$0.27
GPT-6 Solmax32.0%$0.34
GPT-6 Lunanone0.3%$0.01
GPT-6 Lunalow1.2%$0.01
GPT-6 Lunamedium9.4%$0.02
GPT-6 Lunahigh14.5%$0.02
GPT-6 Lunaxhigh12.6%$0.02
GPT-6 Lunamax20.7%$0.04
Claude Fable 5.1xhigh21.3%$2.14
Claude Fable 5.1max (with fallback)31.4%$2.45
Claude Fable 5.1max22.4%$2.45
GPT-6 Astranone25.3%$1.28
GPT-6 Astralow30.3%$1.08
GPT-6 Astramedium34.1%$1.28
GPT-6 Astrahigh37.1%$1.45
GPT-6 Astraxhigh39.0%$1.53
GPT-6 Astramax41.4%$1.73

AutomationBench: Sol is the value pick

Zapier's AutomationBench runs end-to-end workflows across 47 business apps and only counts the final state. At high effort, Sol scores 31.2% for $0.24 and Opus 5.5 scores 32.0% for $0.65: the same result for 37% of the price. Sol stops climbing at xhigh (33.2%), and its max is worse.

AutomationBench 1.0.6 score and cost per taskZapier AutomationBench leaderboard
Astramax41.4%$1.73
Opus 5.5max40.0%$1.28
Solxhigh33.2%$0.27
Opus 5.5high32.0%$0.65
Solhigh31.2%$0.24
Lunamax20.7%$0.04

The top of the board belongs to the expensive settings. GPT-6 Astra at max leads with 41.4% for $1.73, and Opus 5.5 at max is right behind with 40.0% for $1.28. This is one of the few tests where Opus 5.5's max effort clearly earns its cost. Research math is the other.

Docs, decks and spreadsheets

On GDPval-AA, where Artificial Analysis judges real work products head to head, Opus 5.5 at medium (1576 Elo for $0.86) beats GPT-6 Astra at max (1542 Elo for $4.53). Artificial Analysis puts Sol at max at 1487 Elo and Luna at 1367 Elo. If your output is a document someone has to read, Opus 5.5 is the pick.

Computer use: no fair head-to-head exists

OpenAI and Anthropic ran different versions of OSWorld 2.0. Sol scores 64.4% at max in OpenAI's version, and Opus 5.5 scores 81.8% in Anthropic's. The one model both ran, Claude Opus 5, scored 70.2% in one and 74.0% in the other. The setups differ by a few points; the gap between Sol and Opus 5.5 is much larger than that. My read is that Opus 5.5 leads, but that's an inference, not a measurement.

Long professional tasks and factual errors

Two OpenAI-only charts round it out. On Agents' Last Exam, Sol at xhigh (55.4% for $1.67) matches Claude Opus 5 at high (55.9% for $7.29) for 23% of the cost. Opus 5.5 isn't on it. And on OpenAI's hard-prompt factuality test, Sol at max makes factual errors on 4.6% of prompts, down from 8.5% for GPT-5.6 Sol. That's the clearest real improvement in the Sol launch.

On research math, Anthropic's system card has Opus 5.5 at 91.2% on ArXivMath at max, against Fable 5.1's 82.9%, and at 67.7% on Humanity's Last Exam with tools, against 65.6%. No GPT-6 model was run on either.

Which model should you use?

Pick the job and what matters most. Each answer shows the numbers behind it, from a single source per row.

Pick a job, get a model

Claude Opus 5.5medium effort

Unusual result: the best score is also one of the cheapest. Opus 5.5 at medium beats every GPT-6 Sol setting on FrontierCode and costs less than Sol at high.

  • Opus 5.5 mediumFrontierCode 1.1 · Cognition FrontierCode leaderboard54.6% · $0.80
  • Sol highFrontierCode 1.1 · Cognition FrontierCode leaderboard47.7% · $1.04
  • Sol maxFrontierCode 1.1 · Cognition FrontierCode leaderboard49.3% · $2.07

My recommendations, read off the published numbers. Scores and costs come from one source per row; I never mix two sources in one comparison.

The pattern across all of it: Opus 5.5 when the task is ambiguous or long, Sol when it's well-specified and repeated, Luna when it's simple and huge. Effort matters as much as the model. Opus 5.5 at medium beats most models at max, and Sol at max is usually money you don't need to spend.

GPT-6 Astra vs Sol vs Luna

Same family, three price points. Sol costs a fifth of Astra per token, and Luna costs a twentieth of Sol.

GPT-6 Luna vs Sol vs Astra, with GPT-5.6 Sol for reference

BenchmarkLunaSolAstra5.6 Sol
Price per 1M tokensinput / output$0.10 / $0.50cached $0.01$2 / $10cached $0.20$10 / $50cached $1$4 / $20cached $0.40
Intelligence IndexArtificial Analysis37.3max · $0.06847.5max · $1.0652.7max · $3.2647.0max · $1.99
FrontierCode 1.1Cognition FrontierCode leaderboard42.4%max · $0.1049.3%max · $2.0753.3%max · $4.5947.5%max · $5.19
Terminal-Bench 4.0Artificial Analysis12.6%max · $0.02443.9%max · $0.4059.6%xhigh · cost n/a39.9%max · $0.81
AutomationBench 1.0.6Zapier AutomationBench leaderboard20.7%max · $0.0433.2%xhigh · $0.2741.4%max · $1.7328.8%max · $0.67
DeepSWE 1.1OpenAI launch post66.6%max · $0.2268.8%max · $2.7474.1%xhigh · $4.4372.7%max · $6.46
OSWorld 2.0OpenAI launch post52.7%max · $0.2764.4%max · $3.2573.5%max · $9.0766.2%max · $7.71
Agents' Last ExamOpenAI launch post50.9%max · $0.1556.4%max · $2.9359.3%max · $6.2353.6%xhigh · $5.08
Factual errors (lower is better)OpenAI launch post7.6%max · $0.0124.5%xhigh · $0.133.9%high · $0.488.4%xhigh · $0.39

Underlined: the row leader. Green dot: one neutral party ran every model the same way. Amber dot: a vendor's own chart, so compare within the row, not across rows. Terminal-Bench costs are only published for the launch models and for max effort on the rest. * A result from a different source, shown for reference.

Luna is the model to put under something bigger. Use it for routing, extraction, first drafts and sub-agent reading, always at max effort, which is its best setting on every chart I pulled and still costs cents.

Sol is OpenAI's workhorse. It doesn't beat GPT-5.6 Sol on intelligence (within 0.6 points at every effort), but it costs 45% to 53% as much per task. Default to high or xhigh. If you already run GPT-5.6 Sol in an agent, like my Hermes Agent setup with no API bill, Sol gets you the same scores for about half the cost per task.

Astra scores higher than Sol on every OpenAI chart, by the widest margin on Terminal-Bench (59.1% vs 43.9%) and AutomationBench. Pay for it on long, multi-app runs and computer use, where a failed attempt costs more than the tokens. I covered its launch in GPT-6 Astra release date, price and benchmarks.

Claude Fable 5.1 vs Opus 5.5

Opus 5.5's best beats Fable 5.1's best on 9 of 9 benchmarks, at 40% of the per-token price.

Claude Opus 5.5 vs Fable 5.1, with Opus 5 for reference

BenchmarkOpus 5.5Fable 5.1Opus 5
Price per 1M tokensinput / output$4 / $20cached $0.20$10 / $50cached $0.25$5 / $25cached $0.50
Intelligence IndexArtificial Analysis57.6max · $5.9853.4max · $7.6350.8max · $5.86
FrontierCode 1.1Cognition FrontierCode leaderboard54.6%med · $0.8050.9%med · $3.2853.4%med · $4.31
Terminal-Bench 4.0Artificial Analysis59.6%xhigh · $0.8855.1%xhigh · cost n/a49.0%max · $1.93
CursorBench 4.0Cursor CursorBench 4.0 leaderboard57.8%max · $13.4351.8%max · $17.2846.6%max · $11.95
AutomationBench 1.0.6Zapier AutomationBench leaderboard40.0%max · $1.2822.4%max · $2.4526.9%max · $3.05
GDPval-AA v2.1Anthropic launch post1846 Elomax · $8.921735 Elomax · $9.591708 Elomax · $6.76
OSWorld 2.0Claude Opus 5.5 system card81.8%max · $8.4680.7%max · $12.4074.4%xhigh · $10.60
Humanity's Last ExamClaude Opus 5.5 system card67.7%max · $2.1165.6%max · $3.4063.6%max · $2.47
ArXivMath (Aug 2026)Claude Opus 5.5 system card91.2%max · $3.9482.9%max · $12.9078.1%max · $7.49

Underlined: the row leader. Green dot: one neutral party ran every model the same way. Amber dot: a vendor's own chart, so compare within the row, not across rows. Fable 5.1 was only run at xhigh and max on AutomationBench. * A result from a different source, shown for reference.

There's one real difference: Fable 5.1 does more with less thinking. At low effort it beats Opus 5.5 at low on 8 of 8 shared charts. But Opus 5.5 at medium beats Fable 5.1 at low on 8 of 8, and it's cheaper on all 7 charts that publish a cost. Example: on the Intelligence Index, Fable 5.1 at low scores 46.8 for $2.37, and Opus 5.5 at medium scores 51.2 for $1.34.

So the practical advice is to move Fable 5.1 workloads to Opus 5.5 and re-run your own evals. If something you care about gets worse, keep Fable for that one job. The upgrade from Opus 5 is an easier call: Opus 5.5 scores higher on every benchmark both were run on, at 20% lower token prices.

How they work together

The cheapest way to use frontier models is to use several. Route each step to the cheapest model that can do it, and escalate what fails.

A four-tier routing ladder

Cost range: Intelligence Index task, low to max effort

Send it here

  • Multi-file changes in a real codebase, where scope discipline matters
  • Terminal and infrastructure work with long tool loops
  • Docs, decks and spreadsheets people will actually read
  • Planning the job and reviewing what the cheaper models produced

Why

  • Opus 5.5 mediumFrontierCode54.6% · $0.80
  • Opus 5.5 mediumTerminal-Bench52.5% · $0.40
  • Opus 5.5 mediumGDPval-AA1576 Elo · $0.86
Example flow Luna triages Opus 5.5 plans Sol executes each step Opus 5.5 reviews Astra takes the failures

Plan with Opus, build with Sol

The planner runs once and the executor runs many times. Let Opus 5.5 at medium break the job into specs, then hand each step to Sol at high. You pay Opus prices on a fraction of the tokens.

Luna as the sub-agent

Reading files, searching, and summarizing logs eat most of an agent's tokens. Give that work to Luna at max and send only the summary up to the expensive model.

Escalate on failure

For workflows with a clear pass or fail, start with Sol at xhigh. Send only the failures to Opus 5.5 or Astra at max. You get close to top-model reliability at mostly Sol prices.

These are my routing suggestions from the numbers. I found no launch-day report of anyone running this exact combination yet. It is the same idea as the three-tier model cascade I use to run Hermes Agent for $8 a month. If you want one API key across labs, OpenRouter listed GPT-6 Sol and Luna on launch day: what OpenRouter is and how it works.

What it actually costs

Per token, Opus 5.5 costs twice what Sol does. Per task, the gap is often smaller, and sometimes Opus is cheaper.

Two things close the gap. First, cache reads cost the same $0.20 per million tokens on Opus 5.5 and Sol, so long agent loops that mostly re-read context get much closer than 2x. Second, models use different amounts of tokens. On FrontierCode, Opus 5.5 at low costs $0.40 a task and Sol at low costs $0.43, despite the 2x price per token. And Opus 5.5 at medium beats Sol at max for 39% of the cost.

Price a workload

List prices per million tokens, standard tier

On this mix Opus 5.5 costs 1.69x GPT-6 Sol per call. Cache reads cost the same $0.20 per million on both, so heavy caching narrows the gap from 2x.

  • Luna$116/mo
  • Sol$2,320/mo
  • Opus 5.5$3,920/mo
  • 5.6 Sol$4,640/mo
  • Opus 5$5,800/mo
  • Fable 5.1$8,900/mo
  • Astra$11,600/mo

Same token counts for every model, which flatters wordy models. In the benchmarks, models spend very different amounts of tokens on the same task: on FrontierCode, Opus 5.5 at medium writes about 18.5k output tokens per task and GPT-6 Sol at medium about 11.5k. Use the curves above for measured cost per task, and this for list-price math.

Opus 5.5 is also cheaper than the model it replaces: $4 and $20 per million tokens against Opus 5's $5 and $25, and Sol is half of GPT-5.6 Sol's $4 and $20. Luna is 20 times cheaper than Sol on every token type. If you want a cheaper route for smart-enough coding work, I tested one in MiniMax M3 in Claude Code.

The full scoreboard

Every benchmark I could find for these models, in one grid. Toggle to independent sources only to see the fair fights.

Every benchmark, one grid

Benchmark and sourceOpus 5.5SolLunaAstraFable 5.1
Intelligence IndexArtificial Analysis57.6max · $5.9847.5max · $1.0637.3max · $0.06852.7max · $3.2653.4max · $7.63
FrontierCode 1.1Cognition FrontierCode leaderboard54.6%med · $0.8049.3%max · $2.0742.4%max · $0.1053.3%max · $4.5950.9%med · $3.28
Terminal-Bench 4.0Artificial Analysis59.6%xhigh · $0.8843.9%max · $0.4012.6%max · $0.02459.6%xhigh · no cost55.1%xhigh · no cost
AutomationBench 1.0.6Zapier AutomationBench leaderboard40.0%max · $1.2833.2%xhigh · $0.2720.7%max · $0.0441.4%max · $1.7322.4%max · $2.45
CursorBench 4.0Cursor CursorBench 4.0 leaderboard57.8%max · $13.43not runnot runnot run51.8%max · $17.28
DeepSWE 1.1OpenAI's chart, with cost74.2%*other source68.8%max · $2.7466.6%max · $0.2274.1%xhigh · $4.43not run
GDPval-AA v2.1Anthropic's chart, with cost1846 Elomax · $8.921487 Elo*other source1367 Elo*other source1542 Elomax · $4.531735 Elomax · $9.59
OSWorld 2.0OpenAI's run (offline set)not run64.4%max · $3.2552.7%max · $0.2773.5%max · $9.07not run
OSWorld 2.0Anthropic's run (Sept 10 files)81.8%max · $8.46not runnot runnot run80.7%max · $12.40
Agents' Last ExamOpenAI's chart, with costnot run56.4%max · $2.9350.9%max · $0.1559.3%max · $6.23not run
Factual errors, lower is betterOpenAI's chartnot run4.5%xhigh · $0.137.6%max · $0.0123.9%high · $0.48not run
Humanity's Last ExamAnthropic system card67.7%max · $2.11not runnot runnot run65.6%max · $3.40
ArXivMath (Aug 2026)Anthropic system card, no tools91.2%max · $3.94not runnot runnot run82.9%max · $12.90
Rows led8 of 90 of 80 of 86 of 90 of 9

"Rows led" counts only rows the model was actually run on, and only against the models shown. Vendor charts (amber dot) mostly leave out the rival lab's newest model, which is why so many cells say "not run". * From a different source than the row, for reference only.

What developers are saying

Launch-day reactions from people who used the models, paraphrased and linked. Tags flag who used both, and who has a stake.

  • Dan Shipper, EveryX and Every's Vibe Check

    Switched to GPT-6 Sol as the daily driver in Codex: not quite Astra, but close enough for most everyday work, faster, and half the price of GPT-5.6 Sol. Every's public summary of the head-to-head with Opus 5.5 calls Sol the better step-by-step collaborator and gives Opus 5.5 the higher ceiling for long autonomous builds.

    SolOpus 5.5Used bothFull piece paywalled
    Source
  • Katie Parrott, EveryEvery's Vibe Check

    Reported that Every staffers who had moved to Codex are drifting back to Claude, because Opus 5.5 gives Fable-level output for much less and fixed some of the personality issues that pushed them away from Opus 5.

    Opus 5.5Used bothTwo people at one company
    Source
  • Simon WillisonHacker News

    Ran the pelican-on-a-bicycle SVG test on Opus 5.5 at max effort twice. Both times the model spent the whole 128k-token budget reasoning and never produced the drawing.

    Opus 5.5Independent testerReproducible
    Source
  • Several Hacker News commentersGPT-6 Sol and Luna launch thread

    The loudest theme in the thread: Sol and Luna read as a price cut, not a capability jump. One commenter said they'd rather have paid double for a real performance gain. Another said they get better value from Opus 5.5.

    SolLunaTop comments only
    Source
  • dom96 and jadboxHacker News

    Cross-vendor skeptics. One said the open-weight MiMo V2.6 Pro beat Sol on quality and price in their own benchmark. The other said Gemini 3.8 Flash beats Sol and Luna on price and coding in their use.

    SolLunaAnecdotal
    Source
  • Artificial AnalysisX

    Found that Sol and Luna roughly halve the cost of GPT-5.6 Sol and Luna while their Intelligence Index and Coding Agent Index scores stay about level: better on some evals, worse on others.

    SolLunaIndependent benchmarker
    Source
  • sidewndr46Hacker News

    Said Opus 5 at max kept failing their real tasks by hitting tool-call limits, while high effort got the work done. Offered as a caution for Opus 5.5 at max too.

    Opus 5About the previous Opus
    Source
  • Alex McFarland, Unite.AIBuyer's guide

    Recommends Opus 5.5 for ambiguous, multi-step work where mistakes are expensive, like codebase migrations, and Sol for clear, bounded, high-volume work that's easy to check.

    Opus 5.5SolAnalysis, not hands-on
    Source
  • the-decoderNews

    Early testers found Opus 5.5's writing clearer and quicker to get to the point. Anthropic marketed exactly this as a fix for 'Claudish' prose.

    Opus 5.5Filtered through Anthropic's messaging
    Source
  • tedsandersHacker News

    Pointed out that Opus 5.5 at high and GPT-6 Astra at high cost about the same per task on Artificial Analysis, and pushed back on claims that OpenAI's API had degraded. The commenter works at OpenAI.

    Opus 5.5AstraOpenAI employee
    Source
  • Anthropic docs and The New StackPlatform docs, trade press

    Anthropic documents that most flagged cybersecurity requests to Opus 5.5 get routed to Opus 4.8, and a new biology classifier joins the cyber one. The New Stack ran the headline that your agent calls might quietly land on an older model.

    Opus 5.5Documented behavior
    Source
  • GitHub, OpenRouter, VercelPlatform rollouts

    Opus 5.5 shipped in GitHub Copilot on day one. OpenRouter and Vercel's AI Gateway listed all three launches the same day. Availability, not a verdict.

    Opus 5.5SolLunaAdoption signal
    Source

My read on the first-use signal

It's launch day, so treat all of this as early. I looked for people who used more than one of these models, then weighed each theme by how many independent sources back it and whether the benchmarks agree.

SignalStrengthWhat it rests onWhat to do with it
Sol is GPT-5.6 Sol at half the priceStrongArtificial Analysis's numbers, the loudest HN theme, and a tech-press analysis all point the same way.If you already run GPT-5.6 Sol, switch for the bill. Don't expect new capability.
Opus 5.5 has the higher ceiling on hard, long workSolidIndependent benchmarks agree. Hands-on reports are few but point the same direction (Every, Unite.AI).Default to Opus 5.5 for repo-scale and terminal work. Re-check in two weeks when more people have shipped with it.
Opus 5.5 at max effort often wastes moneySolidA reproducible failure from Simon Willison, matching HN advice, and flat or falling curves from xhigh to max on FrontierCode and Terminal-Bench.Run medium by default and xhigh for hard tickets. Save max for multi-app automations, where the curve still climbs.
Luna is the volume tier, not a rivalSolidNobody in the sample pitched Luna against the flagships. The cost data backs the framing.Use it under a bigger model, not instead of one.
Coders are drifting back from Codex to ClaudeThinTwo named people at one company, in one outlet's piece.A story worth watching, not a trend yet.
Cheaper rivals beat Sol at codingThinTwo anecdotal HN comments naming MiMo V2.6 Pro and Gemini 3.8 Flash.If price is the whole reason you'd pick Sol, test those two as well.
Flagged security prompts quietly change modelsDocumentedAnthropic's own docs describe the reroute to Opus 4.8. How often it happens in normal use is unknown.If you build security tooling on Opus 5.5, log which model actually answered.

The gaps matter too. I couldn't read Reddit threads, found no launch-day YouTube first impressions, and only saw the top comments in each large Hacker News thread. Among the few people who tried both, the pattern is consistent: Opus 5.5 for serious building, Sol as the cheaper daily driver for well-defined work.

Where the vendor charts disagree

9 cases where the same model at the same effort gets different numbers from two sources. This is why every chart above sticks to one source at a time.

  • Terminal-Bench 4.0

    Anthropic's own run puts Opus 5.5 6.8 points higher than Artificial Analysis does, at 8.4x the cost per task.

    Different harness, time limits and trial counts. Compare models inside one source, never across them.

    Opus 5.5xhigh
    Anthropic launch chart
    66.4%$7.35
    Opus 5.5xhigh
    Artificial Analysis
    59.6%$0.88
    Opus 5.5max
    Anthropic launch chart
    64.8%$11.24
    Opus 5.5max
    Artificial Analysis
    59.6%$1.31
  • Terminal-Bench 4.0

    One model, three published Terminal-Bench scores: 53.9%, 52.3% and 49.0%.

    Anthropic's page cites the public board at 51.8% for Opus 5; the live board showed 53.9% when I checked. Leaderboards move.

    Opus 5xhigh
    tbench.ai public board (live)
    53.9%no cost
    Opus 5max
    Anthropic launch chart
    52.3%$15.83
    Opus 5max
    Artificial Analysis
    49.0%$1.93
  • DeepSWE 1.1

    OpenAI's chart copies Datacurve's GPT-5.6 Sol scores exactly but lists each run 23% to 24% cheaper.

    Same scores, different cost math. Datacurve reports cost per scored attempt; OpenAI does not say how it re-priced the runs.

    5.6 Solmax
    OpenAI launch chart
    72.7%$6.46
    5.6 Solmax
    Datacurve leaderboard
    72.7%$8.39
  • DeepSWE 1.1

    For GPT-5.6 Luna the two sources don't even agree on the score: 62.2% on OpenAI's chart, 67.2% on Datacurve's board, at 5.7x the cost.

    OpenAI ran its own models itself and took Claude numbers from Datacurve. Luna's cheap DeepSWE result has no independent run yet.

    5.6 Lunamax
    OpenAI launch chart
    62.2%$0.53
    5.6 Lunamax
    Datacurve leaderboard
    67.2%$3.03
  • OSWorld 2.0

    OpenAI and Anthropic ran different versions of OSWorld 2.0. Claude Opus 5 at max scores 70.2% in one and 74.0% in the other, at $24.11 vs about $11.60 a task.

    GPT-6 Sol's 64.4% and Opus 5.5's 81.8% come from these two different tests. They are not a head-to-head.

    Opus 5max
    OpenAI launch chart (offline set)
    70.2%$24.11
    Opus 5max
    Anthropic system card (Sept 10 files)
    74.0%$11.60
  • AutomationBench 1.0.6

    OpenAI's chart prices Claude Opus 5 at 2.2x to 2.5x what Zapier's own leaderboard lists, low through xhigh effort.

    Anthropic's chart matches Zapier here. When a vendor plots a rival, check the rival's cost against the neutral board.

    Opus 5med
    OpenAI launch chart
    23.9%$2.22
    Opus 5med
    Zapier leaderboard
    23.9%$0.89
  • AutomationBench 1.0.6

    Anthropic lists its own Opus 5.5 max run at $1.37 a task; Zapier's board says $1.28.

    Small, and in the unflattering direction for Anthropic. Not every gap favors the vendor.

    Opus 5.5max
    Anthropic launch chart
    40.0%$1.37
    Opus 5.5max
    Zapier leaderboard
    40.0%$1.28
  • FrontierCode 1.1

    Anthropic's FrontierCode chart shows Fable 5.1 at low effort scoring 52.8%. Cognition's leaderboard says 49.8%.

    Every other Claude point on Anthropic's chart matches Cognition. This one doesn't, and nothing on the page explains it.

    Fable 5.1low
    Anthropic launch chart
    52.8%$2.47
    Fable 5.1low
    Cognition leaderboard
    49.8%$2.38
  • FrontierCode 1.1

    OpenAI prices its own GPT-6 Sol runs 3% to 5% above Cognition's figures.

    Cognition rounds to the cent, which explains part of it. The direction is against OpenAI, so this is not flattering math.

    Solmax
    OpenAI launch chart
    49.3%$2.14
    Solmax
    Cognition leaderboard
    49.3%$2.07

FAQ

Is Claude Opus 5.5 better than GPT-6 Sol?

On most hard work, yes. On the independent leaderboards, Opus 5.5 at medium effort scores 54.6% on FrontierCode for $0.80 a task, above every GPT-6 Sol setting, and 52.5% on Artificial Analysis's Terminal-Bench run for the same $0.40 that Sol at max needs to reach 43.9%. Sol is the better value on business automations, where it roughly matches Opus 5.5 for about 40% of the cost.

Is GPT-6 Sol better than GPT-5.6 Sol?

Not smarter, but much cheaper. On the Artificial Analysis Intelligence Index, GPT-6 Sol lands within 0.6 points of GPT-5.6 Sol at every effort level for 45% to 53% of the cost per task. It also makes about half as many factual errors on OpenAI's own hard-prompt test.

What is GPT-6 Luna good for?

High-volume, simple work: triage, extraction, tagging, summaries and sub-agent reading. Luna at max scores 37.3 on the Intelligence Index for $0.068 a task, the cheapest model on that board. It struggles with long tool loops: 12.6% on Terminal-Bench.

Should I use GPT-6 Astra or GPT-6 Sol?

Sol for most work, Astra for escalation. Astra at max leads AutomationBench at 41.4% against Sol's best of 33.2%, and it scores higher on every OpenAI chart. It also costs five times as much per token, so send it the long multi-app workflows and the tasks that already failed once.

Is Claude Opus 5.5 better than Claude Fable 5.1?

On every benchmark both were run on (9 of 9), Opus 5.5's best score beats Fable 5.1's best, at 40% of Fable's per-token price. Fable 5.1 does score higher at low effort, but Opus 5.5 at medium beats Fable 5.1 at low on all 8 shared charts and costs less on all 7 that publish a cost.

What effort level should I use for Claude Opus 5.5?

Medium, the default. It is Opus 5.5's best FrontierCode setting and gets most of the Terminal-Bench score. Use xhigh for hard terminal and infrastructure work. Use max only for long business workflows across apps, where AutomationBench still climbs from 34.4% at xhigh to 40.0% at max.

What effort level should I use for GPT-6 Sol?

High or xhigh. Max rarely pays off: on AutomationBench it scores lower than xhigh, and on FrontierCode it adds under a point for about 1.6x the cost. The exception is terminal work, where Sol jumps from 30.3% at xhigh to 43.9% at max.

How much do Claude Opus 5.5, GPT-6 Sol and GPT-6 Luna cost?

Per million tokens: Claude Opus 5.5 is $4 input and $20 output, GPT-6 Sol is $2 and $10, and GPT-6 Luna is $0.10 and $0.50. Cached input costs $0.20 on both Opus 5.5 and Sol, and $0.01 on Luna. GPT-6 Astra and Claude Fable 5.1 are both $10 and $50.

Can I compare OpenAI's and Anthropic's launch charts directly?

No. Neither vendor plotted the other's new model, and their cost math differs from the neutral leaderboards. For example, Anthropic's own Terminal-Bench run puts Opus 5.5 at xhigh at 66.4%, while Artificial Analysis measures 59.6%. Use leaderboards that run every model the same way, like Cognition's FrontierCode, Zapier's AutomationBench and Artificial Analysis.

Is there a DeepSWE 1.3 or a GPT-6 Sol run on the public leaderboards?

Not yet. DeepSWE's latest version is v1.1, and its board had not run any of the September 22 launches when I captured it. The public Terminal-Bench 4.0 and OSWorld 2.0 boards hadn't added them either. On launch day, all three new models were on Cognition's FrontierCode, Zapier's AutomationBench and Artificial Analysis. Cursor's CursorBench had Opus 5.5 only.

How I built this

Every number on this page comes from a public source captured on September 22, 2026. Nothing is estimated unless the chart says so.

Launch-page charts were read from their embedded chart data and checked by hovering each point. A few system-card charts had no embedded data, so I read those off the image, accurate to about half a percent on cost. I re-checked the independent boards around 4 PM Pacific to pick up late additions. Each chart carries its source badge, and no chart mixes two sources. You can download all 471 data points as a CSV.

Moe Lueker builds AI systems for creators and small businesses and tests new models on the channel. Mechanical engineer, then venture capital. More about Moe. If you'd rather talk to these models than type, try Rambleproof, my free Mac dictation app.

Key: Opus 5.5 Fable 5.1 Astra Sol Luna. Hollow markers on dashed lines are the previous generation.