GPT-6 Astra Release Date, Price, and Benchmarks (September 2026)
OpenAI released GPT-6 Astra on September 3, 2026. The release date, the API price, the benchmarks OpenAI buried, and the line in the system card nobody is quoting.
Anthropic and OpenAI shipped three models on September 22. I pulled every chart from both launches and five independent leaderboards, then plotted score against cost at every effort level. Here's where each model wins, what it costs, and how to use them together.
Opus 5.5 is the best model of the three and, per task, often the cheapest way to get a good result. Sol is GPT-5.6 Sol at half the price. Luna is the pennies tier. Most teams should run all three.
Your default for hard work: real codebases, terminal and infrastructure, documents, anything ambiguous. Run it at medium.
Well-specified work at volume: business automations, tickets with tests, anything you ran on GPT-5.6 Sol. Run it at high or xhigh.
Triage, extraction, summaries and sub-agent grunt work. Always run it at max: it's still pennies.
The escalation model for long workflows across many apps and for computer use.
Opus 5.5 beat it on every benchmark both were run on. Keep it only where your own tests say otherwise.
Each line is one model. Each dot is one effort setting, from low to max. Up is smarter, left is cheaper, and the x-axis is logarithmic, so every gridline is a big jump in price.
Both labs launched with charts like these, and neither could plot the other's new model. So I rebuilt them from the sources that run every model the same way: Cognition's FrontierCode leaderboard, Zapier's AutomationBench, Cursor's CursorBench, Datacurve's DeepSWE board and Artificial Analysis. Where only a vendor chart exists, the badge on the chart says so. Click a legend entry to hide a model, hover a dot for its numbers, and click it to see the cheapest way each rival matches it.
FrontierCode 1.1
Hover a point for its score and cost. Click one to see what it costs per solved task and what each rival has to spend to match it. Click a model in the legend to hide it.
Cognition ran every model and priced every run the same way. Claude models ran in Claude Code, GPT models in Codex CLI. Costs are rounded to the cent by the leaderboard. Source: Cognition FrontierCode leaderboard.
| Model | Effort | Score | Cost per task | Output tokens |
|---|---|---|---|---|
| Claude Opus 5.5 | low | 47.3% | $0.40 | 9.1k |
| Claude Opus 5.5 | medium | 54.6% | $0.80 | 18.5k |
| Claude Opus 5.5 | high | 54.0% | $1.09 | 26.1k |
| Claude Opus 5.5 | xhigh | 51.4% | $2.25 | 60.2k |
| Claude Opus 5.5 | max | 54.4% | $6.19 | 166k |
| GPT-6 Sol | low | 37.3% | $0.43 | 6.5k |
| GPT-6 Sol | medium | 45.9% | $0.77 | 11.5k |
| GPT-6 Sol | high | 47.7% | $1.04 | 16.5k |
| GPT-6 Sol | xhigh | 48.4% | $1.32 | 22.8k |
| GPT-6 Sol | max | 49.3% | $2.07 | 41.3k |
| GPT-6 Luna | low | 25.7% | $0.02 | 7.3k |
| GPT-6 Luna | medium | 35.5% | $0.05 | 21.2k |
| GPT-6 Luna | high | 37.3% | $0.06 | 29.0k |
| GPT-6 Luna | xhigh | 37.1% | $0.07 | 33.2k |
| GPT-6 Luna | max | 42.4% | $0.10 | 56.6k |
| Claude Fable 5.1 | low | 49.8% | $2.38 | 19.1k |
| Claude Fable 5.1 | medium | 50.9% | $3.28 | 26.1k |
| Claude Fable 5.1 | high | 50.3% | $5.27 | 39.9k |
| Claude Fable 5.1 | xhigh | 48.7% | $9.27 | 67.7k |
| Claude Fable 5.1 | max | 50.3% | $12.83 | 91.7k |
| GPT-6 Astra | low | 45.3% | $1.70 | 6.8k |
| GPT-6 Astra | medium | 48.8% | $2.43 | 10.6k |
| GPT-6 Astra | high | 50.9% | $3.01 | 14.3k |
| GPT-6 Astra | xhigh | 50.6% | $3.28 | 17.0k |
| GPT-6 Astra | max | 53.3% | $4.59 | 30.1k |
Curves are good for shape. For a decision, flip the question: if you can spend a fixed amount per task, which model gets the most done? Drag the slider. On most boards the leader changes as the budget grows, usually from Luna to Sol to Opus 5.5.
Four coding benchmarks, four different angles: mergeable pull requests, command-line tasks, long engineering projects and real Cursor sessions.
FrontierCode grades real pull requests on whether a maintainer would merge them, so sprawling changes lose points. Opus 5.5 peaks at its default medium effort with 54.6% for $0.80 a task. That's the top score on the whole board, above GPT-6 Astra at max (53.3% for $4.59). Sol's best is 49.3% at max for $2.07.
More thinking doesn't help Opus here: xhigh and max both score below medium. My guess is the grader, which penalizes sprawling changes, and longer runs tend to produce bigger diffs. Luna is the surprise: 42.4% at max for $0.10 beats Sol at low (37.3% for $0.43) and GPT-5.6 Sol at medium (39.9% for $2.69).
Artificial Analysis ran all seven models on Terminal-Bench 4.0 in one harness and priced the launch models at every effort. At the same $0.40 per task, Opus 5.5 at medium scores 52.5% and Sol at max scores 43.9%. Opus 5.5 at xhigh (59.6%) ties GPT-6 Astra at max (59.1%) for about the same money, and Luna tops out at 12.6%.
Anthropic's own chart shows higher scores and much higher costs for the same model, because it uses a different harness: 66.4% at xhigh for $7.35. Both sources agree on the shape. Opus 5.5 gains nothing from xhigh to max, and Sol is the one model where max is worth it on this test: it climbs from 30.3% at xhigh to 43.9%.
DeepSWE is 113 long engineering tasks written from scratch. There is no version 1.3. The latest is v1.1, and Datacurve's public board hadn't run any of the September 22 models when I checked. The only Sol and Luna numbers come from OpenAI's launch chart: Sol at max scores 68.8% for $2.74, and Luna at max scores 66.6% for just $0.22. Anthropic's system card lists Opus 5.5 at max at 74.2% with no cost, in line with GPT-6 Astra's best of 74.1%.
Treat Luna's DeepSWE number with care. OpenAI ran its own models itself and copied Claude numbers from Datacurve, and on GPT-5.6 Luna the two sources don't even agree on the score. Until Datacurve runs Luna, it's one vendor's claim.
Cursor's own board had Opus 5.5 at every effort and no GPT-6 model of any size. Opus 5.5 at medium (52.5% for $2.91) beats Claude Fable 5.1 at max (51.8% for $17.28). Even Opus 5.5 at low (43.7% for $1.17) beats GPT-5.6 Sol at max (41.7% for $8.23). Unlike FrontierCode, effort keeps paying inside Cursor, up to 57.8% at max.
If you want to try these side by side in your own editor, Claude Code can point at any provider. I wrote up the two-line switch in how to change the model in Claude Code.
Business automations, office deliverables, computer use and long professional tasks. The picture is closer, and it's where Sol and Astra look best.
AutomationBench 1.0.6
Hover a point for its score and cost. Click one to see what it costs per solved task and what each rival has to spend to match it. Click a model in the legend to hide it.
Zapier ran and priced every model the same way. The unconnected Fable 5.1 point used Opus 5 as a fallback on about 40% of tasks, and its cost excludes the fallback tokens. Source: Zapier AutomationBench leaderboard.
| Model | Effort | Score | Cost per task |
|---|---|---|---|
| Claude Opus 5.5 | low | 23.3% | $0.46 |
| Claude Opus 5.5 | medium | 28.6% | $0.60 |
| Claude Opus 5.5 | high | 32.0% | $0.65 |
| Claude Opus 5.5 | xhigh | 34.4% | $0.80 |
| Claude Opus 5.5 | max | 40.0% | $1.28 |
| GPT-6 Sol | none | 9.1% | $0.25 |
| GPT-6 Sol | low | 21.2% | $0.19 |
| GPT-6 Sol | medium | 26.9% | $0.21 |
| GPT-6 Sol | high | 31.2% | $0.24 |
| GPT-6 Sol | xhigh | 33.2% | $0.27 |
| GPT-6 Sol | max | 32.0% | $0.34 |
| GPT-6 Luna | none | 0.3% | $0.01 |
| GPT-6 Luna | low | 1.2% | $0.01 |
| GPT-6 Luna | medium | 9.4% | $0.02 |
| GPT-6 Luna | high | 14.5% | $0.02 |
| GPT-6 Luna | xhigh | 12.6% | $0.02 |
| GPT-6 Luna | max | 20.7% | $0.04 |
| Claude Fable 5.1 | xhigh | 21.3% | $2.14 |
| Claude Fable 5.1 | max (with fallback) | 31.4% | $2.45 |
| Claude Fable 5.1 | max | 22.4% | $2.45 |
| GPT-6 Astra | none | 25.3% | $1.28 |
| GPT-6 Astra | low | 30.3% | $1.08 |
| GPT-6 Astra | medium | 34.1% | $1.28 |
| GPT-6 Astra | high | 37.1% | $1.45 |
| GPT-6 Astra | xhigh | 39.0% | $1.53 |
| GPT-6 Astra | max | 41.4% | $1.73 |
Zapier's AutomationBench runs end-to-end workflows across 47 business apps and only counts the final state. At high effort, Sol scores 31.2% for $0.24 and Opus 5.5 scores 32.0% for $0.65: the same result for 37% of the price. Sol stops climbing at xhigh (33.2%), and its max is worse.
The top of the board belongs to the expensive settings. GPT-6 Astra at max leads with 41.4% for $1.73, and Opus 5.5 at max is right behind with 40.0% for $1.28. This is one of the few tests where Opus 5.5's max effort clearly earns its cost. Research math is the other.
On GDPval-AA, where Artificial Analysis judges real work products head to head, Opus 5.5 at medium (1576 Elo for $0.86) beats GPT-6 Astra at max (1542 Elo for $4.53). Artificial Analysis puts Sol at max at 1487 Elo and Luna at 1367 Elo. If your output is a document someone has to read, Opus 5.5 is the pick.
OpenAI and Anthropic ran different versions of OSWorld 2.0. Sol scores 64.4% at max in OpenAI's version, and Opus 5.5 scores 81.8% in Anthropic's. The one model both ran, Claude Opus 5, scored 70.2% in one and 74.0% in the other. The setups differ by a few points; the gap between Sol and Opus 5.5 is much larger than that. My read is that Opus 5.5 leads, but that's an inference, not a measurement.
Two OpenAI-only charts round it out. On Agents' Last Exam, Sol at xhigh (55.4% for $1.67) matches Claude Opus 5 at high (55.9% for $7.29) for 23% of the cost. Opus 5.5 isn't on it. And on OpenAI's hard-prompt factuality test, Sol at max makes factual errors on 4.6% of prompts, down from 8.5% for GPT-5.6 Sol. That's the clearest real improvement in the Sol launch.
On research math, Anthropic's system card has Opus 5.5 at 91.2% on ArXivMath at max, against Fable 5.1's 82.9%, and at 67.7% on Humanity's Last Exam with tools, against 65.6%. No GPT-6 model was run on either.
Pick the job and what matters most. Each answer shows the numbers behind it, from a single source per row.
The pattern across all of it: Opus 5.5 when the task is ambiguous or long, Sol when it's well-specified and repeated, Luna when it's simple and huge. Effort matters as much as the model. Opus 5.5 at medium beats most models at max, and Sol at max is usually money you don't need to spend.
Same family, three price points. Sol costs a fifth of Astra per token, and Luna costs a twentieth of Sol.
Luna is the model to put under something bigger. Use it for routing, extraction, first drafts and sub-agent reading, always at max effort, which is its best setting on every chart I pulled and still costs cents.
Sol is OpenAI's workhorse. It doesn't beat GPT-5.6 Sol on intelligence (within 0.6 points at every effort), but it costs 45% to 53% as much per task. Default to high or xhigh. If you already run GPT-5.6 Sol in an agent, like my Hermes Agent setup with no API bill, Sol gets you the same scores for about half the cost per task.
Astra scores higher than Sol on every OpenAI chart, by the widest margin on Terminal-Bench (59.1% vs 43.9%) and AutomationBench. Pay for it on long, multi-app runs and computer use, where a failed attempt costs more than the tokens. I covered its launch in GPT-6 Astra release date, price and benchmarks.
Opus 5.5's best beats Fable 5.1's best on 9 of 9 benchmarks, at 40% of the per-token price.
There's one real difference: Fable 5.1 does more with less thinking. At low effort it beats Opus 5.5 at low on 8 of 8 shared charts. But Opus 5.5 at medium beats Fable 5.1 at low on 8 of 8, and it's cheaper on all 7 charts that publish a cost. Example: on the Intelligence Index, Fable 5.1 at low scores 46.8 for $2.37, and Opus 5.5 at medium scores 51.2 for $1.34.
So the practical advice is to move Fable 5.1 workloads to Opus 5.5 and re-run your own evals. If something you care about gets worse, keep Fable for that one job. The upgrade from Opus 5 is an easier call: Opus 5.5 scores higher on every benchmark both were run on, at 20% lower token prices.
The cheapest way to use frontier models is to use several. Route each step to the cheapest model that can do it, and escalate what fails.
The planner runs once and the executor runs many times. Let Opus 5.5 at medium break the job into specs, then hand each step to Sol at high. You pay Opus prices on a fraction of the tokens.
Reading files, searching, and summarizing logs eat most of an agent's tokens. Give that work to Luna at max and send only the summary up to the expensive model.
For workflows with a clear pass or fail, start with Sol at xhigh. Send only the failures to Opus 5.5 or Astra at max. You get close to top-model reliability at mostly Sol prices.
These are my routing suggestions from the numbers. I found no launch-day report of anyone running this exact combination yet. It is the same idea as the three-tier model cascade I use to run Hermes Agent for $8 a month. If you want one API key across labs, OpenRouter listed GPT-6 Sol and Luna on launch day: what OpenRouter is and how it works.
Per token, Opus 5.5 costs twice what Sol does. Per task, the gap is often smaller, and sometimes Opus is cheaper.
Two things close the gap. First, cache reads cost the same $0.20 per million tokens on Opus 5.5 and Sol, so long agent loops that mostly re-read context get much closer than 2x. Second, models use different amounts of tokens. On FrontierCode, Opus 5.5 at low costs $0.40 a task and Sol at low costs $0.43, despite the 2x price per token. And Opus 5.5 at medium beats Sol at max for 39% of the cost.
Opus 5.5 is also cheaper than the model it replaces: $4 and $20 per million tokens against Opus 5's $5 and $25, and Sol is half of GPT-5.6 Sol's $4 and $20. Luna is 20 times cheaper than Sol on every token type. If you want a cheaper route for smart-enough coding work, I tested one in MiniMax M3 in Claude Code.
Every benchmark I could find for these models, in one grid. Toggle to independent sources only to see the fair fights.
Launch-day reactions from people who used the models, paraphrased and linked. Tags flag who used both, and who has a stake.
Switched to GPT-6 Sol as the daily driver in Codex: not quite Astra, but close enough for most everyday work, faster, and half the price of GPT-5.6 Sol. Every's public summary of the head-to-head with Opus 5.5 calls Sol the better step-by-step collaborator and gives Opus 5.5 the higher ceiling for long autonomous builds.
Reported that Every staffers who had moved to Codex are drifting back to Claude, because Opus 5.5 gives Fable-level output for much less and fixed some of the personality issues that pushed them away from Opus 5.
Ran the pelican-on-a-bicycle SVG test on Opus 5.5 at max effort twice. Both times the model spent the whole 128k-token budget reasoning and never produced the drawing.
The loudest theme in the thread: Sol and Luna read as a price cut, not a capability jump. One commenter said they'd rather have paid double for a real performance gain. Another said they get better value from Opus 5.5.
Cross-vendor skeptics. One said the open-weight MiMo V2.6 Pro beat Sol on quality and price in their own benchmark. The other said Gemini 3.8 Flash beats Sol and Luna on price and coding in their use.
Found that Sol and Luna roughly halve the cost of GPT-5.6 Sol and Luna while their Intelligence Index and Coding Agent Index scores stay about level: better on some evals, worse on others.
Said Opus 5 at max kept failing their real tasks by hitting tool-call limits, while high effort got the work done. Offered as a caution for Opus 5.5 at max too.
Recommends Opus 5.5 for ambiguous, multi-step work where mistakes are expensive, like codebase migrations, and Sol for clear, bounded, high-volume work that's easy to check.
Early testers found Opus 5.5's writing clearer and quicker to get to the point. Anthropic marketed exactly this as a fix for 'Claudish' prose.
Pointed out that Opus 5.5 at high and GPT-6 Astra at high cost about the same per task on Artificial Analysis, and pushed back on claims that OpenAI's API had degraded. The commenter works at OpenAI.
Anthropic documents that most flagged cybersecurity requests to Opus 5.5 get routed to Opus 4.8, and a new biology classifier joins the cyber one. The New Stack ran the headline that your agent calls might quietly land on an older model.
Opus 5.5 shipped in GitHub Copilot on day one. OpenRouter and Vercel's AI Gateway listed all three launches the same day. Availability, not a verdict.
It's launch day, so treat all of this as early. I looked for people who used more than one of these models, then weighed each theme by how many independent sources back it and whether the benchmarks agree.
| Signal | Strength | What it rests on | What to do with it |
|---|---|---|---|
| Sol is GPT-5.6 Sol at half the price | Strong | Artificial Analysis's numbers, the loudest HN theme, and a tech-press analysis all point the same way. | If you already run GPT-5.6 Sol, switch for the bill. Don't expect new capability. |
| Opus 5.5 has the higher ceiling on hard, long work | Solid | Independent benchmarks agree. Hands-on reports are few but point the same direction (Every, Unite.AI). | Default to Opus 5.5 for repo-scale and terminal work. Re-check in two weeks when more people have shipped with it. |
| Opus 5.5 at max effort often wastes money | Solid | A reproducible failure from Simon Willison, matching HN advice, and flat or falling curves from xhigh to max on FrontierCode and Terminal-Bench. | Run medium by default and xhigh for hard tickets. Save max for multi-app automations, where the curve still climbs. |
| Luna is the volume tier, not a rival | Solid | Nobody in the sample pitched Luna against the flagships. The cost data backs the framing. | Use it under a bigger model, not instead of one. |
| Coders are drifting back from Codex to Claude | Thin | Two named people at one company, in one outlet's piece. | A story worth watching, not a trend yet. |
| Cheaper rivals beat Sol at coding | Thin | Two anecdotal HN comments naming MiMo V2.6 Pro and Gemini 3.8 Flash. | If price is the whole reason you'd pick Sol, test those two as well. |
| Flagged security prompts quietly change models | Documented | Anthropic's own docs describe the reroute to Opus 4.8. How often it happens in normal use is unknown. | If you build security tooling on Opus 5.5, log which model actually answered. |
The gaps matter too. I couldn't read Reddit threads, found no launch-day YouTube first impressions, and only saw the top comments in each large Hacker News thread. Among the few people who tried both, the pattern is consistent: Opus 5.5 for serious building, Sol as the cheaper daily driver for well-defined work.
9 cases where the same model at the same effort gets different numbers from two sources. This is why every chart above sticks to one source at a time.
Anthropic's own run puts Opus 5.5 6.8 points higher than Artificial Analysis does, at 8.4x the cost per task.
Different harness, time limits and trial counts. Compare models inside one source, never across them.
| Opus 5.5xhigh Anthropic launch chart | 66.4% | $7.35 |
| Opus 5.5xhigh Artificial Analysis | 59.6% | $0.88 |
| Opus 5.5max Anthropic launch chart | 64.8% | $11.24 |
| Opus 5.5max Artificial Analysis | 59.6% | $1.31 |
One model, three published Terminal-Bench scores: 53.9%, 52.3% and 49.0%.
Anthropic's page cites the public board at 51.8% for Opus 5; the live board showed 53.9% when I checked. Leaderboards move.
| Opus 5xhigh tbench.ai public board (live) | 53.9% | no cost |
| Opus 5max Anthropic launch chart | 52.3% | $15.83 |
| Opus 5max Artificial Analysis | 49.0% | $1.93 |
OpenAI's chart copies Datacurve's GPT-5.6 Sol scores exactly but lists each run 23% to 24% cheaper.
Same scores, different cost math. Datacurve reports cost per scored attempt; OpenAI does not say how it re-priced the runs.
| 5.6 Solmax OpenAI launch chart | 72.7% | $6.46 |
| 5.6 Solmax Datacurve leaderboard | 72.7% | $8.39 |
For GPT-5.6 Luna the two sources don't even agree on the score: 62.2% on OpenAI's chart, 67.2% on Datacurve's board, at 5.7x the cost.
OpenAI ran its own models itself and took Claude numbers from Datacurve. Luna's cheap DeepSWE result has no independent run yet.
| 5.6 Lunamax OpenAI launch chart | 62.2% | $0.53 |
| 5.6 Lunamax Datacurve leaderboard | 67.2% | $3.03 |
OpenAI and Anthropic ran different versions of OSWorld 2.0. Claude Opus 5 at max scores 70.2% in one and 74.0% in the other, at $24.11 vs about $11.60 a task.
GPT-6 Sol's 64.4% and Opus 5.5's 81.8% come from these two different tests. They are not a head-to-head.
| Opus 5max OpenAI launch chart (offline set) | 70.2% | $24.11 |
| Opus 5max Anthropic system card (Sept 10 files) | 74.0% | $11.60 |
OpenAI's chart prices Claude Opus 5 at 2.2x to 2.5x what Zapier's own leaderboard lists, low through xhigh effort.
Anthropic's chart matches Zapier here. When a vendor plots a rival, check the rival's cost against the neutral board.
| Opus 5med OpenAI launch chart | 23.9% | $2.22 |
| Opus 5med Zapier leaderboard | 23.9% | $0.89 |
Anthropic lists its own Opus 5.5 max run at $1.37 a task; Zapier's board says $1.28.
Small, and in the unflattering direction for Anthropic. Not every gap favors the vendor.
| Opus 5.5max Anthropic launch chart | 40.0% | $1.37 |
| Opus 5.5max Zapier leaderboard | 40.0% | $1.28 |
Anthropic's FrontierCode chart shows Fable 5.1 at low effort scoring 52.8%. Cognition's leaderboard says 49.8%.
Every other Claude point on Anthropic's chart matches Cognition. This one doesn't, and nothing on the page explains it.
| Fable 5.1low Anthropic launch chart | 52.8% | $2.47 |
| Fable 5.1low Cognition leaderboard | 49.8% | $2.38 |
OpenAI prices its own GPT-6 Sol runs 3% to 5% above Cognition's figures.
Cognition rounds to the cent, which explains part of it. The direction is against OpenAI, so this is not flattering math.
| Solmax OpenAI launch chart | 49.3% | $2.14 |
| Solmax Cognition leaderboard | 49.3% | $2.07 |
On most hard work, yes. On the independent leaderboards, Opus 5.5 at medium effort scores 54.6% on FrontierCode for $0.80 a task, above every GPT-6 Sol setting, and 52.5% on Artificial Analysis's Terminal-Bench run for the same $0.40 that Sol at max needs to reach 43.9%. Sol is the better value on business automations, where it roughly matches Opus 5.5 for about 40% of the cost.
Not smarter, but much cheaper. On the Artificial Analysis Intelligence Index, GPT-6 Sol lands within 0.6 points of GPT-5.6 Sol at every effort level for 45% to 53% of the cost per task. It also makes about half as many factual errors on OpenAI's own hard-prompt test.
High-volume, simple work: triage, extraction, tagging, summaries and sub-agent reading. Luna at max scores 37.3 on the Intelligence Index for $0.068 a task, the cheapest model on that board. It struggles with long tool loops: 12.6% on Terminal-Bench.
Sol for most work, Astra for escalation. Astra at max leads AutomationBench at 41.4% against Sol's best of 33.2%, and it scores higher on every OpenAI chart. It also costs five times as much per token, so send it the long multi-app workflows and the tasks that already failed once.
On every benchmark both were run on (9 of 9), Opus 5.5's best score beats Fable 5.1's best, at 40% of Fable's per-token price. Fable 5.1 does score higher at low effort, but Opus 5.5 at medium beats Fable 5.1 at low on all 8 shared charts and costs less on all 7 that publish a cost.
Medium, the default. It is Opus 5.5's best FrontierCode setting and gets most of the Terminal-Bench score. Use xhigh for hard terminal and infrastructure work. Use max only for long business workflows across apps, where AutomationBench still climbs from 34.4% at xhigh to 40.0% at max.
High or xhigh. Max rarely pays off: on AutomationBench it scores lower than xhigh, and on FrontierCode it adds under a point for about 1.6x the cost. The exception is terminal work, where Sol jumps from 30.3% at xhigh to 43.9% at max.
Per million tokens: Claude Opus 5.5 is $4 input and $20 output, GPT-6 Sol is $2 and $10, and GPT-6 Luna is $0.10 and $0.50. Cached input costs $0.20 on both Opus 5.5 and Sol, and $0.01 on Luna. GPT-6 Astra and Claude Fable 5.1 are both $10 and $50.
No. Neither vendor plotted the other's new model, and their cost math differs from the neutral leaderboards. For example, Anthropic's own Terminal-Bench run puts Opus 5.5 at xhigh at 66.4%, while Artificial Analysis measures 59.6%. Use leaderboards that run every model the same way, like Cognition's FrontierCode, Zapier's AutomationBench and Artificial Analysis.
Not yet. DeepSWE's latest version is v1.1, and its board had not run any of the September 22 launches when I captured it. The public Terminal-Bench 4.0 and OSWorld 2.0 boards hadn't added them either. On launch day, all three new models were on Cognition's FrontierCode, Zapier's AutomationBench and Artificial Analysis. Cursor's CursorBench had Opus 5.5 only.
Every number on this page comes from a public source captured on September 22, 2026. Nothing is estimated unless the chart says so.
Launch-page charts were read from their embedded chart data and checked by hovering each point. A few system-card charts had no embedded data, so I read those off the image, accurate to about half a percent on cost. I re-checked the independent boards around 4 PM Pacific to pick up late additions. Each chart carries its source badge, and no chart mixes two sources. You can download all 471 data points as a CSV.
OpenAI released GPT-6 Astra on September 3, 2026. The release date, the API price, the benchmarks OpenAI buried, and the line in the system card nobody is quoting.
Claude Code will talk to almost any model, not just Anthropic's. Here's the two-line config that switches it, how to keep the change scoped to one folder, and how to check which model you're actually running.
I built the same website twice, generated its cover images and a voiceover, and it used 13% of my weekly limit. The real usage numbers from running MiniMax M3 inside Claude Code, including where it lost to Claude.