Back to blog
AI Tools193 data points · 10 benchmarks · 5 models

GPT-6.1 Sol vs GPT-6 Astra vs Claude Opus 5.5: Is it really near-Astra at a fifth of the price? Benchmarks and cost per task, at every effort level

OpenAI says GPT-6.1 Sol nearly matches GPT-6 Astra for a fifth of the token price. I checked that claim against every chart in the launch and every independent leaderboard that had already run it, then put Claude Opus 5.5 next to both. Here's where the claim holds, where it breaks, and which model to use for what.

Score vs cost per task, low to max effortArtificial Analysis Intelligence Index
30405060$0.10$0.20$0.50$1$2$5$10cost per task, log scaleGPT-6 SolAstra52.7 · $3.26 maxOpus 5.557.6 · $5.98 max6.1 Sol51.8 · $0.72 max

Verdict: the near-Astra claim, checked

Mostly true. GPT-6.1 Sol lands within 3.5 points of GPT-6 Astra, or beats it, on 8 of the 10 charts that include both, at 8% to 23% of Astra's cost per task. It breaks on science: 11.1 points behind. Claude Opus 5.5 still wins on coding and documents, but Sol now matches it on terminal work for about a third of the cost.

  • GPT-6.1 SolBest value · new

    Your default for well-specified work at volume: automations, reviews, computer use, terminal jobs with tests. Run it at medium, and go to max for shell and desktop work.

    • Intelligence Index 51.8 for $0.72, against Astra's 52.7 for $3.26
    • Beats Astra's best on OpenAI's DeepSWE chart: 75.2% at high for $0.65
    • Weak spot: science terminal work, 11.1 points behind Astra
    $2 / $10in / out
    per 1M tokens
  • Claude Opus 5.5Best overall

    Still the pick when quality decides: real codebases, front ends, documents people read, and anything long and ambiguous. Run it at medium.

    • Top FrontierCode score on the board: 54.6% at medium for $0.80
    • Highest Intelligence Index of the five models here: 57.6
    • GDPval-AA: 1846 Elo against Sol's 1575 Elo
    $4 / $20in / out
    per 1M tokens
  • GPT-6 AstraThe escalation model

    Keep it for the jobs 6.1 Sol measurably loses: scientific computing, the hardest multi-app automations, and runs that already failed once on Sol.

    • Terminal-Bench-Science: 68.1% against Sol's 57.0%
    • AutomationBench: 41.4% against Sol's 36.1%
    • Costs 5x Sol per input and output token
    $10 / $50in / out
    per 1M tokens
  • GPT-6 SolReplace it

    Same token price as 6.1 Sol and worse on almost every chart. Switch the model ID.

    • Terminal-Bench at max: 43.9% to 56.1%, and cheaper per task
    • OSWorld 2.0 at max: 64.4% to 71.4%
    • Cached input drops from $0.20 to $0.10 per million tokens
    $2 / $10in / out
    per 1M tokens

Is GPT-6.1 Sol really near-Astra?

OpenAI's headline is near-Astra performance at one-fifth the price. The price part is exact. The performance part depends on the job.

The token price is a clean one-fifth: $2 and $10 per million input and output tokens against Astra's $10 and $50. Cost per task isn't a fixed ratio, because the two models use different amounts of tokens and tools. On the charts below, Sol costs 8% to 23% of what Astra costs for the same task.

The headline number comes from Artificial Analysis, which runs every model the same way. Sol at max scores 51.8 on its Intelligence Index for $0.72 a task. Astra at max scores 52.7 for $3.26. That's 98% of the score for 22% of the cost. It's an index, not a percentage of intelligence, but it's the fairest single comparison on launch day. For Astra's own launch numbers, see GPT-6 Astra release date, price and benchmarks.

Is it really near-Astra? Every chart that has both

Green = ahead or tied · amber = within 3.5 points · red = further behind
Benchmark and source6.1 SolAstraGapSol's cost vs Astra
Intelligence IndexBroad reasoning composite · max vs maxArtificial Analysis51.8max · $0.7252.7max · $3.260.8 behind22%
Terminal-Bench 4.0Terminal and command-line work · max vs maxArtificial Analysis56.1%max · $1.8259.1%max · $8.503.0 pts behind21%
GDPval-AAKnowledge-work deliverables · max vs maxArtificial Analysis1575 Elomax · $0.891542 Elomax · $4.5333 Elo ahead20%
FrontierCode 1.1Agentic coding · best vs bestCognition FrontierCode leaderboard50.2%med · $0.3653.3%max · $4.593.1 pts behind8%
DeepSWE 1.1Long engineering projects · best vs bestOpenAI's chart75.2%high · $0.6574.1%xhigh · $4.431.1 pts ahead15%
OSWorld 2.0 (offline)Computer use · best vs bestOpenAI's chart71.4%max · $1.2773.5%max · $9.442.1 pts behind13%
GDP.pdfDocuments and office work · best vs bestOpenAI's chart32.0%high · $0.3532.2%xhigh · $1.910.2 pts behind18%
AutomationBench 1.0.6Business workflows · best vs bestOpenAI's chart36.1%max · $0.3041.4%max · $1.735.3 pts behind17%
Factual errors, lower is betterHallucination on hard prompts · best vs bestOpenAI's chart4.1%xhigh · $0.103.9%high · $0.480.2 pts behind21%
Terminal-Bench-Science 0.1Scientific terminal work · best vs bestOpenAI's chart57.0%max · $5.4768.1%max · $23.8011.1 pts behind23%

Independent rows compare both models at max, the setting each source ran for Astra. OpenAI's charts run both models at every effort, so those rows compare each model's best setting. Cost is what each source measured per task, not the token price.

Where the claim holds

Computer use is the strongest case. On OpenAI's OSWorld 2.0 chart, Sol at max scores 71.4% for $1.27 and Astra at max scores 73.5% for $9.44. On DeepSWE, OpenAI's long engineering benchmark, Sol at high (75.2%) beats Astra's best setting (74.1% at xhigh) for about a seventh of the cost. On GDPval-AA, the one independent test of real work products that has both, Sol at max outscores Astra at max: 1575 Elo against 1542 Elo.

Where it breaks

Scientific computing. On OpenAI's own Terminal-Bench-Science chart, Sol at max scores 57.0% and Astra at max scores 68.1%. Even at its best setting, Sol trails by 11.1 points. Business automations are the other soft spot: Sol tops out at 36.1% on AutomationBench against Astra's 41.4%. And on FrontierCode, Astra at max still beats Sol's best by 3.1 points, though it costs $4.59 a task to Sol's $0.36.

Same method for every modelBroad reasoning composite

Intelligence Index

Score vs cost per index task
30354045505560$0.10$0.20$0.50$1$2$5$10Cost per index task (USD, log scale)Indexlowmedhighxhighmaxlowmedhighmax

Hover any dot to read its score and cost. Click it to pin it and see the cost per solved task, plus the cheapest setting where each other model catches up. Use the legend to toggle models on and off.

Index v4.3.2. GPT-6.1 Sol captured Sept 29 at every effort. Rival efforts come from the Sept 22 capture; every rival's max point was re-captured Sept 29 and is unchanged. Source: Artificial Analysis.

View this chart as a table
ModelEffortIndexCost per index task
GPT-6.1 Sollow42.1$0.13
GPT-6.1 Solmedium47.8$0.21
GPT-6.1 Solhigh50.2$0.32
GPT-6.1 Solxhigh51.0$0.39
GPT-6.1 Solmax51.8$0.72
GPT-6 Astralow45.8$0.82
GPT-6 Astramedium49.6$1.54
GPT-6 Astrahigh50.9$1.73
GPT-6 Astraxhigh52.4$2.31
GPT-6 Astramax52.7$3.26
Claude Opus 5.5low42.3$0.55
Claude Opus 5.5medium51.2$1.34
Claude Opus 5.5high53.6$1.82
Claude Opus 5.5xhigh56.0$3.46
Claude Opus 5.5max57.6$5.98
Claude Fable 5.1low46.8$2.37
Claude Fable 5.1medium48.9$2.98
Claude Fable 5.1high51.2$3.91
Claude Fable 5.1xhigh53.2$5.98
Claude Fable 5.1max53.4$7.63

See the difference: launch-day builds, side by side

Charts give you averages. These are real builds people posted on September 29, most of them the same task on two or more models. Each is the creator's own test, with what they did and didn't disclose.

GPT-6.1 Sol vs Claude Opus 5.5

Opus 5.5 still has the higher ceiling. What changed this week is how close Sol gets for the money, especially on terminal work.

Coding: Opus 5.5 keeps the top score

Cognition's FrontierCode grades real pull requests on whether a maintainer would merge them. Opus 5.5 at its default medium effort scores 54.6% for $0.80 a task, still the best result on the board. Sol's best is 50.2%, also at medium, for $0.36. That beats Opus 5.5 at low (47.3% for $0.40) for less money, so for small, well-scoped tickets Sol is the cheaper good-enough option.

FrontierCode 1.1 score and cost per taskCognition FrontierCode leaderboard
Opus 5.5med54.6%$0.80
Astramax53.3%$4.59
6.1 Solmed50.2%$0.36
Opus 5.5low47.3%$0.40
6.1 Solmax47.6%$0.85

Terminal work: Sol closes the gap for a third of the cost

Last week this was Opus 5.5's biggest win. Now, on Artificial Analysis's Terminal-Bench 4.0 run, Sol at max scores 56.1% for $1.82 a task, and Opus 5.5 at high scores 56.6% for $5.12. Opus 5.5 still has the top score (59.6%), but it pays $13.11 a task for it.

Terminal-Bench 4.0 score and cost per taskArtificial Analysis
Opus 5.5max59.6%$13.11
Astramax59.1%$8.50
Opus 5.5high56.6%$5.12
6.1 Solmax56.1%$1.82
Opus 5.5med52.5%$4.04
6.1 Solhigh51.5%$0.83

Documents, automations and science

On GDPval-AA, where judges compare real work products head to head, Opus 5.5 is far ahead: 1846 Elo against Sol's 1575 Elo. If a person has to read the output, start with Opus. On OpenAI's AutomationBench chart, which reuses Zapier's published Opus 5.5 results with fallbacks, Sol at xhigh (35.5% for $0.25) matches Opus 5.5 at xhigh (35.8% for $0.89). Opus 5.5 at max tops that chart at 42.5%. On OpenAI's office-document test, GDP.pdf, Sol at high scores 32.0% against Opus 5.5's best of 28.8%, on OpenAI's run of it. On science, Opus 5.5 at max is 63.3% to Sol's 57.0%.

Same method for every modelTerminal and command-line work

Terminal-Bench 4.0

Score vs cost per task
010203040506070$0.50$1$2$5$10$20Cost per task (USD, log scale)Score (%)lowmedhighxhighmaxlowmedhighxhighmax

Hover any dot to read its score and cost. Click it to pin it and see the cost per solved task, plus the cheapest setting where each other model catches up. Use the legend to toggle models on and off.

AA's own Terminal-Bench 4.0 run, not the tbench.ai leaderboard (which has no GPT-6.1 Sol yet). Cost is AA's real cost per Terminal-Bench task. Rival efforts from Sept 22; max points re-checked Sept 29. Source: Artificial Analysis.

View this chart as a table
ModelEffortScoreCost per taskOutput tokens
GPT-6.1 Sollow30.8%$0.3817.4k
GPT-6.1 Solmedium48.0%$0.6132.0k
GPT-6.1 Solhigh51.5%$0.8343.6k
GPT-6.1 Solxhigh54.0%$1.0356.9k
GPT-6.1 Solmax56.1%$1.82108k
GPT-6 Astramax59.1%$8.50
Claude Opus 5.5low31.3%$2.08
Claude Opus 5.5medium52.5%$4.04
Claude Opus 5.5high56.6%$5.12
Claude Opus 5.5xhigh59.6%$8.78
Claude Opus 5.5max59.6%$13.11
Claude Fable 5.1max52.0%$19.22

I compared Opus 5.5 with last week's GPT-6 Sol in Claude Opus 5.5 vs GPT-6 Sol and Luna, and tested Opus 5.5 hands-on in my Claude Opus 5.5 review. If you want to run both side by side in one editor, Claude Code can point at any provider: how to change the model in Claude Code.

What changed from GPT-6 Sol

GPT-6 Sol launched on September 22. Seven days later, 6.1 Sol has the same input and output price and better numbers on almost every chart.

One week, same price per token: what changed

Same source and effort in every row
Benchmark, effort and source6 Sol6.1 SolChangeCost change
Intelligence Index, maxArtificial Analysis47.551.8+4.3−31%
Terminal-Bench 4.0, maxArtificial Analysis43.9%56.1%+12.1 pts−54%
FrontierCode, mediumCognition FrontierCode leaderboard45.9%50.2%+4.3 pts−53%
FrontierCode, maxCognition FrontierCode leaderboard49.3%47.6%−1.7 pts−59%
DeepSWE, highOpenAI's chart65.3%75.2%+10.0 pts+1%
OSWorld 2.0, maxOpenAI's chart64.4%71.4%+7.0 pts−62%
Terminal-Bench-Science, maxOpenAI's chart27.6%57.0%+29.4 pts−55%
AutomationBench, xhighOpenAI's chart33.2%35.5%+2.3 pts−9%
Factual errors, low (lower is better)OpenAI's chart11.4%7.7%−3.7 pts−9%
Factual errors, max (lower is better)OpenAI's chart4.6%4.6%0.0 pts−27%

The biggest jumps are in long tool loops: Terminal-Bench at max goes from 43.9% to 56.1% while getting cheaper per task, and OpenAI's science chart goes from 27.6% to 57.0%. FrontierCode at medium improves from 45.9% to 50.2% at less than half the cost.

Two things didn't improve. FrontierCode at max slips from 49.3% to 47.6%. And the factual-error rate at max is flat (4.6% to 4.6%). The error-rate improvement is real at low effort, from 11.4% to 7.7%, so don't read it as a blanket drop in hallucinations. OpenAI's launch post doesn't say GPT-6 Sol is being retired; if you pinned gpt-6-sol, switch the model ID to gpt-6.1-sol yourself.

Same method for every modelTerminal and command-line work

Terminal-Bench 4.0

Score vs cost per task
010203040506070$0.50$1$2$5$10$20Cost per task (USD, log scale)Score (%)lowmedhighxhighmaxlowmedhighxhighmax

Hover any dot to read its score and cost. Click it to pin it and see the cost per solved task, plus the cheapest setting where each other model catches up. Use the legend to toggle models on and off.

AA's own Terminal-Bench 4.0 run, not the tbench.ai leaderboard (which has no GPT-6.1 Sol yet). Cost is AA's real cost per Terminal-Bench task. Rival efforts from Sept 22; max points re-checked Sept 29. Source: Artificial Analysis.

View this chart as a table
ModelEffortScoreCost per taskOutput tokens
GPT-6.1 Sollow30.8%$0.3817.4k
GPT-6.1 Solmedium48.0%$0.6132.0k
GPT-6.1 Solhigh51.5%$0.8343.6k
GPT-6.1 Solxhigh54.0%$1.0356.9k
GPT-6.1 Solmax56.1%$1.82108k
GPT-6 Astramax59.1%$8.50
GPT-6 Sollow9.1%$0.49
GPT-6 Solmedium18.7%$1.12
GPT-6 Solhigh26.3%$1.60
GPT-6 Solxhigh30.3%$1.91
GPT-6 Solmax43.9%$4.00

Which effort level to use

More thinking isn't always better with 6.1 Sol. On two coding charts, the top score comes before max.

On FrontierCode, Sol scores 45.5%, 50.2%, 48.0%, 49.3%, 47.6% from low to max. Medium is the peak, and max costs more than twice as much for a lower score. On OpenAI's DeepSWE chart, high (75.2%) beats both xhigh and max (71.9%). Readers on r/codex noticed the same shape on OpenAI's chart on launch day.

Terminal and computer-use work is the exception. Sol keeps climbing to max on Terminal-Bench, OSWorld and the science chart. My rule: start at medium, the default. Go to high for long engineering tasks, and max for shell and desktop agents. Drag the slider to see what each model gets you at a fixed budget per task.

Set a budget, see who wins

Highest published score each model reaches without going over it
Benchmark
$1.00per task
  • 6.1 Sol51.8max · $0.72
  • Astra45.8low · $0.82
  • 6 Sol44.1xhigh · $0.53
  • Opus 5.542.3low · $0.55
  • Fable 5.1needs $2.37+

At $1.00 a task, GPT-6.1 Sol at max effort gets the most done on Intelligence Index. Source: Artificial Analysis. Same method for every model

GPT-6.1 Sol price, access and Ultrafast

Official prices captured on launch day, September 29, 2026. Check OpenAI's pricing page before you budget; rollout and plan limits change.

WhatGPT-6.1 SolNotes
API, standard$2 input · $0.10 cached input · $10 output per 1M tokensAstra: $10 · $1 · $50. Opus 5.5: $4 · $0.20 · $20.
Cache writes$2.50 per 1M tokensAstra: $12.50
Long promptsAbove 272K input tokens, the whole request bills at 2x input and 1.5x outputApplies to the full request, not just the tokens past the threshold
Batch, Flex, FastBatch and Flex are half price; Fast is 2x standardFast is not the upcoming Ultrafast mode
Context and output1.05M-token context, 128K max outputEffort levels: low, medium (default), high, xhigh, max
ChatGPT and CodexEligible paid plans, in ChatGPT Work and CodexRolling out by account and workspace. Not in the regular chat picker
Plus, $20/monthIncludedOpenAI's Codex docs estimate 15 to 160 local 6.1 Sol messages per five hours, not guaranteed
Pro, $100 to $500/monthIncludedNo five-hour limit today; weekly allowances still apply
Business$20 per seat per month billed yearly, $25 monthly, two seats minimumEnterprise and Edu: custom
OpenRouterListed at launch at the same $2 / $10Also batch and Pro aliases
UltrafastAnnounced, not live at launchUp to 8x faster generation in Codex. Price not published

Ultrafast isn't out yet. OpenAI announced a GPT-6.1 Sol Ultrafast mode at DevDay for the coming days, with up to 8x faster token generation in Codex. Its own pages disagree on the API figure (up to 6x in the DevDay recap, up to 8x in the API guide), and there is no Sol-specific price yet. Today's Astra Ultrafast is limited to the $500 Pro tier and eligible Enterprise and Edu plans. Faster tokens also don't guarantee faster tasks: Artificial Analysis measured 6.1 Sol at max generating slower than GPT-6 Sol at max.

What a month of GPT-6.1 Sol costs

Same token mix on every model, standard API prices

On this mix GPT-6.1 Sol costs 17% of GPT-6 Astra and 50% of Opus 5.5 per call. Cached input is where Sol pulls further ahead: $0.10 per million, against $1 on Astra and $0.20 on Opus 5.5.

  • 6.1 Sol$1,960/mo
  • 6 Sol$2,320/mo
  • Opus 5.5$3,920/mo
  • Fable 5.1$8,900/mo
  • Astra$11,600/mo

Same token counts for every model, which flatters wordy models. In the benchmarks, models spend very different amounts of tokens on the same task: on Artificial Analysis's Terminal-Bench run, GPT-6.1 Sol at high writes about 43.6k output tokens per task and spends $0.83, while Opus 5.5 at high spends $5.12. Use the curves above for measured cost per task, and this for list-price math.

If you want one API key across OpenAI and Anthropic models, OpenRouter listed 6.1 Sol on launch day: what OpenRouter is and how it works.

Pick a model for the job

Pick the job and what matters most. Each answer shows the numbers behind it, one source per row.

What are you using it for?

GPT-6.1 Solmedium effort

Sol's best FrontierCode score comes at medium, and it costs less than Opus 5.5 at low while scoring higher. Going above medium makes it worse here, not better.

  • 6.1 Sol mediumFrontierCode 1.1 · Cognition FrontierCode leaderboard50.2% · $0.36
  • Opus 5.5 lowFrontierCode 1.1 · Cognition FrontierCode leaderboard47.3% · $0.40
  • 6.1 Sol maxFrontierCode 1.1 · Cognition FrontierCode leaderboard47.6% · $0.85

My recommendations, read off the published numbers. Scores and costs come from one source per row; I never mix two sources in one comparison.

The pattern: 6.1 Sol for well-specified work you run a lot, Opus 5.5 when quality decides, Astra only where a chart shows it winning. One split worth trying is the one Theo describes in his review: let Opus 5.5 build and let Sol review and audit the result.

Launch-day reactions, weighed

Launch-day reactions from people who actually used it, weighed by how many independent sources back each theme and whether the benchmarks agree.

SignalStrengthWhat it rests onWhat to do with it
Close to Astra on most tests, for a fraction of the costStrongArtificial Analysis, Cognition and seven of OpenAI's own charts. Lumina, who ran both on the same prompt, says they feel very similar.Move Astra workloads to 6.1 Sol and keep Astra for the jobs your own tests say it wins.
Science and long terminal work still favor AstraSolidOpenAI's own Terminal-Bench-Science chart shows an 11-point gap. Artificial Analysis's separate run shows a smaller one.Don't route research computing or long shell jobs to Sol without testing first.
Opus 5.5 still makes the more polished front endEarlyjjcm's side-by-side LCARS pages and BridgeBench's ocean scene both favor Opus. GDPval-AA, a judged deliverables test, has Opus far ahead.For UI and anything a person will read, start with Opus 5.5.
Sol for reviews, Opus for buildingThinOne early-access reviewer (Theo) with a detailed video. No controlled test yet.A good split to try: let Sol audit and review what Opus writes.
More effort doesn't always score higherSolidFrontierCode peaks at medium and DeepSWE at high. Readers on r/codex spotted the same shape on OpenAI's chart.Start at medium, the default, and only go up where a chart shows it pays.
Faster tokens don't mean faster tasksEarlyBridgeBench timed 6.1 Sol at 97 seconds on a scene GPT-6 Sol did in 44. Ultrafast isn't out yet.Judge speed on your own tasks, not on tokens per second.

Treat all of this as early. Most launch-day posts talked about the claims and older models rather than new tests, and nobody has published a controlled head-to-head yet. The gaps matter too: I couldn't read the AP and WIRED originals, and I didn't count search-preview comments that weren't in the threads I opened.

Every result in one grid

All ten benchmarks, each model's best or max setting. Switch to independent sources only to hide OpenAI's own charts.

All 10 benchmarks side by side

Benchmark and source6.1 SolAstraOpus 5.5Fable 5.1
Intelligence IndexArtificial Analysis51.8max · $0.7252.7max · $3.2657.6max · $5.9853.4max · $7.63
FrontierCode 1.1Cognition FrontierCode leaderboard50.2%med · $0.3653.3%max · $4.5954.6%med · $0.8050.9%med · $3.28
Terminal-Bench 4.0Artificial Analysis56.1%max · $1.8259.1%max · $8.5059.6%xhigh · $8.7852.0%max · $19.22
GDPval-AAArtificial Analysis1575 Elomax · $0.891542 Elomax · $4.531846 Elomax · $8.921735 Elomax · $9.77
AutomationBench 1.0.6Zapier AutomationBench leaderboardnot run41.4%max · $1.7342.5%max · $1.44not run
AutomationBench 1.0.6OpenAI launch chart36.1%max · $0.3041.4%max · $1.7342.5%max · $1.44not run
DeepSWE 1.1OpenAI launch chart75.2%high · $0.6574.1%xhigh · $4.43not runnot run
OSWorld 2.0 (offline)OpenAI launch chart71.4%max · $1.2773.5%max · $9.44not runnot run
GDP.pdfOpenAI launch chart32.0%high · $0.3532.2%xhigh · $1.9128.8%high · $0.83not run
Terminal-Bench-Science 0.1OpenAI launch chart57.0%max · $5.4768.1%max · $23.8063.3%max · $23.21not run
Factual errors, lower is betterOpenAI launch chart4.1%xhigh · $0.103.9%high · $0.48not runnot run
Rows led1 of 104 of 116 of 80 of 4

"Rows led" counts only rows the model was actually run on, and only against the models shown. Rows with an amber dot are OpenAI's own charts, which only include Claude where OpenAI copied a public result, so several Claude cells say "not run". * From a different source than the row, for reference only.

When two sources give different numbers

5 cases where the same model at the same effort gets different numbers from two sources. This is why every chart above sticks to one source at a time.

  • Terminal-Bench-Science 0.1

    GPT-6 Astra at max: 68.1% on OpenAI's chart and the official board, 63.3% in Artificial Analysis's run

    OpenAI's number matches the official board. AA's separate run lands 4.8 points lower. The cause isn't published, so compare within one source.

    Astramax
    OpenAI launch chart
    68.1%$23.80
    Astramax
    Terminal-Bench-Science board
    68.1%no cost
    Astramax
    Artificial Analysis
    63.3%no cost
  • Terminal-Bench-Science 0.1

    Claude Opus 5.5 at max: 63.3% on OpenAI's chart and the official board, 59.0% in Artificial Analysis's run

    OpenAI's number matches the official board. AA's separate run lands 4.3 points lower. The cause isn't published, so compare within one source.

    Opus 5.5max
    OpenAI launch chart
    63.3%$23.21
    Opus 5.5max
    Terminal-Bench-Science board
    63.3%no cost
    Opus 5.5max
    Artificial Analysis
    59.0%no cost
  • Terminal-Bench-Science 0.1

    GPT-6.1 Sol at max: 57.0% on OpenAI's chart, 58.1% in Artificial Analysis's run

    Here AA scores Sol slightly higher than OpenAI does, while AA scores Astra and Opus lower. In AA's run the gap to Astra shrinks from about 11 points to about 5.

    6.1 Solmax
    OpenAI launch chart
    57.0%$5.47
    6.1 Solmax
    Artificial Analysis
    58.1%no cost
  • GDP.pdf

    Claude Opus 5.5 at max on GDP.pdf: 26.2% on OpenAI's chart, 30.6% on Surge's board

    Surge's own board scores it 4.4 points higher than OpenAI's chart. Surge hasn't added GPT-6.1 Sol, so there is no independent GDP.pdf number for it yet.

    Opus 5.5max
    OpenAI launch chart
    26.2%$1.55
    Opus 5.5max
    Surge GDP.pdf board
    30.6%no cost
  • GDP.pdf

    GPT-6 Astra at max on GDP.pdf: 31.0% on OpenAI's chart, 34.2% on Surge's board

    Surge's own board scores it 3.2 points higher than OpenAI's chart. Surge hasn't added GPT-6.1 Sol, so there is no independent GDP.pdf number for it yet.

    Astramax
    OpenAI launch chart
    31.0%$2.08
    Astramax
    Surge GDP.pdf board
    34.2%no cost

GPT-6.1 Sol FAQ

Is GPT-6.1 Sol as good as GPT-6 Astra?

Close on most tests, not all. On the Artificial Analysis Intelligence Index, GPT-6.1 Sol at max scores 51.8 against Astra's 52.7, for 22% of the cost per task. It lands within 3.5 points of Astra, or ahead, on 8 of the 10 charts that include both. The big exception is scientific terminal work, where it trails by 11.1 points on OpenAI's own chart.

Is GPT-6.1 Sol better than Claude Opus 5.5?

Not overall. Opus 5.5 scores higher on the Intelligence Index (57.6 vs 51.8), on FrontierCode coding (54.6% vs 50.2%) and by a wide margin on GDPval-AA work products. Sol is the value pick: on Terminal-Bench it gets within a point of Opus 5.5 at high for 36% of the cost per task, and on business automations it matches Opus 5.5 at xhigh for under a third of the price.

What changed from GPT-6 Sol to GPT-6.1 Sol?

Same token price, better results, one week apart. The Intelligence Index at max rises from 47.5 to 51.8 while the cost per task drops from $1.06 to $0.72. Terminal-Bench at max jumps from 43.9% to 56.1%, and OpenAI's Terminal-Bench-Science chart roughly doubles, from 27.6% to 57.0%. Cached input got cheaper too: $0.10 per million tokens instead of $0.20.

How much does GPT-6.1 Sol cost?

In the API, $2 per million input tokens, $0.10 per million cached input tokens and $10 per million output tokens, one fifth of GPT-6 Astra's $10 and $50. Requests above 272K input tokens are billed at 2x input and 1.5x output for the whole request. Batch and Flex are half price, and the Fast tier is double. In ChatGPT it comes with paid plans, starting with Plus at $20 a month.

Is GPT-6.1 Sol free?

No. At launch OpenAI made it available in the API and to eligible paid ChatGPT plans in ChatGPT Work and Codex, rolling out by account and workspace. There is no free-tier access in the launch materials, and it is not in the regular ChatGPT chat model picker.

What is GPT-6.1 Sol Ultrafast?

A faster serving mode OpenAI announced at DevDay 2026 as coming soon, with up to 8x faster token generation in Codex. It was not available at launch, and OpenAI had not published its price. Today's Astra Ultrafast bills 8x included usage or 6x purchased credits, but OpenAI hasn't said the Sol version works the same way.

What effort level should I use for GPT-6.1 Sol?

Medium, the default, for most coding: it is Sol's best setting on FrontierCode (50.2% for $0.36), and every higher setting scores lower there. Use high for long engineering tasks, where OpenAI's DeepSWE chart peaks. Use max for terminal work and computer use, where the curves keep climbing.

Why is there no GPT-6.1 Astra?

OpenAI's launch post doesn't mention one. TechCrunch reported that OpenAI dropped a planned GPT-6.1 Astra release over safety concerns raised during internal testing. That is press reporting, not an OpenAI statement I could read directly.

Is GPT-6.1 Sol on the independent leaderboards yet?

Partly. On launch day, Artificial Analysis had it at every effort level and Cognition's FrontierCode had it at every effort. Zapier AutomationBench, Cursor's CursorBench, Datacurve's DeepSWE, tbench.ai, the OSWorld board, ARC Prize, LMArena and Surge's GDP.pdf board had not added it yet, so several numbers on this page come from OpenAI's own charts, and each chart says so.

Sources, method and what I didn't check

Every number on this page comes from a public source captured on launch day, September 29, 2026. Nothing is estimated.

OpenAI's launch charts were read from the chart data embedded in the page, and every point was checked against the rendered chart labels. Artificial Analysis and Cognition numbers come from their published data, re-parsed from the saved pages. Rival effort curves on Artificial Analysis come from my September 22 capture; I re-checked every rival's max point on September 29 and none had changed. Terminal-Bench costs are Artificial Analysis's real cost per task. Every chart is labeled with its one source; none combines two. You can download all 193 data points as a CSV.

Community builds and reactions are attributed to their creators and linked. I didn't rerun them, and I only describe what each creator disclosed. If a leaderboard adds 6.1 Sol later, I'll update this page and its date.

Reuse the data: the compiled dataset is licensed CC BY 4.0. Use it anywhere, including commercially, if you credit Moe Lueker, link back here and to the license, and say what you changed. The underlying benchmark results still belong to the organizations that published them.

Written by Moe Lueker. I test every major model launch on my YouTube channel and build AI tools for creators and small teams, after starting out in mechanical engineering and venture capital. About me. I dictate most of my prompts with Rambleproof, a Mac app I made.

Chart key: 6.1 Sol Astra Opus 5.5 Fable 5.1 6 Sol. Hollow markers on a dashed line are GPT-6 Sol, the generation 6.1 Sol replaces.