Back to blog
AI Tools11 min read

Claude Opus 5.5 vs GPT-6 Astra: Same Prompt, Two Games, Real Cost

Claude Opus 5.5 vs GPT-6 Astra, Sol and Luna on the same two game prompts. Opus built the best games both times, but the runner cost $15.25 to Astra's $7.09.

By Moe Lueker

Claude Opus 5.5 is cheaper per token than GPT-6 Astra. It still cost me more than twice as much to build the same game.

I gave four models one identical prompt, twice: Opus 5.5 and all three GPT-6 sizes, Astra, Sol and Luna. Nobody got help and nobody got fixes. Opus built the best game in both rounds. On the first round it also ran up the biggest bill and took the longest. On the second round, the bill flipped.

When Opus 5.5 and GPT-6 Sol launched on the same day, I expected the real fight to be Opus against Sol. After this test it's clearly Opus against Astra, and the answer is less simple than "use the best one."

How I kept it fair#

Every model got an empty folder with its own git repository, so none of them could see another model's work. The three GPT-6 models each ran in their own Codex session. Opus 5.5 ran in the Claude Code app. I pasted the exact same prompt into all four and walked away.

Two tests:

  • A runner game with a flying level. Simple on purpose.
  • A cozy island builder. More systems, more room to be lazy.

Luna ran at max effort. The other three ran at xhigh on the runner and high on the island. The costs below are the build turn only, pulled from my local Claude Code and Codex logs and priced at API list rates. Follow-up prompts don't count.

This isn't how I build real apps. Normally I go back and forth with a model for hours. One prompt is just the fairest way to see what each model does straight out of the box.

Round 1 runner game lineup from one prompt: GPT-6 Luna built Lumen Run, GPT-6 Sol built Skyline Relay, GPT-6 Astra built Tailwind and Claude Opus 5.5 built Embertail
Round 1 runner game lineup from one prompt: GPT-6 Luna built Lumen Run, GPT-6 Sol built Skyline Relay, GPT-6 Astra built Tailwind and Claude Opus 5.5 built Embertail

Round 1: the runner game#

GPT-6 Luna built Lumen Run. It works, it's easy to play, and it feels a lot like the Chrome T-Rex game. The second level lets you float and press X to open a gate, which is a nice touch. Solid start. Not exciting.

GPT-6 Sol built Skyline Relay. It has sound, the glide feels responsive, and shift bursts you through glass. But it's mostly jumping from platform to platform with no real obstacles.

GPT-6 Astra built Tailwind, and this is where GPT-6 starts to look good. Real art, a parallax background, and a glide level where you can climb while you glide. If I had to build this game with GPT-6, I'd pick Astra.

Claude Opus 5.5 built Embertail, and it wasn't close. You can jump, dive, brake and sprint, which no other model thought of. It added checkpoints, a mushroom boost, a glider you launch by holding the button, and a swing on shift. It even composed its own background music. None of that was in the prompt.

Embertail, the runner game Claude Opus 5.5 built from one prompt: a fox diving through a purple valley at sunset with score popups and a speed meter
Embertail, the runner game Claude Opus 5.5 built from one prompt: a fox diving through a purple valley at sunset with score popups and a speed meter

The code tells the same story. Opus shipped 2,405 lines of game code. Astra shipped 358.

And that's exactly why it cost more.

Runner game build cost: Claude Opus 5.5 is 2.5x cheaper per token than GPT-6 Astra but wrote 7.2x the output, 256K vs 36K tokens, so it cost $15.25 vs $7.09 and took 1 hour 2 minutes vs 24 minutes
Runner game build cost: Claude Opus 5.5 is 2.5x cheaper per token than GPT-6 Astra but wrote 7.2x the output, 256K vs 36K tokens, so it cost $15.25 vs $7.09 and took 1 hour 2 minutes vs 24 minutes

Opus lists at $4 per million input tokens and $20 per million output. Astra lists at $10 and $50. So Opus is 2.5x cheaper per token. But on this build it wrote 256K output tokens to Astra's 36K, about 7x more. Final bill: $15.25 for Opus, $7.09 for Astra. Opus also took an hour and two minutes. Astra finished in 24.

Sol came in at $2.21 in 28 minutes. Luna cost 17 cents, but it took 38 minutes, longer than Astra. The cheapest model wasn't the fastest one.

You pay per task, not per token. A model that does more work per prompt costs more per prompt, even at a lower rate.

Round 2: the cozy island builder#

Round 2 cozy island builder lineup from one prompt: GPT-6 Luna built Little Tide, GPT-6 Sol built Little Haven, GPT-6 Astra built Tiny Tides and Claude Opus 5.5 built Tiny Tides
Round 2 cozy island builder lineup from one prompt: GPT-6 Luna built Little Tide, GPT-6 Sol built Little Haven, GPT-6 Astra built Tiny Tides and Claude Opus 5.5 built Tiny Tides

GPT-6 Luna let you place cottages, pine trees, pebbles, flower patches, harbors and lanterns. That's about all. I wasn't impressed.

GPT-6 Sol had the nicest 3D art of the GPT-6 builds and finished in about nine minutes for 68 cents. As a concept, it's a good start. As a game, it's too lazy.

GPT-6 Astra looked plainer than Sol but had slightly better mechanics: more kinds of trees and houses. Then it stopped. I was hoping for more.

Claude Opus 5.5 built a different kind of game, and yes, it picked the same name as Astra: Tiny Tides. The island grows as you build. There's a night mode where the houses light up. Placement matters: a house next to a path gives you three hearts instead of two. Little boats sail around, little people walk around, and a whale breaches off the coast.

Tiny Tides island builder by Claude Opus 5.5: a grown island with colorful cottages, paths, a windmill, two lighthouses, villagers and a sailboat, with a build toolbar
Tiny Tides island builder by Claude Opus 5.5: a grown island with colorful cottages, paths, a windmill, two lighthouses, villagers and a sailboat, with a build toolbar

Here's the part that surprised me. On this round Opus was the cheaper one.

Island builder build cost: Claude Opus 5.5 $5.40 in 30 minutes vs GPT-6 Astra $6.88 in 27 minutes, and GPT-6 Sol built Little Haven in 9 minutes for $0.68
Island builder build cost: Claude Opus 5.5 $5.40 in 30 minutes vs GPT-6 Astra $6.88 in 27 minutes, and GPT-6 Sol built Little Haven in 9 minutes for $0.68

Opus cost $5.40 in about 30 minutes. Astra cost $6.88 in 27. Opus still wrote about 3x more output than Astra, but not 7x, and at 2.5x cheaper tokens that was enough to flip the bill. So "Opus is the expensive one" isn't a rule. It depends on how much the model decides the task deserves.

What the whole test cost#

Same prompt, different bill: total API cost across both games was Claude Opus 5.5 $20.65, GPT-6 Astra $13.97, GPT-6 Sol $2.89 and GPT-6 Luna $0.23
Same prompt, different bill: total API cost across both games was Claude Opus 5.5 $20.65, GPT-6 Astra $13.97, GPT-6 Sol $2.89 and GPT-6 Luna $0.23

ModelRunnerIslandBoth gamesTotal build time
Claude Opus 5.5$15.25$5.40$20.651 h 32 min
GPT-6 Astra$7.09$6.88$13.9751 min
GPT-6 Sol$2.21$0.68$2.8937 min
GPT-6 Luna$0.17$0.06$0.2359 min

Opus built the best game both times for about $7 more than Astra across the whole test. Whether that's worth it depends on what you're building. For something people will see and play, I think it clearly is.

These are two real builds, not a benchmark suite. If you want the benchmark side, I put Opus 5.5, Astra, Sol and Luna on the same cost-per-task charts at every effort level in Claude Opus 5.5 vs GPT-6 Sol and Luna: Benchmarks and Cost. The short version: neither company puts the other's new model in its launch charts, so you have to line them up yourself.

Putting the winner online#

Most one-prompt demos on X only run on the builder's laptop. I wanted you to be able to play the winner.

The game I put online is Sundown Courier, the Opus 5.5 runner from my launch-day Claude Opus 5.5 review, grown into four levels with a handful of follow-up prompts. You get a dash that breaks glass, a glider, a vine hook you can swing from, a bolt that takes out enemies, and a belt puzzle that opens gates. Those follow-up prompts aren't in any of the cost numbers above.

Sundown Courier, the Claude Opus 5.5 runner game, live on a Hostinger site with a level select for Glasswind Mesas, Skyreach Isles, Thornwood Canopy and Stormcrown Spire
Sundown Courier, the Claude Opus 5.5 runner game, live on a Hostinger site with a level select for Glasswind Mesas, Skyreach Isles, Thornwood Canopy and Stormcrown Spire

Play Sundown Courier here and try to beat my score.

Hosting it with Hostinger#

Hostinger sponsored this video, and it was an easy fit: I already run my AI agents on a Hostinger VPS, including the $8 a month Hermes Agent setup I use every day. For a game or a website you don't need a VPS at all. Regular hosting plus the Hostinger connector is enough.

Here's the whole flow:

  1. Open the Hostinger connector page and pick a plan. I went with Unlimited, which covers unlimited websites and a domain. Code MOELUEKER takes an extra 10% off at checkout.
  2. Connect it to Claude Code. Click "Install for Claude Code," run the install command in your terminal, or open Claude Code, go to Settings, then Extensions, browse, search "Hostinger," and connect it.
  3. Go back to the session where you built the game and say: "Use the Hostinger connector to publish this game to my Hostinger hosting so it's playable online."

That's it. One prompt later the game was live on a Hostinger site, responsive, with every mechanic working. In the Hostinger panel you can then see the site, manage a database if you add one, open the file manager, clear the cache, or point your own domain at it. I wrote up the connector itself in more detail in What Is Hostinger Connector and How to Use It.

Hostinger plan picker showing the Unlimited plan with unlimited websites, a free domain for one year and AI tools next to the Cloud Startup plan
Hostinger plan picker showing the Unlimited plan with unlimited websites, a free domain for one year and AI tools next to the Cloud Startup plan
Host what you build on Hostinger
Opens the Hostinger connector for Claude Code. Pick a plan, connect it, and publish your game or site with one prompt. Code MOELUEKER takes an extra 10% off.

How I split the work now#

Diagram of how Moe uses the models together: Claude Opus 5.5 plans and builds the front end, GPT-6 Astra builds the back end, and each reviews the other's code
Diagram of how Moe uses the models together: Claude Opus 5.5 plans and builds the front end, GPT-6 Astra builds the back end, and each reviews the other's code

After both rounds, this is how I use them:

  • Claude Opus 5.5 for anything people see: front end, design, animations, 3D and games. I've already used it this week to redesign some of my apps.
  • GPT-6 Astra for back end work, long debugging sessions, and finishing a big project without hitting a limit.
  • GPT-6 Sol as the everyday worker when cost matters.
  • GPT-6 Luna for small, repetitive jobs. I wouldn't let it drive a whole app.

The best results come from using them together. Opus plans the project and builds the front end. Astra builds the back end. Then whichever model didn't write the code reviews it. A model checking its own work tends to miss its own small mistakes.

You could get closer results out of the cheaper models with a much more detailed prompt, and that's a fair point. It's also a good way to save money: let Opus write the detailed plan, then hand it to Sol to build. Plans that long are faster to say than to type. I talk mine out with Rambleproof, the Mac dictation app I built and use every day.

Rambleproof dictation app for Mac showing dictated prompts to an AI agent in the history view, with total words, average words per minute and minutes saved
Rambleproof dictation app for Mac showing dictated prompts to an AI agent in the history view, with total words, average words per minute and minutes saved
Talk your prompts with Rambleproof
Push-to-talk dictation for Mac. Hold the shortcut, speak your plan, and get clean text at your cursor in Claude Code, Codex or any app. Free plan available.

And Fable 5.1? I've basically stopped using it. It's expensive, I kept hitting my limits, and Opus 5.5 does the job. Anthropic says Sonnet 5.5 and Haiku 5.5 are coming in the next few weeks. I'm hoping Fable 5.5 isn't far behind.

Pick the model for the job, not a side. The best one here cost the most on one round and the least on the other.

Watch the full walkthrough on YouTube: https://youtu.be/swcMZtSp0Xc

Some links above are affiliate links, so I may earn a commission at no extra cost to you. I only recommend tools I actually use.

Moe Lueker

Moe shares tool walkthroughs and lessons from real projects. Mechanical engineer, then venture capital, now building AI tools for creators and small businesses. More about Moe

claude opus 5.5 vs gpt-6 astraopus 5.5 vs astragpt-6 astraclaude opus 5.5gpt-6 solgpt-6 lunaai codingbuild a game with ai

Get new videos in your inbox

Weekly AI workflows. No fluff.

No spam. Unsubscribe anytime.

Want more guides like this?

Subscribe for new videos every week.

Subscribe on YouTube