GPT-6 Sol and Luna Ranked: The New OpenAI Models Compared

2026-09-23
GPT-6 Sol and Luna bring Astra-class ability to cheaper tiers. Here's the full GPT-6 lineup ranked by price, benchmarks and everyday value.
OpenAI added two models to the GPT-6 family on September 22, 2026, and the timing tells you most of what you need to know. Barely three weeks after shipping GPT-6 Astra as its smartest release yet, the company went back to the same workbench and carved out two cheaper versions. GPT-6 Sol and GPT-6 Luna aren't upgrades to Astra. They're a way to get most of Astra's ability without paying Astra's bill.
That raises a practical question. If you build with an OpenAI GPT-6 model, or you just want the better ChatGPT update, which of the three tiers actually belongs in your day? This ranking sorts the lineup by what each model is good at, not by which one looks best on a slide.
The criteria here are narrow on purpose. I weighed four things: API price per million tokens, benchmark results on coding and agentic tasks, the per-task cost against the closest rival, and how each model is handed out to real users. Capability matters, but cost per useful answer matters just as much, because that's the number that shows up on your invoice.
The GPT-6 Family at a Glance
The fastest way to read this AI model update is to line all three models up together. Astra sits on top for sheer capability. Sol is the middle tier built for professional work. Luna is the budget option for volume.
| Rank | Model | Key Feature | Why It's Here |
|---|---|---|---|
| 1 | GPT-6 Astra | Highest overall capability and strongest computer use | Flagship; the pick when quality outranks cost |
| 2 | GPT-6 Sol | Near-Astra reliability at half the old price | Best balance of power and cost for coding and agents |
| 3 | GPT-6 Luna | Cheapest per token, fastest responses | Best for high-volume, cost-sensitive work |
Astra takes the top slot because OpenAI itself still calls it its best model in every respect, built for the hardest and most important projects. Sol and Luna are ranked by how much usable capability they deliver per dollar. That's the difference most readers will feel, and it's the one worth arguing about.
Price runs through all three. Against the promotional rates for the GPT-5.6 line, Sol and Luna cut API costs by roughly half. Sol now lands at $2 per million input tokens and $10 per million output. Luna drops far lower, to $0.10 input and $0.50 output. OpenAI credits better caching and inference efficiency, and says it's passing the savings straight to developers instead of pocketing them.
GPT-6 Astra: Still the Ceiling
Astra keeps first place, and the reason is simple. When the people who made all three models tell you to reach for Astra if you want the best results without compromise, that settles the capability question. It's aimed at the projects where a wrong answer is expensive and you'd rather not find out the hard way.
Launched September 3, 2026, Astra was pitched as OpenAI's smartest and most aligned model to date, with fresh ground covered in computer and browser use, software engineering, cybersecurity, science, and professional work. It also remains the model OpenAI says is strongest at operating a computer. Sol and Luna inherit that ability at a lower price rather than beat it.
The catch is cost. Astra's API rates sit well above the two newcomers, so running it on every routine request in a busy service gets expensive fast. Think of it as the specialist you bring in when a job genuinely needs top-tier reasoning and computer use, while Sol or Luna handles the daily grind around it.
If the best possible output is the only thing that matters, Astra is still the answer. For everyone else, the next two entries are where the real choices live.
GPT-6 Sol: The Best Balance of Power and Cost
Sol takes second overall and first for value, because it pulls off the hardest trick of the three. It gets close to flagship reliability while charging about half of what the previous generation cost. OpenAI describes its error rate on real-world conversations as roughly half of its predecessor, and that gap is the kind that matters once you deploy a large language model on work a customer will actually see.
The benchmarks back it up. On AutomationBench, a cross-app business process test, Sol at extra-high reasoning effort outperforms Claude Opus 5 while costing about 9% of what Opus 5 charges per task. On the agentic Last Exam, Sol reaches 56.4% at maximum effort, above Opus 5's best score, again at roughly 60% lower per-task cost. For coding, DeepSWE v1.1 puts Sol at 68.8%, a whisker under Claude Fable 5's 69.9% top mark, for about 80% less per task.
Computer use tells the same story. On OSWorld 2.0 offline, Sol at extra-high effort hits 60.5%, close to Claude Opus 5's 60.3% at medium effort, while costing roughly 80% less per task. On FrontierCode 1.1 Main, OpenAI says Sol improves clearly over GPT-5.6 Sol and matches Claude Fable 5.1 at extra-high effort for less money.
That package is why Sol outranks Luna. It's the model you pick when you need an AI chatbot or a coding agent to be accurate and dependable, but you've also got a budget to protect. In the GPT-6 tier structure, it's the default for professional work, programming, and agent-style automation, the jobs where a mistake costs more than the tokens did.
GPT-6 Luna: The Cheapest Path to Chatbot Scale
Luna ranks third on raw capability and first on value for volume, and for plenty of teams that's the row that decides the purchase. Its API pricing is the headline: $0.10 per million input tokens and $0.50 per million output tokens. At that rate you can push large language model work through at a scale that would be reckless on a premium tier, and the model still counts as serious artificial intelligence rather than a toy.
The performance isn't filler, either. At high reasoning effort, Luna improves 5.4% over its predecessor while cutting per-task cost by 58% on the AutomationBench cross-app test. Better results at a lower bill is an unusual pairing, and it's the real selling point here. On computer-use tasks it inherits the same approach OpenAI built into Astra, delivered at a fraction of the cost. It also picks up Astra's cleaner communication style, so the answers read better than a stripped-down model usually manages.
On software engineering, Luna reaches 66.6% on DeepSWE v1.1 at maximum effort, landing near the medium-effort results of pricier rivals while costing a small fraction of what they charge. For a model positioned as the budget pick, that's a more-than-respectable showing.
The catch is raw capability. Luna trails Sol on the hardest reasoning and coding benchmarks, which is exactly why it sits below it. But for high-frequency jobs like bulk content, support drafts, and automation that runs all day, Luna is the OpenAI model that keeps the math sane. If your pipeline fires thousands of requests an hour, that difference in per-token price stops being a rounding error and starts deciding what you can afford to ship.
Which Tier Fits Your Work
The decision usually comes down to one question: does the task need the best answer, or a good answer at the lowest cost?
If you're writing code, chaining agent steps, or handling anything where a mistake hurts, Sol is the pick. Its near-Astra reliability at half the previous price is the whole point of this OpenAI update.
If you're generating content at volume, running routine automation, or watching a tight budget, Luna wins on economics. That 50% cut across the API makes it the natural fit for workloads that scale.
And if the job genuinely sits at the edge of what generative AI can do, Astra stays the model to reach for. Ranked by raw power it's first. Ranked by value it slips, and that trade is the entire story of the GPT-6 lineup.
A practical middle path exists, too. Many teams will run Sol as the default and route only their heaviest, most error-sensitive work to Astra, while sending routine or bulk jobs to Luna. OpenAI's per-model pricing makes that split cost-effective in a way it simply wasn't a generation ago.
If you're not wiring up an API at all, the free and Go tiers still matter. Free and Go users get GPT-6 Luna on the desktop app, while Plus, Pro, Business, Enterprise, and Edu users get both Sol and Luna inside ChatGPT Work and Codex. One wrinkle worth flagging: neither model is live in the standard ChatGPT chat window yet, so that rollout is still coming. Until then, the ChatGPT app on Android is the easiest place to poke at the new models for yourself.
How Sol and Luna Compare to the Wider Field
Ranking the GPT-6 pair only means something against the rest of the market, because developers rarely choose an OpenAI model in a vacuum. Against Anthropic's current line, the pattern holds steady: Sol lands close to the top tier on the benchmarks that count, then undercuts it sharply on cost. It's the same playbook OpenAI used with the earlier Sol and Luna, just pushed further.
It isn't a clean sweep, either. On some frontier scores, rivals still edge ahead, which is exactly why Astra remains first for the hardest work rather than Sol. The honest read is that the GPT-6 pair competes on value and reliability for everyday tasks far more than on absolute peak ability.
Anyone weighing the latest AI models faces a narrower question than the marketing suggests. Which tier does the task actually need? If a cheaper model lands within a few points on your own workload, the premium one rarely pays for itself. That's the trade this release was engineered around, and it's the one worth testing before you commit a budget.
Final Verdict: The GPT-6 Tier List
Ranked purely on capability, the order is Astra, then Sol, then Luna. Ranked on value for money, that order flips almost entirely, and that reversal is the actual news. OpenAI didn't replace its flagship. It spread the flagship's best traits down into two cheaper tiers, then priced them to pull developers toward the middle.
For most readers, Sol is the model worth trying first. It carries the coding and agentic strengths that made GPT-6 Astra notable, at a cost that won't punish you for leaning on them. Luna is the one to reach for once you know a task works and you want to run it a million times without watching the meter.
The cleanest move is to test both against your own work. Download the ChatGPT app on your Android phone, see which GPT-6 model fits how you actually build, and let your own results settle the ranking.