OpenAI's GPT-6 Sol and Luna: half the price, same pitch — intelligence is getting cheaper, not smarter
On the same day Anthropic shipped Claude Opus 5.5, OpenAI completed its GPT-6 lineup with two models whose headline feature is the price tag. The benchmarks look fine; the independent read is more complicated.

September 22 was a two-frontier-lab day. While Anthropic put out Claude Opus 5.5, OpenAI announced two additions to the GPT-6 family: Sol, aimed at complex coding, reasoning, and agentic workflows, and Luna, built for focused, high-volume, low-latency tasks. GPT-6 Astra remains the flagship; there is no GPT-6 Terra. The portfolio now spans three tiers, and the interesting part of this launch isn't the tiers — it's the receipts.
The framing, in OpenAI's own announcement, is about access rather than capability: "GPT-6 Astra introduced a new generation of intelligence; these models extend its benefits by making that intelligence more efficient and accessible." Read plainly, that's a promise that frontier-adjacent intelligence is about to cost much less. Whether it is also better is a different question, and the evidence doesn't fully agree with itself.
The naming is worth a pause, because it's the whole product philosophy. Sol and Luna are not "GPT-6 lite" and "GPT-6 nano" — they're workload-shaped models. OpenAI has stopped asking which single model is smartest and started asking which model is cheapest for the job you actually run a million times. That's the same segmentation Anthropic has been pushing from the other direction with its Claude family tiers, and it reflects how the frontier labs now see the market: not one leaderboard, but many workloads, each with its own price-performance curve.
The price halved#
The headline is unambiguous. Both models launch at roughly half the API price of their GPT-5.6 equivalents, and OpenAI credits inference-efficiency improvements for the cut rather than a smaller model.
| Model | Input (per M tokens) | Output (per M tokens) | Predecessor pricing |
|---|---|---|---|
| GPT-6 Sol | $2 | $10 | GPT-5.6 Sol: $4 / $20 |
| GPT-6 Luna | $0.10 | $0.50 | GPT-5.6: $0.20 / $1.20 |
Luna at ten cents per million input tokens sits firmly in the volume-pricing tier that used to belong to distilled small models, except Luna isn't positioned as a small model — it's positioned as the fast one. That distinction matters for the economics of agents, where a planner might fire off hundreds of short calls per task and the per-call overhead was always the binding constraint.
The numbers — OpenAI's version#
On the company's own benchmarks, Sol's story is told in cost per completed task, not in raw leadership — and that choice of denominator is the tell.
| Claim | Figure | Caveat |
|---|---|---|
| AutomationBench | Sol at xhigh effort: 33.2% vs Claude Opus 5 at max: 26.9%; $0.27 per task (~9% of Opus 5's cost) | OpenAI-run eval; "cost per task" mixes price and success rate |
| DeepSWE v1.1 | Sol at max: 68.8%, ~1.1 pts behind Fable 5 at xhigh at ~80% lower cost | Vendor benchmark; the rival ran at a different effort tier |
| DeepSWE v1.1 (Luna) | Luna at max: 66.6% | Vendor benchmark; no independent reproduction yet |
| Factual errors | Sol makes ~half as many as GPT-5.6 Sol on an internal test | Sample drawn from conversations where users flagged errors — error-seeking, not representative traffic |
The reliability claim deserves a second look because it sounds better than it is. Cutting factual errors in half within the subset of conversations users already flagged as wrong is a strange denominator: it measures improvement where the model was already failing, not how often it fails in the wild. It's progress, not a clean bill of health.
The models also inherit GPT-6 Astra's newer communication style and alignment behavior, which OpenAI says means fewer warning circumventions and less coding deception. Both of those are process claims about training that no outside lab has yet verified.
The independent view: cheaper, not smarter#
Here's where the launch gets honest. Independent analysis from Artificial Analysis, published the same day, puts Sol on its AA Intelligence Index and finds only modest gains over GPT-5.6 Sol. More striking: on long-form knowledge-work evaluations, Sol went backward — scoring 1,487 on GDPval-AA v2.1 versus 1,588 for GPT-5.6 Sol, with a drop on AA-Briefcase as well, and presentation quality slipping alongside. To put that in context: GDPval-style evaluations are designed to approximate the knowledge work people actually pay for — research, analysis, drafting — and a hundred-point regression there is the kind of thing that shows up as "the new model feels dumber" in real usage, even while the benchmark sheet says it's faster and cheaper.
That pattern — better at agentic coding tasks, worse at long-form knowledge work — reads like a model optimized for where the money is: coding agents. The honest read, then: at half the price, Sol doesn't need to be smarter to be a better deal. But the story of this launch is cost efficiency, not a universal capability jump, and anyone benchmarking Sol against GPT-5.6 Sol on writing-heavy work may find it has regressed.
Why this matters#
Step back and the September 22 pairing tells the industry's story in miniature. Anthropic launched Opus 5.5 at a higher price with stronger knowledge-work numbers and a safety narrative; OpenAI launched Sol and Luna at half price with an efficiency narrative. One lab is racing capability up, the other is racing price down — and both are aimed at the same customer: the enterprise running agents at scale.
The price-down race has a predictable second-order effect, the one named for the economist behind the Jev launch: cheaper intelligence expands use rather than shrinking the bill. Halving the cost of a coding agent doesn't halve the inference budget; it quadruples the number of agents a team is willing to deploy. That is exactly what OpenAI wants — more tokens flowing through its APIs — and exactly what makes the knowledge-work regression worth watching. If the trade is "better coding agents, worse analysts," the industry is quietly choosing which jobs get automated first.
Where to find them#
Both models are rolling out to Codex and ChatGPT Work for Plus, Pro, Business, Enterprise, and Edu users, with Luna available to Free and Go users in the desktop app. They are not yet in traditional ChatGPT chat, and enterprise admins have to enable them — so if your org's rollout lags, the toggle is probably the reason.
The Luna-in-the-desktop-app detail is the one to watch for consumer strategy. Giving free-tier users Luna but not Sol, and only in the desktop app rather than chat, is OpenAI steering casual users toward its cheapest inference tier while keeping the heavy lifting in the paid products. It's usage-based segmentation dressed up as a feature rollout — and if it works, expect the same treatment for every future mid-tier model.
What to watch#
- Independent reproduction. Every headline number so far comes from OpenAI's own evaluations. Third-party runs on public benchmarks — and on knowledge work, not just coding — will set the real picture.
- The cost-per-task math. Sol's strongest claim is efficiency, not intelligence. If independent cost-per-task measurements hold up, it becomes the default coding-agent model on economics alone.
- The knowledge-work regression. The GDPval-AA drop is the finding OpenAI didn't lead with. If it reproduces, it's the first sign that the frontier labs are trading general capability for workload-specific efficiency.
- What Anthropic does. Opus 5.5 launched the same day at a higher price with stronger knowledge-work numbers. The two launches frame the industry's new axis: OpenAI is racing price down, Anthropic is racing capability up. One of those bets is wrong about what the market wants.
Intelligence is getting cheaper. Whether it's getting better is a question for the independent benchmarks, not the launch post.