OpenAI halves GPT-6 pricing with Sol and Luna, against rates it calls promotional
OpenAI's GPT-6 Sol and Luna halve API prices against promotional GPT-5.6 rates, with selective vendor benchmarks and a same-evening rival left off the tables.

OpenAI released GPT-6 Sol and GPT-6 Luna on September 22, 2026, completing the family that began with GPT-6 Astra on September 3. The pitch is not a smarter model. OpenAI says the two were "trained with similar methods as GPT-6 Astra," and the news is arithmetic: GPT-6 Sol costs $2 per million input tokens and $10 per million output, down from $4 and $20 for GPT-5.6 Sol, while GPT-6 Luna falls to $0.10 and $0.50 from $0.20 and $1.20. In its launch post, OpenAI leads with cost efficiency rather than capability — unusual for a frontier lab, and a reason to read its numbers carefully.
Half off a price that was already a discount
The headline cut is 50%, and against the rates OpenAI compares to, it holds. It is also measured against promotional pricing. GPT-5.6 Sol's standard list price was $5 input and $30 output; an OpenAI spokesperson told The New Stack that the GPT-5.6 rates were always promotional and the GPT-6 prices are the default. Coursiv's analysis does the arithmetic the launch post skips: against standard rates, GPT-6 Sol is about 60% cheaper on input and roughly 67% cheaper on output. A discount off a temporary price is real for anyone who was paying it; it is not a claim of a 60% list-price cut, and both numbers are now circulating.
| Model | Input / output per 1M | Cached input | Cache write |
|---|---|---|---|
| GPT-6 Sol | $2 / $10 | $0.20 | $2.50 |
| GPT-6 Luna | $0.10 / $0.50 | $0.01 | $0.125 |
| GPT-6 Astra | $10 / $50 | unchanged | unchanged |
What you get for the money
Both new models carry a 1,050,000-token context window and a 128,000-token maximum output, with six reasoning-effort levels: none, low, medium as the default, high, xhigh and max. Above 272,000 input tokens a long-context surcharge applies — $4 and $15 for Sol, $0.20 and $0.75 for Luna. Fast mode doubles the price, batch and Flex halve it, and regional data residency adds 10%. Credit rates in ChatGPT mirror the token cut, with Sol at 50 credits per million input tokens and 250 per million output, and Luna at 2.5 and 12.5.
Availability is broad but incomplete. Sol and Luna ship in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu, with administrators required to enable them; the API ids are gpt-6-sol and gpt-6-luna. Luna also reaches Free and Go users in the desktop app. Neither appears in the main Chat surface yet, and the rollout is gradual. There is no GPT-6 Terra. GPT-5.5 retires from ChatGPT, ChatGPT Work and Codex on October 14, 2026, though not from the API.
The benchmarks OpenAI chose to publish
Every figure below is OpenAI's, from the launch post: vendor numbers, with rival scores drawn from what the post calls publicly available reports.
| Benchmark | GPT-6 Sol | Comparison OpenAI publishes | Cost per task |
|---|---|---|---|
| AutomationBench 1.0.6 | 33.2% (xhigh) | Astra 30.3% (low) | $0.27 vs 3.9x |
| Agents' Last Exam V1 | 56.4% (max) | above Opus 5's highest | 60% lower |
| DeepSWE v1.1 | 68.8% (max) | Fable 5 at 69.9% (xhigh) | about 80% lower |
| OSWorld 2.0 offline | 60.5% (xhigh) | Opus 5 at 60.3% (medium) | about 80% lower |
AutomationBench 1.0.6, a 47-tool benchmark run by Zapier, has Sol at xhigh scoring 33.2% at $0.27 per task against Astra at low on 30.3% for 3.9 times the cost, Claude Opus 5 at max on 26.9% for 11.1 times, and Claude Fable 5.1 with an Opus 5 fallback at 31.4% for more than 8.9 times. A footnote concedes that Fable 5.1's cost omits its fallback runs, which occurred on roughly 40% of tasks — an omission that flatters the gap. On Agents' Last Exam, Sol reaches 56.4% at max, above Opus 5's highest score at 60% lower cost per task, though OpenAI never states Opus 5's effort level. DeepSWE puts Sol's 68.8% within 1.1 points of Fable 5's 69.9% at about 80% less per task, and Luna's 66.6% lands near Opus 5 and Fable 5 at medium effort for 93% and 96% less. On FrontierCode 1.1 Main, OpenAI says Sol matches Fable 5.1 at xhigh but publishes no score.
Two non-benchmark claims remain. On an internal eval built from de-identified conversations where users flagged errors, OpenAI says Sol makes about half as many mistakes as GPT-5.6 Sol — with the caveat that those conversations are not representative of typical use. On alignment, it reports lower rates of misleading claims about its own coding work in adversarial scenarios.
What the tables leave out
The skepticism here is structural, not personal. OpenAI publishes selected comparisons instead of a full table; rival scores arrive "from publicly available reports" under harnesses that may differ from its own; Claude Fable 5 stands in wherever Fable 5.1 was unavailable; and no row carries a confidence interval. The most conspicuous absence is chronological: Claude Opus 5.5, which Anthropic released the same evening at $4 and $20, appears nowhere in OpenAI's comparisons — the rival that most undercuts a price story is the one missing from the price tables. Independent measurement is absent too: Artificial Analysis had no model pages for Sol or Luna at launch, so every capability figure above is downstream of OpenAI's harness choices, the same structural problem we described when Astra shipped a selective benchmark set.
For builders, the model id is not the only change
Both model pages direct built-in tools and function calling to the Responses API. In Chat Completions, function calling works only when reasoning effort is set to none — a pipeline cannot simply swap the model id, particularly if it relies on tool use at a higher effort level. Caching moves in the builder's favour: GPT-6 gets higher default cache hit rates, cached reads are discounted 90%, and changing effort level or tools no longer breaks the cache. OpenAI says GitHub reports these improvements cut the share of prompt tokens needing fresh processing by more than 50% across billions of requests. It also cites provider-level cost context: valued at API prices, daily token usage inside the company exceeds $600 for the median researcher and $7,000 at the 90th percentile.
What would settle it
Three ordinary things. Independent evaluation from a party that runs the same tasks on a disclosed harness, placing Sol and Luna beside Opus 5.5, Fable 5.1 and Fable 5 on one surface. A full benchmark table rather than a curated one, with confidence intervals and the fallback runs counted. And a price comparison stated against standard rates as well as promotional ones. None of that is exotic; OpenAI has the data. Until then, the honest summary of September 22 is narrow and real: a capability class got meaningfully cheaper, and the company announcing it chose the baseline that made the cut sound largest. Treat "half the price" as a promotional rate measured against a promotional rate, and check which one you were actually paying.


