All News
openaigpt-6-1-solpricingagentic-codingdevdayai-news

GPT-6.1 Sol: near-Astra scores at a fifth of Astra's token price

OpenAI's GPT-6.1 Sol promises near-Astra coding and computer-use scores at a fifth of Astra's token price. We check the per-task math behind the headline.

Vlad MakarovVlad Makarovreviewed and published
6 min read
GPT-6.1 Sol: near-Astra scores at a fifth of Astra's token price

OpenAI released GPT-6.1 Sol on September 29, an upgrade to GPT-6 Sol that the company says "nearly matches" GPT-6 Astra's intelligence on agentic coding, computer use and professional work while charging a fifth of Astra's standard token prices. The launch post prices it at $2 per million input tokens, $0.10 per million cached input tokens and $10 per million output, against Astra's $10, $1 and $50. The cheap tier, GPT-6 Luna, stays at $0.10, $0.01 and $0.50.

That the model exists is not in question. What deserves attention is the arithmetic behind "a fifth of the price", because OpenAI's headline compares a token rate rather than a finished job. GPT-6 Sol and Luna arrived on September 22 with their own round of cuts, and Astra's launch set the $10/$1/$50 anchor this model now undercuts. Every performance figure below was produced by OpenAI, and the column comparing it to Anthropic's Opus 5.5 and Claude Fable 5.1 is one the company assembles from "publicly available reports".

What the scorecard says, and who ran it

OpenAI reports progress across five benchmark families: agentic software engineering, document and business-workflow work, computer use, scientific research and factuality. In each, GPT-6.1 Sol is graded beside GPT-6 Sol and, where relevant, GPT-6 Astra and Opus 5.5. The table below compresses the company's launch tables and the system card addendum.

BenchmarkGPT-6.1 SolGPT-6 SolAstra / Opus 5.5
DeepSWE v1.1, agentic codingmatches Astra at roughly one-fifth the costbeats Sol's best by 6.4 pointsAstra, comparable score
GDP.pdf, complex documentsabove Opus 5.5 with fallbacks, under half the per-task costn/aapproaches Astra at one-fifth per-task cost
AutomationBench 1.0.6, 47 tools2.2 points above Opus 5.5 at medium effort4.8 points up from Sol, same settingOpus 5.5, its cost understated
OSWorld 2.0 offline set7 points above Sol at max effort, under half the costn/awithin 2.1 points of Astra
Terminal-Bench Science 0.1more than double Sol at max effortn/aAstra 68.1%, the top score
Factuality, low reasoning effort7.7% of answers hold an error11.4%Astra within 1.9 points

Where the per-token cut meets the per-task bill

The fifth-of-price line is a rate card. Under real workloads the gap depends on how a task is billed, and OpenAI's own per-task rows are the honest comparison. On Terminal-Bench Science 0.1 at maximum reasoning effort, GPT-6.1 Sol averages $5.47 per task against $23.21 for Opus 5.5 and $23.80 for Astra, which OpenAI describes as "over 75% lower" cost. Astra still posts the highest score on that set, 68.1%, and the launch post says Astra "should be used for the most difficult scientific research tasks". On OSWorld 2.0's offline set, GPT-6.1 Sol lands within 2.1 points of Astra at roughly one-seventh the cost per task.

Two footnotes matter. The cached-input price of $0.10 per million tokens is 95% below the model's own standard input rate and 50% below GPT-6 Sol's cached rate, but a cache only pays off when context repeats across requests; a single-turn workload banks none of that saving, so the discount assumes context reuse, not the average call. And the rival column is not neutral. OpenAI's AutomationBench footnote concedes that the Claude Fable 5.1 datapoint "understates its actual cost, as it omits the cost of fallbacks, which occurred on ~40% of tasks", a reminder that cost rows move with failure handling.

The factuality gain points the right way and is narrow in scope. OpenAI measures the share of answers holding a factual error on de-identified conversations "where users had flagged a factual error from a prior model", then notes these prompts "are not representative of typical usage". The drop from 11.4% to 7.7% at low effort, a reduction of roughly 32%, describes a deliberately hard slice of traffic, not everyday answers.

On safety, the addendum reports no observed attempts to bypass an automated safety reviewer, matching Astra and Sol. On an evaluation where the search tool is broken, GPT-6.1 Sol fails to disclose the problem in 2.1% of cases, against 4.9% for GPT-6 Sol, 1.5% for Astra and 28.7% for GPT-6 Luna. OpenAI classifies the model as Critical in cybersecurity and High for biological and chemical capability, and applies Astra's safeguards stack.

Availability is Codex and Work first

From September 29, GPT-6.1 Sol reaches all Plus, Pro, Business, Enterprise and Edu users in ChatGPT Work and Codex, plus the API under the id gpt-6.1-sol. It is not yet in the consumer Chat surface. The DevDay recap blurs that when it calls the model "available to all API, Plus, Pro, Business, Enterprise, and Edu users". "GPT-6.1 Sol Ultrafast", promised "in the coming days", is slated for up to 8x faster token generation in Codex.

DevDay folded the launch into a bigger number

DevDay 2026 ran the same day with more than 20 announcements, and the launch landed inside that frame, not as a standalone event. The new Ultrafast tier advertises up to 8x faster generation, roughly 300 tokens per second, in Codex and up to 6x in the API; GPT-6 Astra Ultrafast shipped immediately in the API and in ChatGPT Work and Codex on the Pro 500 and Enterprise plans. The new Pro 500 plan carries 25 times the ChatGPT Plus usage allowance. Elsewhere the same day came dots, always-on agents; Codex in the cloud; Codex Security Cloud; a Decisions API; and computer use in the Agents API.

Reception split along familiar lines. On Hacker News the launch post sat at roughly 650 points and close to 600 comments, with much of the thread arguing that per-month spending has outrun what the models actually deliver. Threads on Reddit around the new $500 tier ran similarly negative. CNBC's live blog recorded CFO Sarah Friar defending the ladder of higher tiers, recalling that "when we launched our $200 SKU, people thought we'd lost our minds". The pricing argument, in other words, is not settled by the model card.

What would settle it

The comparison OpenAI did not publish is the one that would decide the case: a third-party, per-task evaluation run on identical prompts, with cache-hit rates and fallback costs reported next to the scores. Until that exists, "a fifth of the price" is a rate card and "near-Astra" is a vendor word, both defensible and both unverified outside OpenAI's harness. Three concrete markers to watch: whether GPT-6.1 Sol Ultrafast actually ships in the coming days, what cache-hit rates developers report once the $0.10 cached price meets production traffic, and whether anyone outside OpenAI reproduces the Terminal-Bench Science gap. Also undisclosed: parameter counts, training corpus, the reinforcement-learning recipe, and any independent evaluation. That silence is standard for a frontier lab, and it is why the per-task rows, not the per-token headline, are the place to look.

Related Articles

Scroll down

to load the next article