DeepSeek V4.1 Flash vs GPT-6 Astra: 98% of the score, 1.4% of the cost
Arena-measured data put DeepSeek V4.1 Flash within two points of GPT-6 Astra on design output at 1.4% of the cost, and the ranking flips once speed is weighted.

On September 9, a link post on r/singularity announced that DeepSeek V4.1 Flash had reached "98% of Astra's score at 1.4% of cost on OpenDesign Arena." The thread drew hedged engagement — roughly 790 points, more than 100 comments — and sent readers to a leaderboard most had never opened. Both figures in the headline are real. Neither is the verdict it looks like.
What the arena actually scores
OpenDesign Arena ranks 13 design and coding models on front-end artifacts across five dimensions. Only two of them feed the headline number: requirement fulfillment, worth 30 points, and design quality, worth 70. Cost per artifact is estimated from official list prices multiplied by measured token usage, and it never enters that average. That gap is the whole trick of the Reddit claim. DeepSeek V4.1 Flash scores 81.2 against GPT-6 Astra's 82.7 — 98% — at $0.023 per artifact against $1.61, which is 1.4%. Both ratios are correct arithmetic on the score and the price, and neither says anything about whether the output was finished enough to use.
| Model | Avg score | Cost/artifact | Time | Delivery rate |
|---|---|---|---|---|
| GPT-6 Astra | 82.7 | $1.61 | 11.1 min | 60.0% |
| DeepSeek V4.1 Flash | 81.2 | $0.023 | 5.3 min | 57.7% |
| Claude Fable 5.1 | 80.3 | $3.66 | 12.8 min | 56.7% |
| GPT-5.6 Sol | 77.6 | $0.544 | 3.2 min | — |
| DeepSeek V4 Flash | 70.6 | $0.031 | 11.1 min | — |
Those two dimensions are not the whole profile the site draws. The arena also plots delivery rate, completion time, and cost efficiency — "different shapes mean different strengths," its copy notes — and it filters all five by scenario, from landing pages to dashboards. The spread across the 13 models is wide on every axis. Hunyuan H4 Preview takes 29.2 minutes per artifact, the slowest in the field, while Muse Spark 1.3 finishes in 3.3; Kimi K3 anchors the quality floor at 65.7. Neither model in the Reddit headline is extreme on any axis except price.
Astra holds the axes that decide hand-off
Astra keeps the average, 82.7 to 81.2, and the delivery rate, 60.0% to 57.7% — the share of runs that produced something ready to hand over at all. Designers also rated its output higher on layout, readability, hierarchy, and style fit: 56.2 out of 70 against DeepSeek's 52.8. The arena is careful with that delivery figure, calling it "an observed sample rate, not a vendor guarantee." Read together, the two leaders are close, and both leave roughly two in five runs short of the finish line. On requirement fulfillment DeepSeek actually leads, 28.4 to 26.5: it follows the brief more faithfully, then loses points on how the result looks.
DeepSeek takes price and speed
The cost gap is not marginal. At $0.023 per artifact against $1.61, DeepSeek V4.1 Flash runs about 70 times cheaper, and it finishes in 5.3 minutes against Astra's 11.1. Two caveats come from the arena's own notes: cost estimates "are not invoices and exclude unrecorded charges," and completion time "excludes platform queue time." Still, on a task you intend to generate in volume, the price difference is the line that decides the budget. The rank below the leaders tells the same story from the other direction: Claude Fable 5.1 places third on raw quality at 80.3 and falls to 12th under the site's own weighted view, because it costs $3.66 a run. The floor of the leaderboard is Kimi K3 at 65.7.
The leaderboard belongs to a vendor
Say plainly who runs this: OpenDesign (open-design.ai) sells a design workspace and desktop app, and the arena page ends in a download prompt for it. That does not make the numbers fake — the published method notes are unusually explicit — but it does make them arena-measured rather than independent in the academic sense. Those notes are specific where they go: each task carries "six to nine functional checks" rated met, partly met, or not met and combined into the 30-point requirement score, while design quality comes from human designers scoring layout and readability, hierarchy, style fit, color and contrast, and image relevance. What the page never publishes is the task set, the rubric, or the per-artifact scores, which is why nobody outside OpenDesign can audit the ranking from the raw data. Treat the results as vendor-published, if unusually transparent.
The ranking flips with the weights
The arena includes a "Choose by preference" view that lets you weight quality, cost, and speed. At its default split of 50/30/20, DeepSeek V4.1 Flash ranks first at 78.8. GPT-6 Astra falls to eighth at 57.5, and Claude Fable 5.1 lands 12th at 52.1 — below models it beats on every raw score. The page states that changing preferences "affects recommendation rankings, not the original evaluation results." Both orderings describe the same measurements, which is the honest caveat the Reddit headline skipped.
How to read the claim
What the post gets right is the price, and price is the axis that decides whether a team can afford a thousand landing pages instead of ten. What it flattens is the gap that matters: this is cost per artifact on a narrow slice of front-end design work, not general capability, and the cheaper model still loses on delivery and on how the output looks. DeepSeek's shipping specs — a 552B backbone MoE, a one-million-token context, an MIT license badge — sit in its model card; the companion piece on that release covers them. For a team generating marketing pages at volume, DeepSeek V4.1 Flash is the rational default. For client work where a designer signs off on the hand-off, Astra's edge on quality and delivery still decides.
What would settle the argument is what the arena cannot supply: an independent run of the same tasks on a harness someone else controls. Until that exists, the honest version of the Reddit headline is narrower than it reads — a cost-per-artifact gap on front-end design work, measured by the vendor whose workspace the models are being measured for.


