All News
fireworksember-1kimi-k3reasoning-tokensinference-costresearch-preview

Fireworks' Ember-1 claims Kimi K3 quality on 40% fewer tokens, and loses two rows on its own table

Fireworks Research says Ember-1 gives Kimi K3's quality on 40% fewer tokens, but its own table shows the derivative losing to K3 max on two of five benchmarks.

Vlad MakarovVlad Makarovreviewed and published
6 min read
Fireworks' Ember-1 claims Kimi K3 quality on 40% fewer tokens, and loses two rows on its own table

On September 27 the Fireworks AI post "Introducing Ember-1" reached the front page of Hacker News, where the entry stood at 263 points and 145 comments when it was checked the same day. Ember-1 is a specialized derivative of Moonshot's open-weights Kimi K3, built by Fireworks Research and sold next to the base model as a serving option on Fireworks Serverless. The claim is K3's quality on roughly 40% fewer tokens. Every performance number behind that headline comes from Fireworks, and so does the index the launch leads with.

The mechanism Fireworks trained for

Reasoning models spend most of what they generate on thinking, the post argues — sometimes more than 90% of tokens. In multi-turn agent work every turn replays prior reasoning back to the model, so context grows roughly quadratically and early thinking is re-read and re-billed on later calls. Lowering K3's reasoning effort was tried and rejected: "Lower effort settings gave up too much quality", the post says. Ember-1 is the other route — keep the effort, change what the model spends it on.

What went into the run

Fireworks Research reports more than 50 training experiments and over 200 evaluations, run on the company's own Serverless Training product, with new training algorithms developed along the way. The training collection spans mathematics, coding, instruction following, conversation, search, tool use and software engineering, mixing single problems with extended interactions. Task and environment feedback guides on-policy planning, and the target behaviour includes fewer tokens on failed attempts, not just successful ones.

The vendor's own table

Every figure below is Fireworks'. The cost column is the vendor's arithmetic at K3's published API list prices — uncached input $3 per million tokens, cached input $0.30, output $15 — and K3 max is K3's default reasoning setting.

BenchmarknK3 lowK3 highK3 maxEmber-1 (cost vs K3 max)
Terminal Bench 2.18976.4%77.6%80.9%82.0% (-51.9%)
SWE-bench Verified50080.4%86.0%93.2%92.2% (-15.5%)
SWE-Interact756.7%13.3%21.3%20.0% (-32.5%)
DeepSWE 1.111355.8%62.8%66.4%75.2% (-23.7%)
tau-2 Bench Airline5064%64%64%66% (-5.9%)

Two rows go the other way inside the seller's own table. Ember-1 is below K3 max on SWE-bench Verified, 92.2% against 93.2% over 500 tasks, and on SWE-Interact, 20.0% against 21.3% over 75. It is above K3 max on the other three: Terminal Bench 2.1 by 1.1 points, DeepSWE 1.1 by 8.8, tau-2 Bench Airline by 2.

The post's own heading reads "half the tokens, same answers" and its closing line promises "roughly half the token cost", while the figures in between are 40% fewer tokens in the model's description and 39% total-token reduction in the customer test below.

Two customers, one anecdote

The production evidence is two live A/B tests on customer coding traffic, reported as roughly 35% fewer tokens per task at comparable quality, with completion, success scores and failure rates said to hold or improve. One paired result is given: K3 at 0.751, 23.8 steps and 49.3K output tokens against Ember-1 at 0.753, 21.4 steps and 29.9K — a 71.3% cut in reasoning tokens, 39% overall. One customer now runs Ember-1 in production and plans to replace the base model with it.

Before any customer saw the model, Fireworks ran it on its own developers' work and reports that nobody noticed the switch. That is an anecdote from the seller, not a measurement, and no table in the release can carry it.

The headline chart is Fireworks' index

The launch leads with the Specialized Intelligence Index, Fireworks' own index, introduced days earlier, covering security, legal, finance, customer support, software, healthcare and productivity, with several benchmarks credited to outside authors. One is Doximity's Bedside Bench, 500 physician-validated clinical cases, where Ember-1 is said to set "a new Pareto frontier" on cost per task against GPT-5.6 Sol, GPT-6 Astra and Claude Opus 5. A frontier drawn by the vendor, on the vendor's own index, is a claim about a chart that no customer can rerun.

Same rate per token, a clock on the release

On Fireworks' own model pages the two are priced identically: $3.00 per million input tokens, $0.30 cached, $15.00 output, for Ember-1 exactly as for Kimi K3. The saving is token volume, not a lower rate, making the deltas checkable arithmetic: 39% fewer tokens at an unchanged price is about 39% less money — not the "half" the post reaches for.

The divergence is larger on the other side of the release. K3's page links its Hugging Face weights and lists fine-tuning as supported; Ember-1's page carries neither a weights link nor fine-tuning support, while the post points enterprises at Ember-1 training support in the same breath. It is served, not released. There is also a clock: Ember-1 ships as a Research Preview with two weeks of serverless access, made permanent only "based on community demand".

What the thread got to first

One commenter argued that most inference cost in agentic coding sits in prefill rather than decode — the half of the bill a reasoning-token cut does not touch. Another objected that Moonshot shares weights openly while the improvement returns as closed, and Fireworks' page publishes no weights for Ember-1. A third pointed at the Hugging Face group Ukisai, which sells the same kind of thinking-token reduction for Qwen, and which we covered separately. Comments are evidence of what the crowd thinks, not of what the model does.

What would settle it

Three things. A third-party evaluation of Ember-1 against K3 at matched reasoning effort on the same harnesses, run by someone who sells neither. Per-task cost measured independently on production-shaped traffic, not vendor telemetry from two partners. And the survival of the preview past its two-week window. Until then the checkable part of the launch is narrow: an unchanged list price per token, 40% fewer tokens on the seller's own counters, and two rows of the seller's own table where the derivative loses.

Related Articles

Scroll down

to load the next article