Fireworks squeezes 40% of the tokens out of Kimi K3's reasoning — same quality, smaller bill
Fireworks AI's Ember-1 is a post-trained version of Moonshot's open-weight Kimi K3 that reaches the same answers with much shorter reasoning traces — roughly 40% fewer tokens, which could slash the bill for coding-agent workloads. It's API-only and still just a Research Preview, though.

Reasoning models are wonderful at working through hard problems — and terrible at your API bill. Fireworks AI says it has found a middle path: instead of training a brand-new model, it took Moonshot AI’s open-weight Kimi K3 and retrained it to arrive at the same answers while burning roughly 40% fewer tokens. The result, called Ember-1, is built for coding agents and long agentic workflows, where the reasoning trace — the model’s hidden scratch work — can dwarf the actual answer.
The release, announced on Fireworks’ blog on September 23 and promoted by co-founder Dmytro Dzhulgakov on September 27, is less about raw capability than about a second axis of competition: cost per task.
The problem: reasoning models think too much#
Fireworks reports that models like Kimi K3 can spend more than 90% of their generated tokens on internal reasoning rather than on the answer itself. That cost compounds in multi-turn agentic work: every turn replays earlier reasoning back into the model, so context grows roughly quadratically with the number of turns. A long trace from an early step gets re-read — and re-billed — on every later call.
Customers had already tried the obvious fix: turn down the model’s reasoning effort at inference time. It cut the token bill but gave up too much quality — which defeated the point for teams that wanted K3’s coding performance. So Fireworks’s research team attacked the problem in training instead: teach the model to keep only the reasoning that actually changes the outcome, and drop the repetitive loops and dead-end deliberation that don’t.

What Ember-1 actually does#
Ember-1 is not a new foundation model — it is Kimi K3 put through additional reinforcement learning aimed at one behavior: shorter reasoning traces that keep task accuracy. The work reportedly involved more than 50 training experiments and 200-plus evaluations spanning mathematics, coding, tool use, and software engineering. Fireworks says the training ran on its own serverless infrastructure, used its own data, and involved no customer data; the new training algorithms it developed have not been published.
On Fireworks’s own evaluations against Kimi K3 Max, Ember-1 held roughly steady on accuracy while cutting tokens by 15.5% to 51.9%, depending on the benchmark:
| Benchmark | Ember-1 | K3 Max (Fireworks) | Token/cost cut |
|---|---|---|---|
| Terminal Bench 2.1 | 82.0% | 80.9% | −51.9% |
| SWE-bench Verified | 92.2% | 93.2% | −15.5% |
| DeepSWE 1.1 | 75.2% | 66.4% | −23.7% |
| SWE-Interact | 20.0% | 21.3% | −32.5% |
| τ-2 Bench Airline | 66% | 64% | −5.9% |
Does the math check out?#
The most persuasive number isn’t a benchmark: live A/B tests across two production coding customers reportedly showed about 35% fewer tokens per task at comparable quality. In one published run, output tokens fell from 49.3K to 29.9K — a 71.3% drop in reasoning tokens — while the task score barely moved (0.753 vs 0.751). One customer now runs Ember-1 in production.
There’s also an unusually honest cross-check: GenZTech’s independent coding leaderboard scored Kimi K3 at 93.4% on SWE-bench Verified — just 0.2 points from Fireworks’ own 93.2% baseline. That tight agreement makes Ember-1’s 92.2% look credible. It doesn’t extend to every benchmark, and no outside lab has scored Ember-1 yet.
The price per token hasn’t changed: Ember-1 bills at the same rate as K3 — $3 per million input tokens, $0.30 cached, $15 per million output. Every dollar saved comes from generating fewer tokens, not a cheaper rate.
The catch: it’s rented efficiency#
Ember-1 ships as a Research Preview on Fireworks’ serverless platform, with a roughly two-week window of guaranteed access; whether it survives past that depends on demand. The weights, training code, and algorithms stay closed — nobody is self-hosting this one.

That’s the deeper read. Fireworks is an inference company, not a foundation lab, and it just showed that a major cost cut doesn’t require a new base model — a focused post-training pass on someone else’s open weights did the job. Any open model with a wasteful reasoning profile is a candidate for the same treatment, and rivals will copy the playbook.
What to watch#
- Does the preview go permanent? Fireworks tied Ember-1’s fate to usage — a quiet sunset would say more about real demand than any benchmark.
- Does an outside harness confirm the scores? No lab outside Fireworks has evaluated Ember-1 yet — that independent confirmation is still missing.
- Do rivals copy the playbook? If other inference providers ship their own shortened-reasoning SKUs of open models, this becomes a category, not a stunt.
- Does a next-generation base model moot it? Savings on today’s Kimi K3 matter less if the next open release rethinks the reasoning profile itself.
Sources#
- Fireworks Ember-1: Kimi K3 Quality on 40% Fewer Tokens — GenZTech
- Fireworks AI Releases Ember-1: A Post-Trained Kimi K3 That Uses About 40% Fewer Tokens — MarkTechPost
- Fireworks says its Kimi K3 variant cuts reasoning tokens by 40% — Runtime Wire