Reasoning models are wonderful at working through hard problems — and terrible at your API bill. Fireworks AI says it has found a middle path: instead of training a brand-new model, it took Moonshot AI’s open-weight Kimi K3 and retrained it to arrive at the same answers while burning roughly 40% fewer tokens. The result, called Ember-1, is built for coding agents and long agentic workflows, where the reasoning trace — the model’s hidden scratch work — can dwarf the actual answer.

The release, announced on Fireworks’ blog on September 23 and promoted by co-founder Dmytro Dzhulgakov on September 27, is less about raw capability than about a second axis of competition: cost per task.

The problem: reasoning models think too much#

Fireworks reports that models like Kimi K3 can spend more than 90% of their generated tokens on internal reasoning rather than on the answer itself. That cost compounds in multi-turn agentic work: every turn replays earlier reasoning back into the model, so context grows roughly quadratically with the number of turns. A long trace from an early step gets re-read — and re-billed — on every later call.

Customers had already tried the obvious fix: turn down the model’s reasoning effort at inference time. It cut the token bill but gave up too much quality — which defeated the point for teams that wanted K3’s coding performance. So Fireworks’s research team attacked the problem in training instead: teach the model to keep only the reasoning that actually changes the outcome, and drop the repetitive loops and dead-end deliberation that don’t.

An AI coding assistant beside a developer, with the agent's chain of thought visualized as a shrinking spiral of light
Illustration: an AI coding agent whose long reasoning chain is compressed into a tighter spiral. (AI Frontier Post illustration)

What Ember-1 actually does#

Ember-1 is not a new foundation model — it is Kimi K3 put through additional reinforcement learning aimed at one behavior: shorter reasoning traces that keep task accuracy. The work reportedly involved more than 50 training experiments and 200-plus evaluations spanning mathematics, coding, tool use, and software engineering. Fireworks says the training ran on its own serverless infrastructure, used its own data, and involved no customer data; the new training algorithms it developed have not been published.

On Fireworks’s own evaluations against Kimi K3 Max, Ember-1 held roughly steady on accuracy while cutting tokens by 15.5% to 51.9%, depending on the benchmark:

BenchmarkEmber-1K3 Max (Fireworks)Token/cost cut
Terminal Bench 2.182.0%80.9%−51.9%
SWE-bench Verified92.2%93.2%−15.5%
DeepSWE 1.175.2%66.4%−23.7%
SWE-Interact20.0%21.3%−32.5%
τ-2 Bench Airline66%64%−5.9%

Does the math check out?#

The most persuasive number isn’t a benchmark: live A/B tests across two production coding customers reportedly showed about 35% fewer tokens per task at comparable quality. In one published run, output tokens fell from 49.3K to 29.9K — a 71.3% drop in reasoning tokens — while the task score barely moved (0.753 vs 0.751). One customer now runs Ember-1 in production.

There’s also an unusually honest cross-check: GenZTech’s independent coding leaderboard scored Kimi K3 at 93.4% on SWE-bench Verified — just 0.2 points from Fireworks’ own 93.2% baseline. That tight agreement makes Ember-1’s 92.2% look credible. It doesn’t extend to every benchmark, and no outside lab has scored Ember-1 yet.

The price per token hasn’t changed: Ember-1 bills at the same rate as K3 — $3 per million input tokens, $0.30 cached, $15 per million output. Every dollar saved comes from generating fewer tokens, not a cheaper rate.

The catch: it’s rented efficiency#

Ember-1 ships as a Research Preview on Fireworks’ serverless platform, with a roughly two-week window of guaranteed access; whether it survives past that depends on demand. The weights, training code, and algorithms stay closed — nobody is self-hosting this one.

Server racks in a data center with streams of glowing tokens shrinking as they flow, evoking falling compute costs
Illustration: inference costs falling as tokens shrink. (AI Frontier Post illustration)

That’s the deeper read. Fireworks is an inference company, not a foundation lab, and it just showed that a major cost cut doesn’t require a new base model — a focused post-training pass on someone else’s open weights did the job. Any open model with a wasteful reasoning profile is a candidate for the same treatment, and rivals will copy the playbook.

What to watch#

  • Does the preview go permanent? Fireworks tied Ember-1’s fate to usage — a quiet sunset would say more about real demand than any benchmark.
  • Does an outside harness confirm the scores? No lab outside Fireworks has evaluated Ember-1 yet — that independent confirmation is still missing.
  • Do rivals copy the playbook? If other inference providers ship their own shortened-reasoning SKUs of open models, this becomes a category, not a stunt.
  • Does a next-generation base model moot it? Savings on today’s Kimi K3 matter less if the next open release rethinks the reasoning profile itself.

Sources#