Introducing Ember-1


Ember-1: half the tokens, similar solutions

Ember-1 is a brand new specialised mannequin from Fireworks Research that delivers Kimi K3’s high quality with 40% fewer tokens. Built on Kimi K3, it discovered to chop pointless reasoning whereas preserving the considering that issues. We examined it on exterior benchmarks, in reside buyer A/B assessments, and on our personal coding and agent workloads, and high quality held up in each setting. Available right this moment, Ember-1 kicks off an ongoing sequence of specialised fashions by Fireworks, formed by what builders need subsequent. Ember is simply the beginning of what you possibly can construct with the Fireworks Training platform.

How Fireworks Research constructed Ember-1

We heard from customers that they wanted K3’s coding capabilities at a decrease price, as a result of its lengthy reasoning traces made automated coding costly at scale. Turning down K3’s reasoning effort did not clear up this. Lower effort settings gave up an excessive amount of high quality. To maintain the standard and reduce the tokens, the mannequin needed to study to purpose extra effectively, and that meant coaching it.

Getting there took critical analysis. Our crew ran greater than 50 coaching experiments and over 200 evaluations, and developed new coaching algorithms alongside the way in which to shorten reasoning with out dropping accuracy. We did all of it on Fireworks Serverless Training. Because we didn’t must provision or handle GPUs, we might launch experiments as quickly as we had an thought, pay just for what we ran, and transfer from analysis to launch in a fraction of the same old time and price.

We skilled throughout a broad set of duties so the token financial savings would carry over to many workloads. We then evaluated Ember-1 on the Specialized Intelligence Index, public benchmarks, and reside manufacturing site visitors to verify it used fewer tokens with no drop in high quality. Ember-1 is Fireworks’ personal mannequin and the primary in a sequence of fashions from Fireworks Research.

The downside: considering fashions assume an excessive amount of

Reasoning fashions like Kimi K3 spend the vast majority of their generated tokens, typically greater than 90%, on inside reasoning reasonably than the reply itself. This considering construction is pricey on a single request, however it will get a lot worse in multi-turn agentic workloads. Every flip replays all prior reasoning again to the mannequin, so context grows roughly quadratically with the variety of turns. Long reasoning traces from early turns get re-read (and re-billed) on each subsequent name.

Is all that reasoning truly obligatory? Our experiments mentioned no. The reasoning Kimi K3 emits is much longer than the duty requires, and the surplus will be eliminated with out touching the reply. This was how we created Ember-1, a cost-effective model of Kimi K3 constructed from specialised intelligence.

From an commentary to a premium mannequin

Not all of K3’s reasoning is wasted. Some of it’s self-reflection: revisiting an assumption, responding to suggestions, or tracing an consequence again to an earlier choice can assist the mannequin recuperate from errors. The alternative is to protect this capability whereas lowering pointless reasoning and escaping unproductive loops. We imagine that studying from duties and setting suggestions can train the mannequin to purpose extra effectively whereas sustaining its capabilities.

For agentic duties, this studying extends throughout the interplay. The mannequin explores attainable actions, incorporates new observations, and refines its reasoning because it progresses. Feedback connects selections to their penalties, encouraging helpful reflection all through the duty.

We carried these insights right into a coaching assortment spanning arithmetic, coding, instruction following, dialog, search, device use, and software program engineering, masking each standalone issues and prolonged interactions to implement adaptation to observations and outcomes. Task suggestions guides on-policy planning and studying, with an emphasis on preserving functionality throughout this vary of settings.

Results on public benchmarks and reside A/B assessments assist this course: throughout seven benchmarks and two clients’ manufacturing site visitors, Kimi K3’s reasoning might be shortened by 35–50% with out sacrificing accuracy. The internalized conduct additionally exhibits restrained token use on unsuccessful makes an attempt, lowering extended, unproductive reasoning.

The Specialized Intelligence Index: Ember-1 units a Pareto frontier for Bedside Bench

Earlier this week, we launched the Specialized Intelligence Index (SII) to benchmark open, closed, and specialised fashions towards real-world duties created by trade specialists.

We evaluated Ember-1 on Doximity’s Bedside Bench, a physician-validated benchmark spanning 500 medical instances throughout 10 specialised classes.

The end result? Ember-1 set a brand new Pareto frontier for Bedside Bench throughout each open and closed fashions together with GPT-5.6 Sol, GPT-6 Astra, and Claude Opus 5 on price/job.

Figure 1: Pareto Frontier from SII on Bedside Bench
Figure 2: Score vs. Duration Chart on Bedside Bench SII
Figure 2: Score vs. Duration Chart on Bedside Bench SII

Evaluating Pareto throughout extra trade benchmarks

We additionally evaluated Ember-1 on the quality-vs-cost frontier throughout another trade benchmarks. We computed per-benchmark price utilizing the general public Kimi K3 API pricing (uncached enter $3/M tokens, cached enter $0.30/M, output $15/M) and plotted it towards go fee for 3 arms: K3 at reasoning effort low, K3 at reasoning effort excessive, K3 at reasoning effort max (default), and Ember-1. Across each benchmark with greater than 50 check samples, Ember-1 sits on or close to the Pareto frontier, matching K3-max high quality at a fraction of the price, and strictly dominating K3-low. We additionally analyzed GPT-6 Astra, Claude Opus-5 and GLM 5.3, and located that Ember-1 was a frontrunner on the Pareto frontier.

Figure 3: Average of 5 Industry Benchmarks on Open and Closed Model Cost/Task
Figure 3: Average of 5 Industry Benchmarks on Open and Closed Model Cost/Task

We took a double-click on the outcomes instantly evaluating Ember-1 to the unique K3, and located the next outcomes:

Industry Benchmarks
N K3 Low K3 High K3 max Ember-1 Ember-1 vs. K3 Max
Terminal Bench 2.1

89

76.4%

77.6%

80.9%

82.0%

-51.9% / -23.1 USD

SWE-bench Verified

500

80.4%

86.0%

93.2%

92.2%

-15.5% / -68.1 USD

SWE-Interact

75

6.7%

13.3%

21.3%

20.0%

-32.5% / -60.8 USD

DeepSWE 1.1

113

55.8%

62.8%

66.4%

75.2%

-23.7% / -126.9 USD

τ-2 Bench Airline

50

64%

64%

64%

66%

-5.9% / -0.3 USD

The most price optimized option to run K3 is now not to make it assume much less, however to run Ember-1, the mannequin that discovered to assume effectively.

Customer validation: Live A/B assessments

Benchmarks solely let you know a lot. Like what we discovered within the Specialized Intelligence Index outcomes, we wished to check the mannequin on extra actual workloads, and to check the mannequin utilizing manufacturing site visitors. The actual check is commonly whether or not the mannequin holds up on manufacturing site visitors, in merchandise customers rely upon.

We ran reside A/B assessments with two clients on their manufacturing coding workloads. In each instances, Ember-1 delivered spectacular token financial savings, roughly 35% fewer tokens per job at comparable high quality. Most of the downstream product metrics held or improved, together with job completion, success scores, and failure charges all shifting in the best course at considerably decrease token price. Following the A/B assessments, one buyer is now operating Ember-1 in reside manufacturing, with plans to scale it as much as change the bottom mannequin solely.

Score Steps Output Tokens Reasoning Token discount Total token discount
Kimi K3

0.751

23.8

49.3K

–

–

Ember-1

0.753

21.4

29.9K

71.3%

39%

Internal validation: Our personal builders did not discover

A big a part of Fireworks’ inside coding/cowork site visitors is powered by our personal inference service. Before any buyer noticed the mannequin, we put Ember-1 to work internally and let our personal builders use it for on a regular basis coding work together with issues like vibe testing at scale on actual duties.

The consequence we’re proudest of: no information. No information is nice information. Developers carried on their coding workloads with out noticing the change, whereas consuming considerably fewer tokens. For a mannequin whose complete worth proposition is “similar solutions, fewer tokens,” an invisible rollout on inside site visitors is the strongest attainable sign.

What’s subsequent

Ember-1 is rolling out as a serving choice alongside the bottom Kimi K3 mannequin as a Research Preview launch on Serverless. To assist the quickly rising open-source ecosystem, we’re introducing analysis releases to provide builders two-week serverless entry to new analysis fashions, making them everlasting based mostly on group demand. For agentic coding and different workloads the place reasoning tokens account for a lot of the price, it delivers the identical high quality at roughly half the token price.

Fireworks Research will proceed to push the frontier of mannequin effectivity by bringing specialised intelligence to extra Ember fashions to allow you to deploy probably the most economical fashions, and cut back your token spend. Token effectivity is changing into a theme of Fireworks.

Looking to take Ember-1 one step additional, and optimize it on your use case? We are additionally launching coaching assist for Ember-1, enabling enterprises to construct personalized, token-efficient fashions tailor-made to their wants with their very own knowledge. The way forward for open fashions is specialised fashions skilled in your particular workload.

Trying Ember-1 out in your workloads? We’d love to listen to about your expertise, so tag us on X (@FireworksAI_HQ) and tell us what you are constructing!



Source link