Model efficiency · Ember-1 explained
Ember-1 cuts Kimi K3 reasoning tokens, but teams should verify the savings on their own agents
Fireworks says its specialized Kimi K3 derivative uses roughly 40% fewer tokens at comparable quality. The strongest case is multi-turn agent work, where every long reasoning trace can be carried into later turns.
Ember-1 is a Fireworks-hosted mixture-of-experts model built on Kimi K3. It keeps K3's public token prices and 1.04-million-token context window while aiming to reduce unnecessary reasoning. Fireworks reports benchmark and production A/B evidence, but adopters should reproduce completed-task quality, total token use, latency, and failure behavior on their own traces.
- Claimed token reduction
- About 40% versus Kimi K3
- Context
- 1,040,000 tokens
- Serverless price
- $3 input · $0.30 cached · $15 output per 1M tokens
- Launch status
- Research Preview
What Fireworks changed
Fireworks trained Ember-1 from Kimi K3 to shorten reasoning without using a low-effort setting that reduced quality. The company reports more than 50 training experiments and 200 evaluations across mathematics, coding, instruction following, conversation, search, tool use, and software engineering.
The model page lists Ember-1 as a 2.78-trillion-parameter mixture-of-experts model with text and image input, function calling, serverless access, and a 1.04-million-token context window. The model card says fine-tuning is not supported, although the launch post separately describes enterprise training support as a future or customized path.
Why shorter reasoning matters more in agents
A long reasoning trace costs output tokens on the turn that creates it. When an agent sends earlier messages back on later turns, that trace can also become repeated input. The practical bill therefore depends on the whole trajectory, including retries, tool observations, cache behavior, and how much history the harness preserves.
Reducing reasoning can help twice: fewer generated tokens now and less history to carry forward. It can also reduce latency and unproductive loops. Those gains disappear if the shorter trace causes more retries, missed tool calls, or lower task completion, so token counts should be read beside outcomes.
What the benchmark table does and does not prove
Fireworks reports Ember-1 at or near the quality-cost Pareto frontier across several coding, terminal, tool-use, and clinical benchmarks. The launch table shows results from 50 to 500 tasks depending on the benchmark, and its direction is encouraging: shorter reasoning often preserved or improved the reported score.
The launch evidence is still vendor-run or vendor-presented. The Specialized Intelligence Index states that models share a standardized serving configuration and benchmark harness, while some benchmark data or scores can come from partners. Results remain sensitive to the exact harness, reasoning setting, task distribution, retries, cache assumptions, and price date. They do not establish that Ember-1 will match K3 on every private workload.
Production A/B evidence
Fireworks says two customers tested Ember-1 on live coding traffic and saw about 35% fewer tokens per task at comparable quality. One customer reportedly moved it into production, and Fireworks says its own developers did not notice an internal switch during everyday coding work.
These are useful production signals, but the announcement does not identify the customers, publish their task sets, or provide raw traces and confidence intervals. Treat the result as a reason to test, rather than a transferable guarantee.
Run a matched A/B test before switching
- Replay representative coding and agent tasks against Ember-1 and the exact Kimi K3 configuration you use today.
- Keep prompts, tools, harness, timeouts, retry policy, cache settings, and graders fixed between arms.
- Measure completed-task success, accepted-output quality, input and output tokens, cached tokens, wall time, tool calls, retries, and cost per accepted result.
- Read failed trajectories to check whether shorter reasoning removed useful self-correction or reduced unproductive loops.
- Test multi-turn sessions long enough to expose context replay costs, rather than comparing only single prompts.
- Confirm availability and preview terms before depending on the model in production.