Fireworks AI Ember-1 Release
Ember-1 reduces token overhead without sacrificing reasoning quality
Fireworks Research has released Ember-1, a specialized model designed to deliver the quality of Kimi K3 while using 35% to 50% fewer tokens. The model is specifically optimized for agentic workloads and automated coding, where long reasoning traces typically increase costs and latency, especially in multi-turn interactions where context grows quadratically.
Ember-1 was developed to solve the problem of "over-thinking" in reasoning models. While models like Kimi K3 often spend over 90% of their generated tokens on internal reasoning, Fireworks Research found that much of this reasoning is redundant. By training the model to be more efficient—preserving critical self-reflection and error recovery while eliminating unproductive loops—Ember-1 achieves comparable accuracy with a significantly smaller token footprint.
Training methodology and infrastructure
Ember-1 was built using Fireworks Serverless Training, allowing the research team to iterate rapidly without managing GPU provisioning. The development process involved over 50 training experiments and 200 evaluations.
To ensure the token savings translated across diverse workloads, the model was trained on a comprehensive collection of tasks, including:
- Mathematics and coding
- Instruction following and conversation
- Search and tool use
- Software engineering (both standalone problems and extended interactions)
The training focused on on-policy planning and learning guided by task feedback to ensure the model could adapt to observations and outcomes without reverting to excessive reasoning.
Performance benchmarks and the Pareto frontier
Fireworks evaluated Ember-1 using the Specialized Intelligence Index (SII) and several industry benchmarks. The results indicate that Ember-1 establishes a new Pareto frontier—the optimal balance between cost and performance—across several categories:
Bedside Bench (Clinical Cases)
On Doximity’s Bedside Bench, which consists of 500 physician-validated clinical cases, Ember-1 outperformed other models including GPT-5.6 Sol, GPT-6 Astra, and Claude Opus 5 in terms of cost per task.
Industry Coding Benchmarks
Ember-1 demonstrated superior performance compared to Kimi K3 at various reasoning effort settings (Low, High, Max). In several cases, Ember-1 actually improved the pass rate while reducing costs:
| Benchmark | K3 Max Pass Rate | Ember-1 Pass Rate | Token/Cost Reduction vs K3 Max |
|---|---|---|---|
| Terminal Bench 2.1 | 77.6% | 80.9% | -51.9% tokens / -$23.1 USD |
| SWE-bench Verified | 86.0% | 93.2% | -15.5% tokens / -$68.1 USD |
| SWE-Interact | 13.3% | 21.3% | -32.5% tokens / -$60.8 USD |
| DeepSWE 1.1 | 62.8% | 66.4% | -23.7% tokens / -$126.9 USD |
| τ-2 Bench Airline | 64% | 66% | -5.9% tokens / -$0.3 USD |
Real-world validation and production impact
Customer A/B Testing
In live production tests with two customers on coding workloads, Ember-1 reduced token usage by approximately 35% per task while maintaining comparable quality. This resulted in improved or stable downstream metrics for task completion and success rates.
Internal Developer Testing
Fireworks deployed Ember-1 internally for its own developers' daily coding tasks. The company reports that the rollout was "invisible," meaning developers did not notice a drop in quality despite the significant reduction in tokens consumed.
Community and Developer Feedback
Following the announcement, the developer community raised several points regarding the model's positioning and accessibility:
- Pricing Concerns: Some users noted that while the model uses fewer tokens, the price per token may be higher than the base Kimi K3 model, potentially offsetting the total cost savings for some users.
- Weight Accessibility: There is discussion regarding the "open" nature of the model, as Ember-1 is based on open weights but the resulting tuned model is not being released as open weights.
- Prefill vs. Decode: Some technical critics pointed out that in agentic coding, a significant portion of the cost is often in the prefill (input) rather than the decode (output), which is where Ember-1's token reduction is most impactful.
"The most cost optimized way to run K3 is no longer to make it think less, but to run Ember-1, the model that learned to think efficiently."
Ember-1 is currently available as a Research Preview release on Fireworks Serverless.
Sources
- HNEmber-1
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch