Fireworks Research's new Ember-1 model finished reasoning tasks 3.4 times faster than the Kimi K3 model it's based on, while maintaining nearly identical accuracy across logic, scheduling, and probability problems. The company released Ember-1 on September 23 as a research preview, claiming it delivers the same quality as Moonshot's open-weight Kimi K3 using roughly 40% fewer tokens. A new independent test put those claims under pressure with progressively harder challenges designed to reveal whether cutting reasoning tokens hurts performance on complex problems.

The testing ran both models through identical prompts five times each: three logic puzzles of increasing complexity, a deployment scheduling problem requiring a 17-hour optimal solution, and five probability questions about retry systems with exact fractional answers. Kimi K3 solved all 15 test runs perfectly, while Ember-1 completed 14 out of 15 correctly, with one small arithmetic error on a probability calculation. On logic puzzles, Ember-1 averaged 13,630 reasoning tokens compared to Kimi's 16,679—an 18% reduction—and finished in 3 minutes 46 seconds versus Kimi's 12 minutes 26 seconds. For deployment scheduling, Ember-1 used 6,543 reasoning tokens against Kimi's 7,792, completing the task in 1 minute 29 seconds compared to 4 minutes 46 seconds. The probability test showed the widest gap: Ember-1 averaged 6,242 reasoning tokens and 1 minute 47 seconds, while Kimi used 9,682 tokens and took 6 minutes 48 seconds, though Ember made its only mistake here.

Across all tests, Ember-1 consumed 23% fewer reasoning tokens overall and cost $2.48 total compared to Kimi K3's $3.26 when both ran on Fireworks at $3 per million input tokens and $15 per million output tokens—a 24% savings. The report notes that "reasoning tokens bill as output, so a model that thinks less should cost less," but adds a caveat: other providers sell Kimi for as little as $1 per million input tokens and $9 per million output tokens, which would bring the same test runs down to $1.96, cheaper than Ember-1. The tester confirmed every answer key with two independent methods before running either model and logged reasoning tokens separately from the rest of the output to ensure fair comparison.

The report concludes that speed represents Ember-1's clearest advantage, noting Kimi's slowest run took nearly 20 minutes while Ember consistently finished faster. "If you have time and want to pay less, Kimi K3 wins," the analysis states. "If you want almost identical results to Kimi K3 at a much faster rate, use Ember-1." Fireworks says Ember "learned to cut unnecessary reasoning while keeping the thinking that matters," and the testing suggests that strategy works: the token reduction came closest to the marketed 40% claim on the probability test, the most arithmetically intensive challenge. The cost advantage depends entirely on which provider you choose for Kimi K3, since routing to cheaper endpoints on OpenRouter eliminates Ember's pricing edge while preserving Kimi's perfect accuracy record. Organizations choosing between distilled speed and budget-priced thoroughness now have empirical benchmarks showing where each model trades off performance for efficiency. The competitive pressure on reasoning-model pricing may accelerate as providers realize customers can route around premium endpoints without sacrificing quality.