<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/focusing-on-post-training-2agg0lufi" -->

---
title: Focusing on Post-Training | daily.dev
description: Fireworks&#x27; Ember-1 model, built by post-training Kimi K3, uses roughly 40% fewer tokens while maintaining comparable quality, achieved by training the model to...
canonical: https://daily.dev/posts/focusing-on-post-training-2agg0lufi
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Focusing on Post-Training | daily.dev
og:description: Fireworks&#x27; Ember-1 model, built by post-training Kimi K3, uses roughly 40% fewer tokens while maintaining comparable quality, achieved by training the model to...
og:url: https://daily.dev/posts/focusing-on-post-training-2agg0lufi
og:image: https://api.daily.dev/og/posts/2agG0lufi.png
og:image:alt: Focusing on Post-Training
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Focusing on Post-Training

**[Sebastian Raschka](https://daily.dev/sources/sebastianraschka)** · 2 min read · 0 upvotes · 0 comments

## Summary

Fireworks' Ember-1 model, built by post-training Kimi K3, uses roughly 40% fewer tokens while maintaining comparable quality, achieved by training the model to produce shorter reasoning traces. This exemplifies a broader argument that post-training existing open-weight LLMs is a better investment of limited budgets than duplicating pre-training efforts for frontier models. The technique connects to token-efficient and budgeted reinforcement learning methods where token usage becomes part of the training objective alongside answer correctness.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://sebastianraschka.com/blog/2026/focusing-on-llm-post-training.html>

## Questions this post answers

### How does Fireworks' Ember-1 model reduce token usage compared to Kimi K3?

Ember-1 uses roughly 40% fewer tokens than Kimi K3 while maintaining comparable quality across evaluations. Fireworks achieved this by post-training Kimi K3 to produce shorter reasoning traces, rather than by pre-training a new model from scratch. The exact training recipe hasn't been fully disclosed, but the approach conceptually resembles token-efficient, budgeted reinforcement learning methods that factor response length into the reward alongside answer correctness.

_Teams weighing efficiency gains against capability tradeoffs in reasoning models can follow cases like this on daily.dev._

### Is it worth spending a multi-million dollar budget to pre-train a new frontier LLM from scratch?

No, for a limited multi-million dollar budget, post-training an existing open-weight model is a better use of funds than duplicating pre-training efforts. Frontier models at 500B parameters or larger require far more than a few million dollars to train competitively, and the current selection of strong open-weight models makes starting from one, then investing in post-training, the more practical path.

_Developers deciding how to allocate compute budgets for custom LLM work can track this reasoning on daily.dev._

## Similar posts on daily.dev

- [Token-budget-aware LLM reasoning: cut costs in 2026](https://daily.dev/posts/token-budget-aware-llm-reasoning-cut-costs-in-2026-cgxkcluom) · Redis · 0 upvotes · 0 comments

---

Tags: [#llm](https://daily.dev/tags/llm), [#reinforcement-learning](https://daily.dev/tags/reinforcement-learning)

[View this post on daily.dev](https://daily.dev/posts/focusing-on-post-training-2agg0lufi)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Focusing on Post-Training","url":"https://daily.dev/posts/focusing-on-post-training-2agg0lufi","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/focusing-on-post-training-2agg0lufi"},"datePublished":"2026-09-27T22:26:18.842Z","dateModified":"2026-09-27T22:27:09.224Z","description":"Fireworks' Ember-1 model, built by post-training Kimi K3, uses roughly 40% fewer tokens while maintaining comparable quality, achieved by training the model to...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/b5dfd702af4ba17d9c3803774d085519?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/b5dfd702af4ba17d9c3803774d085519?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"Sebastian Raschka","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Sebastian Raschka","logo":"https://media.daily.dev/image/upload/t_logo,f_auto/v1/logos/eb22d2e73d074c5598d31588ad16a51c","url":"https://daily.dev/sources/sebastianraschka"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/focusing-on-post-training-2agg0lufi","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"llm,reinforcement-learning","timeRequired":"PT2M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Sebastian Raschka","item":"https://daily.dev/sources/sebastianraschka"},{"@type":"ListItem","position":3,"name":"Focusing on Post-Training"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/focusing-on-post-training-2agg0lufi#faq","mainEntity":[{"@type":"Question","name":"How does Fireworks' Ember-1 model reduce token usage compared to Kimi K3?","acceptedAnswer":{"@type":"Answer","text":"Ember-1 uses roughly 40% fewer tokens than Kimi K3 while maintaining comparable quality across evaluations. Fireworks achieved this by post-training Kimi K3 to produce shorter reasoning traces, rather than by pre-training a new model from scratch. The exact training recipe hasn't been fully disclosed, but the approach conceptually resembles token-efficient, budgeted reinforcement learning methods that factor response length into the reward alongside answer correctness. Teams weighing efficiency gains against capability tradeoffs in reasoning models can follow cases like this on daily.dev."}},{"@type":"Question","name":"Is it worth spending a multi-million dollar budget to pre-train a new frontier LLM from scratch?","acceptedAnswer":{"@type":"Answer","text":"No, for a limited multi-million dollar budget, post-training an existing open-weight model is a better use of funds than duplicating pre-training efforts. Frontier models at 500B parameters or larger require far more than a few million dollars to train competitively, and the current selection of strong open-weight models makes starting from one, then investing in post-training, the more practical path. Developers deciding how to allocate compute budgets for custom LLM work can track this reasoning on daily.dev."}}]}
```

