<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/thinking-of-ace-we-can-do-it-with-fewer-tokens-0stmkpisz" -->

---
title: Thinking of ACE? We Can Do It with Fewer Tokens | daily.dev
description: IBM Research compares ALTK-Evolve and ACE (Agentic Context Engineering), two systems that let LLM agents learn from their own past trajectories without weight...
canonical: https://daily.dev/posts/thinking-of-ace-we-can-do-it-with-fewer-tokens-0stmkpisz
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Thinking of ACE? We Can Do It with Fewer Tokens | daily.dev
og:description: IBM Research compares ALTK-Evolve and ACE (Agentic Context Engineering), two systems that let LLM agents learn from their own past trajectories without weight...
og:url: https://daily.dev/posts/thinking-of-ace-we-can-do-it-with-fewer-tokens-0stmkpisz
og:image: https://api.daily.dev/og/posts/0STmkPISz.png
og:image:alt: Thinking of ACE? We Can Do It with Fewer Tokens
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Thinking of ACE? We Can Do It with Fewer Tokens

**[Hugging Face](https://daily.dev/sources/huggingface)** · 8 min read · 1 upvotes · 0 comments

## Summary

IBM Research compares ALTK-Evolve and ACE (Agentic Context Engineering), two systems that let LLM agents learn from their own past trajectories without weight updates. Both systems avoid compressing learned lessons into summaries, instead maintaining countable guidelines or playbooks. The key difference is delivery: ACE injects its full playbook on every inference step, while ALTK-Evolve selectively retrieves only the guidelines relevant to the current task and model capability. On the AppWorld benchmark with 168 tasks, ALTK-Evolve achieves 89.3 TGC vs ACE's 80.4 on DeepSeek-V3.2 at ~40% of ACE's token cost (263K vs 634K tokens/task). On the weaker gpt-oss-120b model, accuracy is roughly tied (56.0 vs 54.8 TGC) but ALTK-Evolve uses about one-seventh the tokens (116K vs 777K). The post argues that calibrated delivery — matching the volume of injected guidelines to what a given model can actually absorb — is the key to both efficiency and accuracy gains.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://huggingface.co/blog/ibm-research/altk-evolve-sldd>

## Questions this post answers

### How does ALTK-Evolve compare to ACE (Agentic Context Engineering) on token cost and accuracy for LLM agents?

On AppWorld with DeepSeek-V3.2, ALTK-Evolve scores 89.3 TGC / 80.4 SGC versus ACE's 80.4 / 73.2, using 263K tokens per task versus ACE's 634K, about 40% of the cost. On gpt-oss-120b, ALTK-Evolve edges ACE 56.0 to 54.8 TGC while using only 116K tokens versus ACE's 777K, roughly one-seventh the cost. The difference comes from delivery: ACE injects its full playbook every step, while ALTK-Evolve selectively retrieves guidelines per task.

_Developers weighing agent memory approaches can track cost-versus-accuracy comparisons like this one on daily.dev._

### What is the difference between ACE's playbook approach and ALTK-Evolve's guideline retrieval for agent memory?

ACE consolidates lessons into one comprehensive playbook via a Generator-Reflector-Curator loop and injects the whole thing at every inference step regardless of model or task. ALTK-Evolve clusters and merges near-duplicate lessons support-conservingly into typed, retrievable guidelines, then delivers a small fixed core plus a per-task selected subset, or the full set only when a model has the headroom to use it.

_Teams designing agent memory pipelines can compare consolidation and delivery strategies discussed on daily.dev._

### Does giving an LLM agent more context or memory always improve its task accuracy?

No. On gpt-oss-120b, a weaker model, injecting ACE's full playbook helped on easy and medium AppWorld tasks but selective, per-task retrieval of guidelines won on hard tasks and on the aggregate score, because a large context can overwhelm a weaker model rather than help it. On the stronger DeepSeek-V3.2 model, more delivered lessons kept helping instead of crowding each other out.

_Engineers tuning how much context to feed an agent can follow benchmark findings like these on daily.dev._

## Similar posts on daily.dev

- [How Much Memory Does Your Agent Actually Need?](https://daily.dev/posts/how-much-memory-does-your-agent-actually-need--w9bapy7xf) · Hugging Face · 1 upvotes · 0 comments
- [Researchers Introduce ACE, a Framework for Self-Improving LLM Contexts](https://daily.dev/posts/researchers-introduce-ace-a-framework-for-self-improving-llm-contexts-isr1ouohi) · InfoQ · 1 upvotes · 0 comments
- [Analytics Context Engineering for LLM](https://daily.dev/posts/analytics-context-engineering-for-llm-aapm7ljoy) · Cisco · 2 upvotes · 1 comments
- [Your Agent Aced the Task. Will It Do It Again?](https://daily.dev/posts/your-agent-aced-the-task-will-it-do-it-again--yetkhtezb) · Hugging Face · 0 upvotes · 0 comments
- [From Prototype to Profit: Solving the Agentic Token-Burn Problem](https://daily.dev/posts/from-prototype-to-profit-solving-the-agentic-token-burn-problem-3am1jbiyc) · Towards Data Science · 1 upvotes · 0 comments

---

Tags: [#ai-agents](https://daily.dev/tags/ai-agents), [#rag](https://daily.dev/tags/rag), [#context-engineering](https://daily.dev/tags/context-engineering)

[View this post on daily.dev](https://daily.dev/posts/thinking-of-ace-we-can-do-it-with-fewer-tokens-0stmkpisz)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Thinking of ACE? We Can Do It with Fewer Tokens","url":"https://daily.dev/posts/thinking-of-ace-we-can-do-it-with-fewer-tokens-0stmkpisz","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/thinking-of-ace-we-can-do-it-with-fewer-tokens-0stmkpisz"},"datePublished":"2026-08-11T13:37:31.075Z","dateModified":"2026-09-13T21:58:13.798Z","description":"IBM Research compares ALTK-Evolve and ACE (Agentic Context Engineering), two systems that let LLM agents learn from their own past trajectories without weight...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/6a7c5e02fc3fc12792d9421941b76bd0?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/6a7c5e02fc3fc12792d9421941b76bd0?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"Hugging Face","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Hugging Face","logo":"https://media.daily.dev/image/upload/t_logo,f_auto/v1/logos/f1f55c67d81a4330acf5b90b26b0c8e1","url":"https://daily.dev/sources/huggingface"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/thinking-of-ace-we-can-do-it-with-fewer-tokens-0stmkpisz","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"ai-agents,rag,context-engineering","timeRequired":"PT8M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Hugging Face","item":"https://daily.dev/sources/huggingface"},{"@type":"ListItem","position":3,"name":"Thinking of ACE? We Can Do It with Fewer Tokens"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/thinking-of-ace-we-can-do-it-with-fewer-tokens-0stmkpisz#faq","mainEntity":[{"@type":"Question","name":"How does ALTK-Evolve compare to ACE (Agentic Context Engineering) on token cost and accuracy for LLM agents?","acceptedAnswer":{"@type":"Answer","text":"On AppWorld with DeepSeek-V3.2, ALTK-Evolve scores 89.3 TGC / 80.4 SGC versus ACE's 80.4 / 73.2, using 263K tokens per task versus ACE's 634K, about 40% of the cost. On gpt-oss-120b, ALTK-Evolve edges ACE 56.0 to 54.8 TGC while using only 116K tokens versus ACE's 777K, roughly one-seventh the cost. The difference comes from delivery: ACE injects its full playbook every step, while ALTK-Evolve selectively retrieves guidelines per task. Developers weighing agent memory approaches can track cost-versus-accuracy comparisons like this one on daily.dev."}},{"@type":"Question","name":"What is the difference between ACE's playbook approach and ALTK-Evolve's guideline retrieval for agent memory?","acceptedAnswer":{"@type":"Answer","text":"ACE consolidates lessons into one comprehensive playbook via a Generator-Reflector-Curator loop and injects the whole thing at every inference step regardless of model or task. ALTK-Evolve clusters and merges near-duplicate lessons support-conservingly into typed, retrievable guidelines, then delivers a small fixed core plus a per-task selected subset, or the full set only when a model has the headroom to use it. Teams designing agent memory pipelines can compare consolidation and delivery strategies discussed on daily.dev."}},{"@type":"Question","name":"Does giving an LLM agent more context or memory always improve its task accuracy?","acceptedAnswer":{"@type":"Answer","text":"No. On gpt-oss-120b, a weaker model, injecting ACE's full playbook helped on easy and medium AppWorld tasks but selective, per-task retrieval of guidelines won on hard tasks and on the aggregate score, because a large context can overwhelm a weaker model rather than help it. On the stronger DeepSeek-V3.2 model, more delivered lessons kept helping instead of crowding each other out. Engineers tuning how much context to feed an agent can follow benchmark findings like these on daily.dev."}}]}
```

