<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/how-much-memory-does-your-agent-actually-need--w9bapy7xf" -->

---
title: How Much Memory Does Your Agent Actually Need? | daily.dev
description: IBM Research's ALTK-Evolve framework lets agents distill reusable behavioral guidelines from their own past trajectories, injecting them back at inference time...
canonical: https://daily.dev/posts/how-much-memory-does-your-agent-actually-need--w9bapy7xf
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: How Much Memory Does Your Agent Actually Need? | daily.dev
og:description: IBM Research's ALTK-Evolve framework lets agents distill reusable behavioral guidelines from their own past trajectories, injecting them back at inference time...
og:url: https://daily.dev/posts/how-much-memory-does-your-agent-actually-need--w9bapy7xf
og:image: https://api.daily.dev/og/posts/W9BApY7xf.png
og:image:alt: How Much Memory Does Your Agent Actually Need?
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# How Much Memory Does Your Agent Actually Need?

**[Hugging Face](https://daily.dev/sources/huggingface)** · 9 min read · 1 upvotes · 0 comments

## Summary

IBM Research's ALTK-Evolve framework lets agents distill reusable behavioral guidelines from their own past trajectories, injecting them back at inference time without weight updates or human annotation. Testing across eight models on the AppWorld benchmark revealed that the ideal amount of injected memory depends on model capability: strong models with headroom benefit most from the full guideline set, weaker models perform better with a compact core plus retrieved task-specific guidelines, and already-saturated models show no measurable gain. Notably, gpt-oss-120b improved task completion by 16.1 percentage points using curated retrieval while only increasing token usage by 5%, making selective retrieval both the most accurate and cheapest option for weaker models. Prompt caching is highlighted as key to keeping full guideline injection affordable in production.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://huggingface.co/blog/ibm-research/altk-evolve-hmm>

## Questions this post answers

### How much does adding a curated retrieval memory strategy improve gpt-oss-120b's task completion on AppWorld, and what's the token cost?

Curated retrieval memory improves gpt-oss-120b's task goal completion by +16.1 percentage points (from 39.9% to 56.0%) on AppWorld's test_normal split, while only increasing token usage per task by about 5% (from 110K to 116K tokens). This outperforms giving it the full guideline set, which gains less accuracy while costing roughly 50% more tokens.

_Anyone tuning agent memory budgets can track findings like this on daily.dev to avoid overspending on tokens._

### Should I give a strong LLM agent its full guideline set or just retrieved guidelines when using agentic memory?

Strong models with headroom, such as DeepSeek-V3.2 (671B MoE) and Claude Opus 4.6, perform best when given the full mined guideline set injected on every step, gaining +9.5 and +4.1 percentage points in task completion respectively. Weaker models like gpt-oss-120b instead do better with a compact high-confidence core plus a few task-relevant guidelines retrieved per task, since the full set drowns them and costs more tokens.

_Developers deciding between memory strategies for their agent stack can follow this research on daily.dev._

### Why did GLM-5 show no improvement from agentic memory guidelines on the AppWorld benchmark?

GLM-5 (745B MoE) showed zero change in both task goal completion (87.5%) and scenario goal completion (80.4%) whether given the full guideline set or the no-memory baseline, a pattern labeled 'saturated.' Possible explanations include the model already being near its performance ceiling on these tasks, the guidelines not addressing its remaining failure modes, or the model not applying the guidance effectively - the cause was not isolated.

_Teams debugging why memory or guideline injection isn't moving their agent's accuracy can dig into cases like this on daily.dev._

## Similar posts on daily.dev

- [Thinking of ACE? We Can Do It with Fewer Tokens](https://daily.dev/posts/thinking-of-ace-we-can-do-it-with-fewer-tokens-0stmkpisz) · Hugging Face · 1 upvotes · 0 comments
- [Retrieval vs. Memory in Agentic AI System](https://daily.dev/posts/retrieval-vs-memory-in-agentic-ai-system-w27x4mch5) · Machine Learning Mastery · 2 upvotes · 0 comments
- [Memory Scaling for AI Agents](https://daily.dev/posts/memory-scaling-for-ai-agents-yneck54it) · databricks · 1 upvotes · 0 comments

---

Tags: [#llm](https://daily.dev/tags/llm), [#ai-agents](https://daily.dev/tags/ai-agents), [#ai-inference](https://daily.dev/tags/ai-inference)

[View this post on daily.dev](https://daily.dev/posts/how-much-memory-does-your-agent-actually-need--w9bapy7xf)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"How Much Memory Does Your Agent Actually Need?","url":"https://daily.dev/posts/how-much-memory-does-your-agent-actually-need--w9bapy7xf","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/how-much-memory-does-your-agent-actually-need--w9bapy7xf"},"datePublished":"2026-08-18T18:10:14.069Z","dateModified":"2026-09-14T07:31:16.833Z","description":"IBM Research's ALTK-Evolve framework lets agents distill reusable behavioral guidelines from their own past trajectories, injecting them back at inference time...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/34752ccab68e21f2f55076fd19a624a9?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/34752ccab68e21f2f55076fd19a624a9?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"Hugging Face","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Hugging Face","logo":"https://media.daily.dev/image/upload/t_logo,f_auto/v1/logos/f1f55c67d81a4330acf5b90b26b0c8e1","url":"https://daily.dev/sources/huggingface"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/how-much-memory-does-your-agent-actually-need--w9bapy7xf","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"llm,ai-agents,ai-inference","timeRequired":"PT9M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Hugging Face","item":"https://daily.dev/sources/huggingface"},{"@type":"ListItem","position":3,"name":"How Much Memory Does Your Agent Actually Need?"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/how-much-memory-does-your-agent-actually-need--w9bapy7xf#faq","mainEntity":[{"@type":"Question","name":"How much does adding a curated retrieval memory strategy improve gpt-oss-120b's task completion on AppWorld, and what's the token cost?","acceptedAnswer":{"@type":"Answer","text":"Curated retrieval memory improves gpt-oss-120b's task goal completion by +16.1 percentage points (from 39.9% to 56.0%) on AppWorld's test_normal split, while only increasing token usage per task by about 5% (from 110K to 116K tokens). This outperforms giving it the full guideline set, which gains less accuracy while costing roughly 50% more tokens. Anyone tuning agent memory budgets can track findings like this on daily.dev to avoid overspending on tokens."}},{"@type":"Question","name":"Should I give a strong LLM agent its full guideline set or just retrieved guidelines when using agentic memory?","acceptedAnswer":{"@type":"Answer","text":"Strong models with headroom, such as DeepSeek-V3.2 (671B MoE) and Claude Opus 4.6, perform best when given the full mined guideline set injected on every step, gaining +9.5 and +4.1 percentage points in task completion respectively. Weaker models like gpt-oss-120b instead do better with a compact high-confidence core plus a few task-relevant guidelines retrieved per task, since the full set drowns them and costs more tokens. Developers deciding between memory strategies for their agent stack can follow this research on daily.dev."}},{"@type":"Question","name":"Why did GLM-5 show no improvement from agentic memory guidelines on the AppWorld benchmark?","acceptedAnswer":{"@type":"Answer","text":"GLM-5 (745B MoE) showed zero change in both task goal completion (87.5%) and scenario goal completion (80.4%) whether given the full guideline set or the no-memory baseline, a pattern labeled 'saturated.' Possible explanations include the model already being near its performance ceiling on these tasks, the guidelines not addressing its remaining failure modes, or the model not applying the guidance effectively - the cause was not isolated. Teams debugging why memory or guideline injection isn't moving their agent's accuracy can dig into cases like this on daily.dev."}}]}
```

