<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/agent-plasticity-measuring-whether-self-improvement-pays-off-szkrm5lwg" -->

---
title: Agent plasticity: measuring whether self-improvement...
description: Researchers at Meta Superintelligence Labs introduce 'agent plasticity', a metric measuring an agent's gain on held-out tasks per dollar spent on learning,...
canonical: https://daily.dev/posts/agent-plasticity-measuring-whether-self-improvement-pays-off-szkrm5lwg
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Agent plasticity: measuring whether self-improvement pays off | daily.dev
og:description: Researchers at Meta Superintelligence Labs introduce 'agent plasticity', a metric measuring an agent's gain on held-out tasks per dollar spent on learning,...
og:url: https://daily.dev/posts/agent-plasticity-measuring-whether-self-improvement-pays-off-szkrm5lwg
og:image: https://api.daily.dev/og/posts/szkrm5lWG.png
og:image:alt: Agent plasticity: measuring whether self-improvement pays off
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Agent plasticity: measuring whether self-improvement pays off

**[Collections](https://daily.dev/sources/collections)** · 2 min read · 1 upvotes · 0 comments

## Summary

Researchers at Meta Superintelligence Labs introduce 'agent plasticity', a metric measuring an agent's gain on held-out tasks per dollar spent on learning, with model weights frozen and fresh context each run. Testing on chess, Go, Hex, and NetHack shows the highest-scoring model isn't always the most efficient learner: Claude Fable 5 tops final scores in board games while GPT-5.6 Sol gains more per dollar, and only Claude Opus 5.5 shows significant improvement in NetHack. Faster learners tend to reuse their own prior artifacts, while slower learners ignore them.

## Content

Self-improving agents are hard to evaluate. Notes, skills, and tool calls carried between runs all change at once, so it's difficult to say which of them drives any gain. Researchers at Meta Superintelligence Labs propose a single measure for this, which they call agent plasticity.

## What it measures

Agent plasticity is the gain on held-out tasks per dollar spent on learning. Model weights stay frozen, and every run starts from a fresh context. Whatever improvement shows up has to come from what the agent writes down and reuses, not from retraining.

## What they found

The model that scores highest is often not the one that learns most efficiently.

- In chess, Go, and Hex, Claude Fable 5 reaches the highest final score, while GPT-5.6 Sol gains the most per dollar.
- In NetHack, only Claude Opus 5.5 improves significantly: 66 normalized points for about $1,073 of learning.

The researchers also looked at how agents treat their own artifacts. Slow learners often ignore artifacts they already wrote. Faster learners reuse theirs, though they still fail when an artifact is low quality. Writing notes isn't enough. They have to be good, and the agent has to actually read them.

I like that the metric puts a price on learning. A leaderboard score says where a model ended up, but it doesn't say what it cost to get there, and those are different questions.

## Questions this post answers

### What is agent plasticity in AI agent research?

Agent plasticity is a metric measuring the gain an AI agent achieves on held-out tasks per dollar spent on learning, with model weights kept frozen and each run starting from a fresh context. It isolates improvement coming from artifacts the agent builds and reuses, such as notes or skills, rather than from retraining weights. Researchers at Meta Superintelligence Labs proposed it to separate raw model strength from learning efficiency.

_Anyone comparing self-improving AI agents can track metrics like this on daily.dev._

### Which AI model is most efficient at self-improvement in agent benchmarks?

GPT-5.6 Sol gains the most per dollar spent on learning in chess, Go, and Hex, even though Claude Fable 5 reaches the highest final score in those games. In NetHack, only Claude Opus 5.5 shows significant improvement, gaining 66 normalized points for about $1,073 of learning spend.

_Developers weighing AI agent choices can follow these efficiency comparisons on daily.dev._

## Community take

How the wider developer community reacted, aggregated from 2 discussions and 41 comments across x (as of 2026-10-10).

**TL;DR:** People find the core idea (measuring per-dollar learning gains with frozen weights) clever, but much of the discussion zeroes in on the finding that slow learners write artifacts they never reuse, and whether the metric itself (gain per dollar, reuse vs. usefulness) is even measuring what it claims to.

**Sentiment:** 20% positive · 55% mixed · 25% skeptical

**The case for**

- The reframe of measuring gain per dollar instead of raw final score is seen as a genuinely useful way to pick a base model.
- Separating reuse from usefulness is viewed as a helpful methodological result.
- Several see the artifact-reuse finding (fast learners reuse notes, slow learners ignore them) as the most interesting and actionable takeaway.

**The pushback**

- Some argue gain-per-dollar conflates model pricing with actual learning efficiency, so a cheaper-but-worse learner could win artificially.
- Several doubt the setup generalizes beyond games with clear win signals, since real-world tasks lack a scoreboard to validate self-improvement.
- Multiple commenters question whether the ablations actually isolate notes vs. skills vs. tool calls as the source of gains.
- One notes that fast learners can still fail when the reused artifact itself is low quality, so curation matters more than reuse alone.

**By community**

- x (mixed): Reactions split between appreciation for the frozen-weights/per-dollar framing and pointed methodological questions about whether the metric truly measures self-improvement versus artifact reuse or pricing.

**Hottest debate:** Whether gain-per-dollar is really measuring learning efficiency or just rewarding cheaper models and good artifact curation rather than genuine plasticity.

**Open questions**

- Does the artifact reuse gap hold consistently across all three board games or mainly where score gaps are largest?
- Can gains be attributed specifically to notes versus skills versus tool calls?
- What happens if a low-plasticity model inherits a high-plasticity model's artifacts?
- Would normalizing by tokens or samples instead of dollars change the ranking?

**Highlights**

> @omarsar0 gain per dollar bakes model price into the metric. a 10x cheaper model that learns worse per sample still wins it, so that's measuring pricing strategy, not plasticity. normalize by tokens or samples and the ranking probably flips.
> — [scrblanc on x](https://x.com/scrblanc/status/2108582151826321538)

> @omarsar0 fast learners reuse their own artifacts and still fail when the artifact is bad, so the work that matters is curating what the agent keeps. $1,073 for 66 nethack points says bad notes are expensive.
> — [thebasedcapital on x](https://x.com/thebasedcapital/status/2108595395261550852)

> @omarsar0 the artifact reuse finding is the sleeper. slow learners write notes and never read them. most agent memory stacks have the same bug: great write path, retrieval as an afterthought
> — [i\_Am\_Snow\_Flake on x · 1 points, 1 comments](https://x.com/i_Am_Snow_Flake/status/2108587188695003259)

> @ashkaidev @omarsar0 All three, though the size varies by game. Fig. 29 shows the Go/Hex breakdown; Luna reuses artifacts on only ~1% of relevant Go decisions. High-reuse models still differ a lot in score, so reuse is one piece of the learning loop.
> — [dotvion on x](https://x.com/dotvion/status/2108587618976030901)

> @omarsar0 best performer not being the best learner per dollar is the interesting split imo. slow learners ignoring artifacts they already wrote sounds almost fixable. you read that as a prompting problem or a model one?
> — [gol\_tiro on x](https://x.com/gol_tiro/status/2108580919800136140)

**Source threads**

- [x](https://x.com/_akhaliq/status/2108333776774508934) · 0 points · 10 comments
- [x](https://x.com/omarsar0/status/2108580251127439710) · 0 points · 31 comments

## Similar posts on daily.dev

- [An AI meant to learn from its mistakes exploited a mistake in the test](https://daily.dev/posts/an-ai-meant-to-learn-from-its-mistakes-exploited-a-mistake-in-the-test-skqk1r4qs) · The Next Web · 0 upvotes · 0 comments

---

Tags: [#llm](https://daily.dev/tags/llm), [#ai-agents](https://daily.dev/tags/ai-agents), [#claude](https://daily.dev/tags/claude)

[View this post on daily.dev](https://daily.dev/posts/agent-plasticity-measuring-whether-self-improvement-pays-off-szkrm5lwg)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Agent plasticity: measuring whether self-improvement pays off","url":"https://daily.dev/posts/agent-plasticity-measuring-whether-self-improvement-pays-off-szkrm5lwg","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/agent-plasticity-measuring-whether-self-improvement-pays-off-szkrm5lwg"},"datePublished":"2026-10-09T15:28:40.107Z","dateModified":"2026-10-10T03:31:30.286Z","description":"Researchers at Meta Superintelligence Labs introduce 'agent plasticity', a metric measuring an agent's gain on held-out tasks per dollar spent on learning,...","image":"https://pbs.twimg.com/media/HUMukCpaIAAfZwQ.jpg","thumbnailUrl":"https://pbs.twimg.com/media/HUMukCpaIAAfZwQ.jpg","isAccessibleForFree":true,"articleSection":"Collections","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Collections","logo":"https://media.daily.dev/image/upload/s--fk_6ycEi--/f_auto,q_auto/v1780996001/logos/collections?_a=BAMAMiWQ0","url":"https://daily.dev/sources/collections"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/agent-plasticity-measuring-whether-self-improvement-pays-off-szkrm5lwg","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"llm,ai-agents,claude","timeRequired":"PT2M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Collections","item":"https://daily.dev/sources/collections"},{"@type":"ListItem","position":3,"name":"Agent plasticity: measuring whether self-improvement pays off"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/agent-plasticity-measuring-whether-self-improvement-pays-off-szkrm5lwg#faq","mainEntity":[{"@type":"Question","name":"What is agent plasticity in AI agent research?","acceptedAnswer":{"@type":"Answer","text":"Agent plasticity is a metric measuring the gain an AI agent achieves on held-out tasks per dollar spent on learning, with model weights kept frozen and each run starting from a fresh context. It isolates improvement coming from artifacts the agent builds and reuses, such as notes or skills, rather than from retraining weights. Researchers at Meta Superintelligence Labs proposed it to separate raw model strength from learning efficiency. Anyone comparing self-improving AI agents can track metrics like this on daily.dev."}},{"@type":"Question","name":"Which AI model is most efficient at self-improvement in agent benchmarks?","acceptedAnswer":{"@type":"Answer","text":"GPT-5.6 Sol gains the most per dollar spent on learning in chess, Go, and Hex, even though Claude Fable 5 reaches the highest final score in those games. In NetHack, only Claude Opus 5.5 shows significant improvement, gaining 66 normalized points for about $1,073 of learning spend. Developers weighing AI agent choices can follow these efficiency comparisons on daily.dev."}}]}
```

