<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/agents-refactor-300k-lines-in-three-weeks-and-practitioners-ask-what-it-proves-qtk43gheh" -->

---
title: Agents Refactor 300K Lines in Three Weeks, and...
description: CodeScene published a case study describing coding agents refactoring a 300,000-line C codebase (a Street Fighter III decompilation) over three weeks for about...
canonical: https://daily.dev/posts/agents-refactor-300k-lines-in-three-weeks-and-practitioners-ask-what-it-proves-qtk43gheh
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Agents Refactor 300K Lines in Three Weeks, and Practitioners Ask What It Proves | daily.dev
og:description: CodeScene published a case study describing coding agents refactoring a 300,000-line C codebase (a Street Fighter III decompilation) over three weeks for about...
og:url: https://daily.dev/posts/agents-refactor-300k-lines-in-three-weeks-and-practitioners-ask-what-it-proves-qtk43gheh
og:image: https://api.daily.dev/og/posts/qTk43gheH.png
og:image:alt: Agents Refactor 300K Lines in Three Weeks, and Practitioners Ask What It Proves
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Agents Refactor 300K Lines in Three Weeks, and Practitioners Ask What It Proves

**[InfoQ](https://daily.dev/sources/infoq)** · 5 min read · 0 upvotes · 0 comments

## Summary

CodeScene published a case study describing coding agents refactoring a 300,000-line C codebase (a Street Fighter III decompilation) over three weeks for about $4,000 in tokens, raising Code Health from 5.6 to 10.0 across 2,903 commits and 726 files. Two mechanisms enabled the work: a deterministic Code Health MCP scoring server and a frame-by-frame replay-trace harness for correctness. Agents also accumulated a reusable refactoring playbook of 22 recipes. Claude Opus outperformed Codex with Sol at capturing patterns; smaller models plateaued. Practitioner reaction split sharply: some called it a rigorous higher bar than green tests, while skeptics questioned scope (open-source game vs. production code), whether results were merged (yes, via 54 pull requests to a fork), unmeasured architecture quality, and whether the harness or the model was really being measured. Two headline figures — 70% fewer AI-induced defects and 45% less token waste — are projected extrapolations from prior research, not measured outcomes of this study, which feeds into a planned Lund University comparison study.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://www.infoq.com/news/2026/09/agentic-refactoring-case-study>

## Questions this post answers

### How much did it cost to refactor a 300,000-line C codebase using AI coding agents?

Roughly $4,000 in token costs over three weeks, according to a CodeScene case study. The work produced 2,903 commits across 726 files, modified 252,055 lines, and raised the codebase's Code Health score from 5.6 to 10.0. The codebase was Street Fighter III: 3rd Strike, from an open-source decompilation.

_Track real-world cost and outcome data for agentic refactoring projects like this one on daily.dev._

### Why did Claude Opus outperform Codex with Sol for AI-driven code refactoring?

Claude Code with Opus was reported as significantly better than Codex with Sol at capturing and documenting emerging refactoring patterns during a large-scale agentic refactoring project. Smaller models tended to plateau on files, appearing to hit a local optimum they couldn't move past, while Opus kept improving Code Health scores throughout the three-week effort.

_Compare model performance on real coding tasks like refactoring before picking a daily driver on daily.dev._

### What is a replay-trace harness and why does it matter for verifying AI code refactoring?

A replay-trace harness compares a rollback state hash frame by frame to confirm behavior is unchanged after each code transformation, offering a stronger correctness check than a passing test suite alone, since tests only confirm the tests survived. It worked in one case because a decompiled game offers deterministic frame-by-frame replay, an oracle most legacy systems lack.

_Weigh how you'd verify correctness before trusting agents with large refactors, and follow the debate on daily.dev._

## Similar posts on daily.dev

- [The Economic Benefit of Refactoring](https://daily.dev/posts/the-economic-benefit-of-refactoring-ogtaw7yn7) · Martin Fowler · 0 upvotes · 0 comments
- [10 tips to improve your coding agent game](https://daily.dev/posts/10-tips-to-improve-your-coding-agent-game-r2zthx2x0) · Pulumi · 3 upvotes · 0 comments

---

Tags: [#ai-coding](https://daily.dev/tags/ai-coding), [#claude-code](https://daily.dev/tags/claude-code), [#technical-debt](https://daily.dev/tags/technical-debt)

[View this post on daily.dev](https://daily.dev/posts/agents-refactor-300k-lines-in-three-weeks-and-practitioners-ask-what-it-proves-qtk43gheh)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Agents Refactor 300K Lines in Three Weeks, and Practitioners Ask What It Proves","url":"https://daily.dev/posts/agents-refactor-300k-lines-in-three-weeks-and-practitioners-ask-what-it-proves-qtk43gheh","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/agents-refactor-300k-lines-in-three-weeks-and-practitioners-ask-what-it-proves-qtk43gheh"},"datePublished":"2026-09-30T06:24:34.817Z","dateModified":"2026-09-30T06:24:59.569Z","description":"CodeScene published a case study describing coding agents refactoring a 300,000-line C codebase (a Street Fighter III decompilation) over three weeks for about...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/3b33f02979a3305b4ee971f3ff008822?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/3b33f02979a3305b4ee971f3ff008822?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"InfoQ","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"InfoQ","logo":"https://media.daily.dev/image/upload/t_logo,f_auto/v1/logos/afc3bced3e1e4b188dd9127017a60e0c","url":"https://daily.dev/sources/infoq"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/agents-refactor-300k-lines-in-three-weeks-and-practitioners-ask-what-it-proves-qtk43gheh","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"ai-coding,claude-code,technical-debt","timeRequired":"PT5M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"InfoQ","item":"https://daily.dev/sources/infoq"},{"@type":"ListItem","position":3,"name":"Agents Refactor 300K Lines in Three Weeks, and Practitioners Ask What It Proves"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/agents-refactor-300k-lines-in-three-weeks-and-practitioners-ask-what-it-proves-qtk43gheh#faq","mainEntity":[{"@type":"Question","name":"How much did it cost to refactor a 300,000-line C codebase using AI coding agents?","acceptedAnswer":{"@type":"Answer","text":"Roughly $4,000 in token costs over three weeks, according to a CodeScene case study. The work produced 2,903 commits across 726 files, modified 252,055 lines, and raised the codebase's Code Health score from 5.6 to 10.0. The codebase was Street Fighter III: 3rd Strike, from an open-source decompilation. Track real-world cost and outcome data for agentic refactoring projects like this one on daily.dev."}},{"@type":"Question","name":"Why did Claude Opus outperform Codex with Sol for AI-driven code refactoring?","acceptedAnswer":{"@type":"Answer","text":"Claude Code with Opus was reported as significantly better than Codex with Sol at capturing and documenting emerging refactoring patterns during a large-scale agentic refactoring project. Smaller models tended to plateau on files, appearing to hit a local optimum they couldn't move past, while Opus kept improving Code Health scores throughout the three-week effort. Compare model performance on real coding tasks like refactoring before picking a daily driver on daily.dev."}},{"@type":"Question","name":"What is a replay-trace harness and why does it matter for verifying AI code refactoring?","acceptedAnswer":{"@type":"Answer","text":"A replay-trace harness compares a rollback state hash frame by frame to confirm behavior is unchanged after each code transformation, offering a stronger correctness check than a passing test suite alone, since tests only confirm the tests survived. It worked in one case because a decompiled game offers deterministic frame-by-frame replay, an oracle most legacy systems lack. Weigh how you'd verify correctness before trusting agents with large refactors, and follow the debate on daily.dev."}}]}
```

