CodeScene has published a case study showing that coding agents refactored a 300,000-line C codebase in three weeks for roughly $4,000 in token costs. The work pushed the codebase's Code Health score from 5.6 to a perfect 10.0, a result that company founder Adam Tornhill described as the first time he'd witnessed what he called superhuman AI performance at scale after 30 years working on large systems. The project used an open-source decompilation of Street Fighter III: 3rd Strike as its test case.

The agents generated 2,903 commits spanning 726 files and modified 252,055 lines of code. The work wasn't done through a fixed set of rules. Instead, the agents built their own refactoring playbook as they went, finishing with 22 recipes and 82 supporting notes. Standard transformations like Extract Function and Guard Clauses appeared alongside patterns unique to this codebase, including Shared Index Range for repeated loops differing only in start and end values, Action Parameter for duplicated control structures that varied mainly in which function they called, and Uniform Step Table to convert heterogeneous calls into table-driven dispatch. The playbook also documented failed attempts, recording transformations that worsened Code Health scores. Model choice proved significant, with the team settling on Claude Opus for most of the work after finding it substantially better than Codex with Sol at capturing and documenting emerging patterns. Smaller models often plateaued, seemingly stuck at local optimums they couldn't escape.

Two mechanisms made the work possible, according to the report. The first was a quality signal: the CodeHealth MCP Server, which gave agents a deterministic score to optimize and to judge whether a transformation had improved the code. The second was correctness verification through a replay-trace harness that compared the rollback state hash frame by frame, allowing behavior to be checked after every change. Tornhill wrote that "automated tests and equivalence checks are absolutely essential safeguards," placing the burden on exactly what unhealthy codebases typically lack. The team's selection process illustrated the challenge. They considered a Gov.UK marine licensing codebase but rejected it as too healthy for the research that would follow, choosing the game partly because they play it and are therefore its users.

Reaction from practitioners has been sharply divided, though the split runs along what the result proves rather than whether it happened. The skeptics concentrated on scope. Tech lead Konrad Otrębski questioned whether this was an experiment on open-source code rather than production code earning money, and later suggested the real test would be offering this refactoring to a well-known open-source project like Grafana, with merging to master as the definition of done. Software architect Tracy Bannon objected to the framing, noting that describing the outcome as perfect is "pretty bold." Denis Baltor challenged the discovered recipes themselves, arguing that DRY concerns duplication of knowledge and intent rather than identical lines of code. The authors raised unanswered questions too. When asked whether non-functional behavior had improved—framerate, memory use, and input latency matter in a game—Daniel Webb, one of the two engineers who did the work, said a performance specialist was being brought in.

The next step is a study with Lund University in which students will implement features in both versions of the codebase—one at Code Health 5.6 and one at 10.0—using frontier models to compare cost and quality. CodeScene projects roughly 70% fewer AI-induced defects and roughly 45% less token waste from the uplift, though both figures are extrapolated from earlier research rather than measured in this case study. The replay-trace harness worked because a decompiled game offers deterministic frame-by-frame replay, an oracle most legacy systems don't have—which is the same reason refactoring them is risky in the first place. For organizations considering similar approaches, the constraint isn't model capability but the availability of reliable correctness checks. The window for autonomous refactoring opens when you can define what correct means in a way a machine can verify on every commit.