A new engineering analysis from The New Stack finds that pull requests merged without any review, human or machine, jumped 31.3% as AI-generated code flooded development teams. The essay, published this week, argues that pointing AI at code diffs misses what code review was always for: not catching bugs, but building shared understanding across a team. The author proposes a five-layer model that separates what machines do well from what only humans can judge.
The report cites data from Faros AI's 2026 engineering report, which tracked 22,000 developers across more than 4,000 teams and found incidents per pull request up 242.7%, bugs per developer up 54%, and work restarts up 13.8%. DORA's 2025 State of AI-assisted Software Development report describes the core tension: AI adoption raises delivery throughput and delivery instability at the same time. One person shows you a velocity chart; another shows you an incident chart, and both are correct.
According to the essay, research from Microsoft in 2013 and Google in 2018 reached the same conclusion: code review is how a team builds a mental model of its system. In the Microsoft study, 44% of developers ranked finding defects as their top reason for reviewing, but only 14% of the actual comments were about defects. The author writes that "AI review looks at the wrong artifact, at the wrong time, with no access to what was decided." A diff can't tell you what was decided, and a model trained on the kind of code it's reviewing will never tell you not to build the thing.
The report explains that the current approach inverts what works: teams pointed the machine at judgment, the one thing it can't do, and left the mechanics to people. The essay proposes five layers to replace line-by-line review theater. First, let agents argue before the pull request exists, when changing course is cheapest, and review the record of what they couldn't settle. Second, capture intent and acceptance criteria as decisions are made during coding sessions, not in a waterfall spec written before the work begins. Third, turn repeated review comments into a "slop registry" of invariants checked on every change—most teams already have them, scattered across pull request comments and stored in engineers' memory. Fourth, debate the decisions the agents surfaced and couldn't resolve, leaving the diff to the machines. Fifth, define who owns the invariants that govern the code, since a model has no reputation, no liability, and can't be asked why.
The essay recommends starting with the cheapest layer: harvest your last 1,000 review comments and turn the top 20 into invariants, since every review comment your team repeats is a guardrail you haven't written yet. Open-source tools already cover most of the agent argument layer: PR-Agent runs reviews on multiple models, Aider's architect mode has one model propose and a different one edit before the pull request exists, and AutoGen handles multi-agent conversation. The author warns that automating review carelessly deletes shared understanding, quoting Turing Award winner Peter Naur's 1985 essay: "The death of a program happens when the programmer team possessing its theory is dissolved." In today's terms, that's your reorg, attrition, and backfill—and if nobody built the understanding in the first place, you get there without a single resignation. The problem isn't whether AI can spot a type error or flag a security risk; it's that teams are automating the wrong meeting and losing the only scheduled moment when engineers talk about the system. The shift from author-owns-code to reviewer-approves-code to team-owns-invariants changes what accountability means, and naming that before an incident forces the conversation is the difference between a deliberate handoff and an unnoticed collapse.

