Continuous integration systems are buckling under a 25-fold increase in job volume driven by AI coding agents, according to a September analysis published by The New Stack. The report synthesizes recent disclosures from Anthropic, Linear, and Depot, arguing that while companies are racing to speed up their pipelines, they're solving the wrong problem: CI still verifies individual code repositories, not the distributed systems those repositories actually run inside.

Anthropic's engineering team reported that their CI job count jumped 25 times in six months, while engineers now ship roughly eight times as much code per quarter compared to the 2021-2025 period. Linear disclosed that its test suite has grown nearly fourfold since January, with agents now authoring the majority of tests. Blacksmith, which provides CI runners, said the number of jobs it processes has climbed between 5% and 10% every week. The pattern repeats across vendors: when a single engineer operates multiple agents in parallel, pull request volume rises by multiples rather than percentages, and each agent waits for pipeline results that can take 20 minutes—long enough to lose working context between failure and retry.

The report argues that faster pipelines don't address the core gap: a repository represents one service among dozens in cloud-native architectures, and unit tests mock everything beyond that boundary. "A change can pass every unit test, pass CI in record time, pass a sandbox built from the branch, and still break the first real request that crosses a service boundary," the analysis states. According to the author, DevOps Research and Assessment found that higher AI adoption correlates with increases in both software delivery speed and instability—a pattern the report describes as "faster code, same verification, more breakage."

The report traces the problem to a 20-year assumption that CI would process human output: a few pull requests per week, with developers already working on the next task by the time results returned. Agents broke that model in two ways—volume and placement. Because CI runs after the pull request exists, an agent that writes code, opens a PR, and waits loses its working context before the pipeline completes. Meanwhile, every coding agent now runs code in some form of sandbox—Cursor uses cloud VMs, GitHub Copilot uses ephemeral Actions environments, and Greptile's TREX attaches logs and screenshots to PRs—but each sandbox contains only the repository, the branch, and setup scripts, not the 39 other services, real message queues, or production-shaped databases. The agent verifies its change against a copy of its own code, then CI repeats the same check faster, and only at the staging merge does anything confirm whether the change works with the rest of the system.

The report recommends moving verification inside the agent's loop, against the full system rather than an isolated repo. It proposes multiplexing: a single Kubernetes cluster runs one stable version of every service and hosts thousands of lightweight test environments on top, each deploying only the changed service while routing tagged requests through the modification and resolving other hops to shared stable versions. A test environment costs roughly one pod and spins up in seconds, allowing 50 agents to share one stable setup instead of cloning it 50 times. The report concludes that agents need governed verification steps—approved action sequences that exercise changes against live systems and record results—so platform teams can ensure safety and produce artifacts that review tools and merge gates can read. The teams that win, the analysis predicts, will stop asking how fast the pipeline can confirm a repo passes its own tests and start asking how early an agent can prove a change works with everything around it. The industry has spent two decades optimizing for human output; the question now is whether infrastructure can catch up before instability becomes the price of speed.