---
title: "Code review is more than a bug hunt, and coding agents miss the point"
description: "A veteran engineer’s answer to the claim that AI has ended human inspection of software"
author: "Albion Grey"
published: 2026-09-28T08:45:54.903Z
modified: 2026-09-28T09:14:33Z
url: https://rews.cc/a/code-review-is-more-than-a-bug-hunt-and-coding-agents-miss-t-0bc2c7
language: en
tags: ["ai", "automation", "software", "productivity", "openai", "tech"]
publisher: "Rews (https://rews.cc)"
---

# Code review is more than a bug hunt, and coding agents miss the point

*A veteran engineer’s answer to the claim that AI has ended human inspection of software*

By Albion Grey · September 28, 2026 · https://rews.cc/a/code-review-is-more-than-a-bug-hunt-and-coding-agents-miss-t-0bc2c7

## In brief

- A paper argues coding agents can serve every stated goal of code review at lower cost and higher throughput
- John Allspaw of Adaptive Capacity Labs calls the argument the “substitution myth”
- He says agents miss human confusion, challenges to a change’s necessity, and defects of omission
- He adds that review calibrates scrutiny to the author, transfers knowledge both ways, and uses context outside repositories
- He argues review is a coordination, sensemaking and governance process, not just detection

Since Michael Fagan formalised code inspection in 1976, having one engineer read another’s work before it ships has been software’s nearest thing to a sacred rite. A recent paper, *The End of Code Review: Coding Agents Supersede Human Inspection*, declares the rite obsolete. Large-language-model agents, its authors argue, have crossed a capability threshold: every stated goal of code review—detection of defects, enforcement of style, transfer of knowledge, general awareness—can be served by a machine “at lower cost and higher throughput”. Worse, they say, the naive arrangement in which agents write code and humans merely review it is a dead end, providing neither meaningful assurance nor the scale that AI-assisted output demands.

John Allspaw, a veteran of running big systems and a founder of Adaptive Capacity Labs, is unpersuaded. In a pointed rebuttal published on his firm’s site he accuses the paper of falling for what he calls the substitution myth: first decompose a human job into measurable functions, then show a machine performing each function, then declare the human redundant. The trick, he notes, always fails at the same place—the human contribution that mattered most was the integration across the functions, the ability to adapt to unplanned circumstances and to carry the social accountability the job implies.

## What the diff doesn’t show

His catalogue of what an agent cannot do is quietly devastating. When an experienced engineer reads a change and says “I don’t understand this”, the confusion is itself the finding—a signal that the code is too complex or its intent unclear. A model will always “understand” code, in the sense of processing it, so it can never deliver that signal. A human reviewer may question whether the change should exist at all—“This solves the symptom, not the problem”—since review is often the last moment anyone asks. And humans notice absence: an API contract that changed while its error handling did not. Absence-blindness, Mr Allspaw observes, is precisely the class of failure at which language models tend to be poor. “The agent reviews what is there; engineers with expertise can easily notice what’s missing.”

Three more functions resist automation, he argues. Reviewers calibrate their scrutiny to the author: a junior engineer’s first commit to a payments module gets a different reading from a veteran’s routine refactoring, because risk lives in the who and when, not just the what. Knowledge transfer is bidirectional—review is a conversation that changes both participants’ understanding, not a one-way transmission that a generated summary can replace. And reviewers carry context that lives outside the repository: last Tuesday’s incident, a downstream team’s plan to deprecate an interface, what legal said about logging a field. The paper assumes the codebase is the complete context, he writes. “It never is.”

Finally, accountability. The paper treats the human reviewer as a compliance artefact—a named signature for legal comfort—and outsources the ethical questions to requirements engineering and post-deployment monitoring, which Mr Allspaw dismisses as kicking the can down the road. Knowing that you personally approved a change shapes how you read it; an agent that signs off a pull request bears no consequences and has no skin in the game. Review is not foremost a detection process, he concludes, but a coordination process, a sensemaking process and a governance process.

The argument resonated with practitioners. One Hacker News commenter, clintonb, who is weighing code review inside his own organisation, put the worry plainly: “The sad truth is that all of my feedback just goes straight to agents.” Another, dimbletimbers, offered the defence Mr Allspaw’s essay implies: taken seriously, review means “at least two people understand how the feature works (even if that number is, on average, trending closer to between one and zero)”. Not everyone fretted: a third commenter pointed out that humans have happily offloaded knowledge before—nobody solders their own microchips at scale any more.

Both sides can be right in part. Agents are plainly useful at catching the mechanical and the stylistic, and pretending otherwise is Luddism. But the paper’s error is to mistake a job’s visible outputs for the job. Code review, like an editor’s pencil or a second pilot, earns its keep on the rare days when something nobody thought to test for is quietly wrong. Those days are rare. They are also the ones that matter.
