Skip to main content

How many agents can you actually review?

Parallel agents multiply what gets generated. What gets reviewed and accepted is capped by your good sittings each day. This page compares the two, projects the backlog or the escaped defects over two working weeks, and ends with a short game about approval fatigue. Everything runs in this tab.

Predict first

With 3 agents each opening 2 pull requests of 300 lines a day, how many agents can you keep fully reviewed?

You review well for 3 sittings a day, about 400 lines each.

The rest of the page opens once you’ve guessed or skipped.

Split the pull request

How many effective sittings a pull request needs, and the reviewable units to cut it into.

300 lines fit in one effective sitting of 400 lines, about 60 minutes. Review it as it is.

Measure your real throughput

Paste a listing of merged pull requests to see how big they really are, how many land a day, and how many come from agents. Then load the numbers into the simulator.

Loading the listing reader…

Feel the tenth approval

Twenty approval prompts, one of them dangerous, somewhere from the tenth on. See whether you catch it, and how your pace changes as you go.

Loading the approval game…

Assumptions and sources

  • Lines per sitting and minutes per sitting are sourced. A case study of code review at Cisco concluded that a review should cover under 200 lines, not to exceed 400, and take under 60 minutes, not to exceed 90. It also found detection was best below about 300 lines an hour, so the default of 400 lines in 60 minutes is at the top of what that study calls good.
  • The follow-up share is sourced. A report on dotnet/runtime found that 52.3% of merged pull requests from the Copilot coding agent got commits from someone other than the author, against 10.3% of merged human pull requests.
  • Three to five agents is a recommendation, not a measurement. Anthropic suggests three to five parallel sessions. Git isn’t what limits that number. Your review is.
  • These are assumptions, labeled as such: pull requests per agent, lines per pull request, fresh sittings per day, the fatigued factor, defects per 1,000 lines, the fresh detection rate, and the minutes of follow-up. Change them to match what you see.
  • What this leaves out: more than one reviewer, how review quality changes through the day beyond fresh and tired, and anything live from your repository.

What to do about it

  • Make the history reviewable. The outline’s advice is that the unit of history should match the unit of change. A pull request that needs three sittings is three changes, and reviewing it in one is the tired review this page models.
  • An agent’s approval is not a human’s. A reviewing agent can catch a lot, but counting its approval as yours only hides the queue. It doesn’t clear it.
  • Run as many agents as you can review. Past that, the extra output either waits or gets a rubber stamp.