AI coding agents are doing exactly what their demos promise: they are producing more code. The problem is that software companies do not ship lines of code. They ship reviewed, tested, integrated and maintainable changes.
A new Harvard study using engineering telemetry from 718 firms and roughly 300 million work events finds that after companies adopt AI coding agents, lines of code rise about 30%, commits rise 20% and pull requests rise 23%. Yet the researchers find little evidence that completed software output rises with them.
The missing productivity shows up downstream. Average pull-request review time rises about 49%, the share of PRs sent back for changes nearly doubles, and reviewer comments increase roughly 35%. The bottleneck has moved from writing code to deciding whether the code should be trusted and merged.
This Is One of the Largest Real-World Studies of AI Coding Yet
Harvard researchers Fiona Chen and James Stratton analyzed a proprietary dataset from engineering analytics company Jellyfish covering GitHub activity, Jira work and other engineering events from 2021 through March 2026. Their August paper, Artificial Intelligence in the Firm: Bottlenecks in Software Production, uses differences in when companies adopted coding assistants and more autonomous coding agents to estimate what changed afterward.
The result is a useful correction to one of the easiest mistakes in the AI-coding debate: code production is not the same thing as software production. A repository can gain more commits and more lines while the organization still closes roughly the same number of meaningful issues and epics.
That distinction also helps explain why GitHub is already rebuilding infrastructure around agent-scale development. BitcoinVersus.Tech recently covered how GitHub says commits jumped roughly fivefold in a year as coding agents created a much denser stream of repository activity.
The Agent Can Finish Writing Before the Team Is Ready to Review
Traditional software teams were partly rate-limited by implementation. A developer might spend hours or days understanding a task, changing code and preparing a pull request. Reviewers had time to absorb the queue because new changes arrived at approximately human speed.
Coding agents break that balance. One engineer can ask multiple agents to work in parallel, creating a sudden burst of pull requests. The organization may have automated tests and static analysis, but senior engineers still carry architecture knowledge, production context and accountability that cannot simply be multiplied on demand.
The Numbers Show a Pipeline Bottleneck, Not a Failure to Generate Code
The Harvard result is not “AI cannot code.” The measured coding stage clearly speeds up. After agent adoption, firms produce substantially more source code, commits and pull requests per worker.
The problem is pass-through. The researchers report no statistically significant increase in the resolution rate of Jira Issues and Epics, which they use as higher-level measures of completed software work. They also report no significant employment reduction attributable to the adoption of coding agents.
At the same time, more employees get pulled into review. Ars Technica’s reporting on the study notes that the share of workers performing code reviews rises about 14% after agent introduction. That is the hidden labor transfer: the machine saves implementation time but creates more verification work downstream.
AI Review Is Growing Too—but Humans Still Carry Most of the Load
The obvious response is to use AI to review AI. Companies are already doing that. By March 2026, the researchers found that roughly 80% of measured firms used some form of AI code review.
But automation had not eliminated the review wall. Agents produced about 23.3% of measured review comments and were associated with 10.8% of pull requests in the dataset, leaving humans responsible for most review activity. That is one reason GitHub is investing so heavily in better code-review agents and benchmarks.
BitcoinVersus.Tech recently examined GitHub ReviewBench, which measures whether AI reviewers can catch meaningful bugs without flooding pull requests with low-value comments. That problem becomes more important as generated-code volume grows.
More Comments Are Not Automatically Better Review
Review quality is difficult to automate because a useful reviewer needs more than syntax knowledge. The reviewer may need to understand the architecture, customer expectations, operational history, security assumptions, undocumented dependencies and which strange-looking behavior is actually intentional.
An AI reviewer can catch obvious bugs or inconsistencies and still miss the one architectural mistake that matters. It can also create the opposite problem: a wall of plausible but low-priority comments that makes the human reviewer slower.
GitHub has publicly described this tuning problem. In July, GitHub engineers reported that giving Copilot code review better general-purpose tools initially made the system worse: review cost increased while fewer issues were found. Rewriting the agent’s instructions around the actual review workflow later cut average review cost by roughly 20% without degrading the quality signal. That is a reminder that agent productivity depends on workflow design, not simply model capability.
The Practical Answer May Be Smaller PRs, Better Tests and Stronger Boundaries
If agents can write faster than humans can validate, the answer is not necessarily to generate less code. It is to restructure the pipeline so each unit of generated work is easier to verify.
- Keep agent tasks narrow. Smaller diffs are easier to understand, test and roll back.
- Make acceptance criteria machine-readable. Agents perform better when expected behavior is explicit before coding begins.
- Increase deterministic verification. Tests, type checking, static analysis, security scanning and reproducible builds should reject obvious failures before a human sees the PR.
- Protect architectural decisions. Agents should know which APIs, modules and boundaries they are allowed to change.
- Measure shipped outcomes. Lines of code and commit counts become increasingly misleading when software can be generated almost instantly.
The lesson is similar to what happened when faster CPUs, faster networks and faster CI systems removed older constraints: the slowest remaining stage becomes more visible. In agentic software development, that slow stage increasingly looks like trusted human judgment.
This Does Not Mean AI Coding Has Failed
The study’s data end in March 2026, and coding agents have improved rapidly since then. The analysis is also observational rather than a randomized experiment. Companies choose when and how to adopt AI, teams differ in maturity, and Jira completion is only one way to measure software output.
There are also important tasks where agent speed may create value even if headline feature output does not immediately rise: writing tests, migrating APIs, upgrading dependencies, generating documentation, investigating failures and reducing maintenance backlog. BitcoinVersus.Tech has already covered examples such as GitHub using AI agents during a major TypeScript-to-Rust runtime rewrite.
The more defensible conclusion is narrower: AI has made code generation cheap enough that review, integration and organizational judgment are becoming the scarce resources.

What Engineering Teams Should Measure Now
If AI coding agents keep getting faster, organizations will need better productivity metrics than “how much code did we produce?” Useful measures are more likely to include issue cycle time, escaped defects, rollback frequency, reviewer hours, test coverage, deployment frequency, change-failure rate and the amount of production-qualified work delivered per engineer.
That is the deeper implication of the Harvard result. The AI coding race is shifting from generation to verification economics. The winning engineering organizations may not be the ones whose agents write the most code. They may be the ones that can establish trust in each change with the least human attention.
Bottom Line
AI coding agents appear to be delivering a genuine productivity gain at the act of writing code. But software delivery is a system, and speeding up one stage simply exposes the next constraint.
Across 718 firms, the new constraint looks increasingly obvious: agents can generate changes faster than humans can confidently review, integrate and own them. Until review and verification scale with generation, more code will not automatically mean more software.
Ars Technica’s October 9 report provides additional analysis of the study’s review and output findings. Read the reporting here.
Editor’s Note
The Harvard paper is observational research using firm-level adoption timing and engineering telemetry, not a randomized controlled trial. Its data end in March 2026, so the findings should not be interpreted as a permanent ceiling on newer coding-agent systems.
BitcoinVersus.Tech is independently maintained. Support options on the site help fund additional technical research, verification and open educational publishing.
BitcoinVersus.tech is not a financial advisor. Content is provided for informational purposes.

Leave a Reply