Many AI-generated code commits converging through a narrow gate into only three shipped releases

Abstract

AI-assisted software development accelerates code production more reliably than it accelerates delivery. Software development is a tractable proxy for other AI-assisted business processes because instructions, revisions and outcomes are digitally recorded, although effects cannot be transferred directly. Reliable measurement is needed to budget model, platform, integration and verification costs and construct time-to-market business cases.

This rapid, non-protocol-registered systematic review examines 115 documents published from 1 January 2024: 59 empirical studies, two secondary syntheses and 54 supporting or contextual documents. Forty productivity and delivery studies contribute 42 quantitative effect estimates; 19 studies examine review, quality, security, maintainability, skills and interaction cost. Evidence is grouped into three operating models:

  1. Assisted development: the model suggests; the developer implements and verifies.
  2. Spec-driven execution: the agent implements a defined outcome; a human reads and approves the result.
  3. Agent-native delivery: agents build, test and perform initial review; humans own intent, architecture and release.

Controlled evidence included a 19% slowdown in familiar mature repositories and gains of approximately 20–30% in larger enterprise trials; higher estimates came from smaller or narrower studies. The independent meta-analysis found a moderate effect with extreme heterogeneity and materially smaller effects outside laboratories. In the largest matched delivery study, more than 100,000 developers produced 180% more commits but only 30% more releases and no additional application usage. Release output increased by 1.1–1.3× across successive tool generations; this review therefore uses 1.2× as a central planning case, not a universal effect.

Reference values rise with automation: approximately 1.25× for assisted development, for bounded spec-driven execution and 4.5× for agent-native delivery. Only the first is supported by multiple controlled comparisons; the second is an interpolation and the third a low-confidence organisational upside scenario. Effects also vary with greenfield or brownfield context, programming-language feedback, developer experience and downstream verification capacity.

Quality is mixed: bounded trials report higher completeness and correctness, while longitudinal studies identify integration, review, complexity, ownership and security costs. Speed without scaled verification can come at the expense of quality. Maintaining quality at higher generation rates requires executable tests, security controls, review capacity, monitoring and accountable ownership, but does not guarantee it. Productivity should therefore be measured from accepted intent to reliable production—not by code volume; a coding-stage gain is not automatically a lead-time or customer-value gain.

Review question and method

The review question was: how does measured software-development productivity change as AI use progresses from prompt assistance to specified delegation and agent-native delivery?

The search covered publications dated from 1 January 2024 to 28 July 2026. Search combinations included AI coding assistant, developer productivity, developer experience, repository familiarity, domain expertise, coding agent, spec-driven development, agentic code review, deployment velocity, software delivery, programming language, static typing, security, maintainability and unit-test generation. References were followed back to peer-reviewed papers, preprints, official research publications and first-party engineering reports. Where work first appeared as an older preprint but received a 2024–2026 peer-reviewed publication, the later publication was used; otherwise pre-2024 work was excluded.

A source was eligible for the productivity synthesis when it reported a quantitative outcome, described an identifiable development workflow and supplied a baseline or comparator. Surveys were retained only as contextual evidence about perceived effects or bottlenecks. Unattributed marketing claims and secondary summaries without an inspectable primary source were excluded. Source traceability, uncertainty and corrections follow Reinvently’s research and editorial standards.

Of the 115 included documents, 59 form the empirical evidence base. The quantitative productivity synthesis contains 42 effect estimates drawn from 40 of those studies; 19 additional studies measure adjacent outcomes. The remaining 56 documents comprise 54 contextual or methodological sources and two secondary syntheses. Two reports each contribute more than one estimate because they describe materially different workflows. Documents, studies and estimates are therefore counted separately.

The evidence chain

ObjectCountMeaningHow it is traced
Document115A paper, preprint, report or webpageDirect URL in one of the three reference registers
Included empirical study59An evaluation contributing measured productivity or an adjacent outcomeStable study identifier, ST-01 to ST-59
Productivity effect estimate42A reported comparison for time, output, workflow activity or deliveryStable estimate identifier, E-01 to E-42
Review conclusionA synthesis made in this articleLinked to the relevant estimate identifiers and studies

The trace is: conclusion → effect estimate → study → source document. A study’s design, the directness of its outcome and confidence in its estimate are recorded as separate properties:

Peer-reviewed, preprint and first-party report are publication types rather than confidence grades. A randomised design can still receive moderate confidence when it is small or at high risk of bias; a large quasi-experiment can provide stronger decision evidence than a weak experiment.

Where the evidence ceiling sits

A productivity-focused systematic review identified 39 peer-reviewed primary studies published through December 2024, using a broader productivity definition than this review’s quantitative comparator rule. A newer preregistered meta-analysis of 14 productivity studies estimated a moderate pooled effect (Hedges’ g = 0.33) with extreme heterogeneity (I2 = 99%); laboratory effects were larger than enterprise and open-source effects. Later field and repository studies expand the pool, but the number of independent causal estimates remains much smaller than the total literature.

The 59-study empirical base includes measured outcomes that determine whether local speed survives into useful delivery. Studies were assigned once, according to their principal outcome:

Included-study groupStudiesRole in this review
Task time, throughput and delivery40Produces the 42 quantitative effect estimates and the three central reference values
Code review and collaboration7Tests whether generation speed displaces work into review or changes pull-request outcomes
Quality, security and maintainability8Measures defects, vulnerability, technical debt and downstream evolution costs
Skills, experience and interaction cost4Measures learning, perceived effectiveness and the hidden work of directing and checking AI

Confidence profile

ConfidenceIncluded studiesInterpretation
High11Credible comparator, sufficiently direct outcome and no identified material risk that reverses the result
Moderate35Useful comparative evidence with residual bias, confounding, indirectness or precision limitations
Low13Historical baseline, self-report, organisational case or counterfactual estimate
Not graded56 documentsContext, mechanism, methods, benchmarks and secondary synthesis rather than included effect studies
Study register: 19 adjacent-outcome studies
Study IDStudy and principal outcomeDesignConfidence
ST-41RovoDev Code Reviewer — review usefulnessLarge-scale online evaluationModerate
ST-42Is Agentic Code Review Helpful? — review qualityComparative evaluationModerate
ST-43When Code Authors Are Agents — pull-request collaborationLarge-scale observational studyModerate
ST-44Comparing AI Coding Agents — pull-request acceptanceTask-stratified observational studyModerate
ST-45Early Adoption by GitHub Projects — adoption and reviewRepository studyModerate
ST-46AI-generated Pull-request Descriptions — review outcomesComparative repository studyModerate
ST-47Automated Code Review in Practice — review behaviourIndustrial observational studyModerate
ST-48GitHub Copilot and Code Quality — correctness and readabilityRandomised controlled trialHigh
ST-49AI-assisted Programming and Maintenance BurdenComparative repository studyModerate
ST-50Debt Behind the AI Boom — technical debtLarge-scale repository studyModerate
ST-51AI-assisted Assessment of Coding PracticesIndustrial deployment evaluationModerate
ST-52AI-assisted Fixes to Review CommentsLarge-scale industrial evaluationModerate
ST-53Asleep at the Keyboard? — security weaknessesControlled security evaluationModerate
ST-54AI Code Assistants and Security API UsageRepository studyModerate
ST-55GoodVibe — generated-code securityControlled security evaluationModerate
ST-56AI Assistance and Coding-skill FormationRandomised controlled trialHigh
ST-57AI Code-assistant Use at IBMEnterprise mixed-method studyModerate
ST-58Dear Diary — workplace experienceRandomised controlled trialHigh
ST-59Reading Between the Lines — interaction costBehavioural modelling studyModerate

Moving substantially beyond this empirical base is possible only by broadening again—to benchmark-only model evaluations, student exercises without a human baseline, additional vendor surveys or narrow task studies. Those documents can explain mechanisms and boundaries, but counting them as equivalent productivity studies would reduce rather than improve the review’s precision.

A meta-analysis was not appropriate because the studies measure different outcomes—task time, completed work, commits, deployment velocity and estimated project duration—and examine materially different systems. Results are therefore synthesised narratively and normalised to output multipliers only where the underlying comparison permits it. Causal interpretation is reserved for estimates with a credible counterfactual and sufficiently direct outcome.

Implementation-stage planning scenarios at three levels of AI autonomy Manual development is the one-times baseline. Assisted development is modelled at 1.25 times with higher evidence confidence, while spec-driven execution at two times and agent-native implementation at 4.5 times rely on low-confidence case evidence. These are not release-level forecasts. Implementation gains grow—and evidence weakens 1.25×4.5× manualprompt assistspec + humanagent-native loop baselinestronger evidencecase evidencecase estimates
Implementation-stage scenarios, not release forecasts. The first level is supported by controlled studies; the larger numbers rely mainly on organisational cases using different metrics. In the largest matched event study, observed release output increased by approximately 1.1×, 1.2× and 1.3× across successive tool generations.

Outcome definitions

A percentage can describe less time, more throughput, more commits or more output judged valuable. Those are not interchangeable.

Commits and lines of code are especially weak proxies. An agent can generate more of both without increasing production value. Deployment velocity, completed work items, task time and escaped defects are more useful, but even these need context.

Results by degree of automation

The findings are grouped by the three operating models defined in the abstract, each representing a different degree of automation. A separate downstream-conversion section then tests how gains in coding activity survive into releases and actual use.

1. Assisted development: task completion and local output

The strongest evidence concerns assistants operating inside otherwise conventional human-led workflows. The measured effects are heterogeneous, and tool choice changes the interaction model without determining the outcome by itself. For a practical comparison of those interaction models, see GitHub Copilot vs Claude Code vs Cursor.

A randomised trial involving 96 Google engineers found that AI assistance reduced time on a complex enterprise task by an estimated 21%, although the confidence interval was wide. Three field experiments at Microsoft, Accenture and a Fortune 100 company covered 4,867 developers and found a 26.08% increase in completed tasks.

A 109-participant controlled experiment found that ChatGPT helped with simple coding puzzles but did not improve efficiency or quality on a more typical software-development task. A later controlled study of 69 participants found no strong overall time effect, but AI users achieved more than twice the median task completeness and scored 12.5% lower on code-ownership questions. Together these results show why task completion, time and understanding should be reported separately.

Two smaller crossover experiments add qualified positive evidence. In a 32-person ICSE study, an in-IDE code-understanding assistant helped participants complete 0.47 more subtasks than web search, but did not significantly change completion time or comprehension. An 18-person JavaScript study found that participants using ClueBot finished faster and made fewer errors; its small student-heavy sample and high risk of bias limit generalisation.

A larger quasi-experiment at Ant Group compared matched teams during the staged introduction of CodeFuse. Across a sample of 1,219 programmers, code output increased by 55%. More conservative specifications found gains of roughly 13–22% in tasks completed and related workflow measures. Effects were concentrated among junior developers, and lines of code remained the principal outcome.

A public-sector study by GovTech Singapore reported 21–28% faster coding and task completion with GitHub Copilot. The direction is consistent with the larger controlled studies, although the disclosed design provides less protection against selection and reporting effects.

METR examined that contrasting setting. Sixteen experienced open-source developers completed 246 real tasks in repositories they had worked in for an average of five years. With early-2025 AI tools available, they were 19% slower. Before the study they predicted a 24% speed-up; afterwards they still believed they had been 20% faster.

A separate Anthropic randomised trial found only a small, statistically non-significant task-speed improvement, while participants using AI scored 17% lower on a subsequent mastery test. This does not establish a long-term productivity loss, but it identifies skill retention as an omitted outcome in most speed studies.

Enterprise studies based mainly on self-report provide useful triangulation but weaker effect estimates. An IBM study of 669 developers found heterogeneous perceived productivity effects from watsonx Code Assistant, while a Microsoft randomised diary study found that 84% reported positive changes in daily work practices without a corresponding change in trust in generated code.

Interaction studies expose work that completion metrics miss: prompting, inspecting, editing and recovering from suggestions. Microsoft’s CUPS study of 21 programmers made those costs observable. A randomised study of proactive programming assistance found benefits but strong interface effects, while a controlled comparison found that agents reduced effort and completed some otherwise inaccessible tasks without removing the need to understand their behaviour (Code with Me or for Me?). A five-day professional field study similarly found that poorly timed proactive assistance could interrupt flow (ProAIDE), and a comparison of GitHub Copilot and Windsurf treated effectiveness as a combination of usefulness, control and coordination rather than generated volume (AI-Assisted Collaboration).

Adoption evidence helps explain these mechanisms but does not establish delivery gains. GitHub’s telemetry-linked study of 2,047 respondents associated higher completion acceptance with higher perceived productivity (Measuring GitHub Copilot’s Impact on Productivity). A peer-reviewed 410-developer usability survey identified fewer keystrokes, faster completion and syntax recall as leading benefits, but control and non-functional requirements as common failures. A 57-developer enterprise survey found that production use depended on repository context, trust and integration requirements (Vukovic et al.).

Smaller controlled studies reinforce the importance of task selection. At ANZ Bank, a six-week controlled study reported 42.3% faster completion on Python exercises. A 24-person within-subject study found higher requirements completed per minute with both conversational and completion interfaces, although its tasks were short and only nine participants were professionals (Weber et al.). In a 2026 brownfield-onboarding experiment with 24 developers, Copilot Agent reduced mean completion time by 61.7% relative to Copilot Ask and workload by 57.4%, without a significant correctness improvement.

Industrial completion systems provide credible day-to-day activity measures, although these are not product outcomes. Meta’s CodeCompose supplied 8% of accepted code for 16,000 users (FSE 2024 deployment study); extending it to multi-line suggestions nearly doubled keystrokes saved from 9% to 17% (large-scale experiment). The UK Government’s trial made 2,500 licences available and combined telemetry with surveys; participants reported approximately one hour saved per working day, but the result was self-reported rather than a causal time study.

2. Spec-driven execution: bounded delegation

At the second level, the engineer defines an outcome, constraints and acceptance checks; the agent performs multi-file implementation; the engineer still reads and owns the change. Evidence here is less controlled because organisations usually adopt several changes at once. How to Choose an AI Development Framework compares the main frameworks used to structure this kind of specification and delegation.

Salesforce reports that work items completed per developer rose 50.8% year over year, while its “Effective Output” measure rose 151.3%. Engineers remained responsible for structuring problems and evaluating results, but the organisation had also standardised tooling, expanded access and changed workflows. The figures are not an isolated test of specification.

An Amazon Prime Video Financial Systems team provides a more specific example. A senior engineer spent three weeks decomposing work and writing detailed requirements before a ten-day sprint. The team used spec-driven development for complex features and direct agent assistance where requirements were already clear. It reduced a 90-week estimate to 24 weeks—approximately 3.75× acceleration—and produced nearly six times the baseline commits.

The study conditions limit external validity: six engineers had no on-call work, no other projects, minimal meetings and no context switching. Specification was part of the system, not the only intervention.

Cisco reports teams achieving up to 3× productivity in technical-debt workflows. Its agent investigates a quality issue, produces a written plan, implements in a fresh session and opens a pull request for human review. The work is unusually deterministic, with SonarQube providing a ready-made issue description and verification layer.

Evidence on specification quality shows why “more detail” is not sufficient by itself. Copilot generated both repayments and reproductions of technical debt when prompted with 1,140 variants derived from TODO comments (O’Brien et al.). By contrast, Meta found that pre-computing context across 4,100 files reduced agent tool calls by 40%, while a separate internal agent reduced approximately ten hours of performance-regression investigation to 30 minutes (capacity-efficiency case). These first-party cases suggest that explicit context and executable feedback reduce search cost; they do not isolate specification as a causal multiplier.

A peer-reviewed ICSE 2026 study examined an expert-led, AI-implemented workflow on PicoScenes, a 1.52-million-line industrial system with a ten-year history. The authors report 68.3% less feature-implementation time, alongside 28.2% lower mean cyclomatic complexity and 50% fewer defects. The comparison concerns one system and a bundled development method, but it provides unusually detailed evidence that specified delegation can extend beyond small greenfield tasks.

Anthropic’s internal mixed-methods study offers a second organisational data point. Engineers self-reported a 50% average productivity gain as Claude use expanded, while internal telemetry showed a 67% increase in merged pull requests per engineer per day. The before-and-after design cannot isolate Claude from team growth, task selection or workflow change.

Review finding: reported outcomes support a 1.5–2.5× range for spec-driven execution with human review. The midpoint is an evidence-informed interpolation, not a controlled effect estimate. Basis: E-21, E-29, E-30, E-38, E-39.

3. Agent-native delivery: production cases

The largest claims concern teams that restructure repositories, planning, testing, review and parallel execution so agents can operate for longer with less synchronous human attention. Where that design extends to several specialised agents, Multi-Agent Orchestration Frameworks Compared maps the main architectural categories and their reliability trade-offs.

AWS reports that, among 25 teams combining new tools with new practices, the median gain was 4.5× normalised deployment velocity, with some teams exceeding 10×. Its Bedrock pathfinder reported 20× normalised commit velocity, but this involved six senior engineers rebuilding a system under highly unusual conditions. Deployment velocity is the more useful number.

OpenAI describes an internal product built with no manually written code: approximately one million lines and 1,500 merged pull requests directed initially by three engineers over five months. The team estimates it took one-tenth of the manual development time. This is a serious production account with active users, but the 10× figure remains the team’s counterfactual estimate rather than an experiment.

Salesforce reports a 33-endpoint migration completed in 13 days rather than an estimated 231 person-days—an 18× task-specific acceleration. It used reference implementations, reusable rules, isolated environments and autonomous build-fix-validate loops. Migration work is repetitive and parallelisable, limiting the result’s generalisability to novel product development.

Amazon reports a broader but similarly task-specific transformation programme: Amazon Q Developer assisted upgrades of tens of thousands of production Java applications, with an estimated 4,500 developer-years saved. The estimate derives from migrated dependencies and assumed manual effort rather than a concurrent control group, so it is evidence of scale rather than a transferable acceleration factor.

Usage telemetry also shows that agentic work remains strongly shaped by human expertise. Anthropic’s analysis of approximately 400,000 Claude Code sessions found that people made most planning decisions while the agent made most execution decisions; greater domain expertise was associated with a higher probability of verifiable success.

Agent benchmarks define a capability ceiling rather than a workplace-productivity effect. SWE-Lancer maps more than 1,400 real freelance tasks to approximately $1 million of historical payouts. The ICLR 2024 SWE-bench study and SWE-agent show that repository tools and execution feedback can materially improve issue resolution. SWE-bench Multimodal exposes weaker performance on visual JavaScript tasks, while GitTaskBench adds workflow-driven repository use and an economic metric. The paired SWE-Skills-Bench result is a useful constraint: 39 of 49 packaged skills produced no pass-rate gain and average improvement was only 1.2%. None of these results can be converted directly into hours saved without a human workflow comparator.

Review finding: reported agent-native workflows cluster around 3–10×. AWS’s 4.5× median is the most suitable central reference in the available case evidence. Confidence remains low because the studies are observational, heterogeneous and frequently vendor-published. Basis: E-29, E-31, E-38E-42.

Study results by degree of automation

The evidence changes in both magnitude and strength across the three operating models. The graph places selected implementation-stage estimates on a logarithmic ratio scale. Where a study reports time saved, the value is converted to potential throughput using the reciprocal relationship defined under outcome definitions. Every label resolves to the corresponding E entry. The values remain heterogeneous and are not pooled.

Reported implementation effects by degree of automation Selected results are grouped into assisted development, spec-driven execution and agent-native delivery. The horizontal axis is logarithmic. Reported effects rise with automation while evidence confidence generally falls. A green vertical band marks the more conservative 1.1 to 1.3 times release-output planning range. Reported gains rise as evidence confidence falls High confidence Moderate Low 1.1–1.3× release range 1. Assisted development E-04 · METR · 0.84× E-18 · CodeFuse · 1.13–1.22× E-02 · field trials · 1.26× E-01 · Google · 1.27× E-28 · GovTech · 1.27–1.39× 2. Spec-driven execution E-30 · Anthropic · 1.67× E-29 · Salesforce · 2.51× E-39 · Cisco · up to 3× E-21 · PicoScenes · 3.15× E-38 · Prime Video · 3.75× 3. Agent-native delivery E-31 · AWS median · 4.5× E-40 · OpenAI estimate · 10× E-41 · Salesforce migration · 18× 0.5× 16× reported implementation effect relative to baseline (log scale)
Selected estimates assigned to the closest operating model in this review. Time reductions are converted to potential throughput where appropriate. Measures include task completion, workflow output, effective output, implementation time, deployment velocity and counterfactual project estimates; they should be compared within operating models and not averaged. The green band is the 1.1–1.3× end-to-end release-output planning range, not a confidence interval. The larger Level 2 and Level 3 values mainly come from bundled organisational cases with low-confidence historical or counterfactual baselines.

Moderators and downstream effects

The observed effect is not determined by the degree of automation alone. Project context, downstream capacity, programming-language feedback, developer experience and verification strength change both the attainable gain and the risk carried into later stages.

Greenfield and brownfield development

Project context is a major source of heterogeneity. Greenfield work begins without legacy constraints, while brownfield work changes an existing system whose behaviour, interfaces and operational history may be only partly documented. The studies do not support a single greenfield-versus-brownfield multiplier.

The mechanism changes with the degree of automation. At Level 1, repository familiarity determines whether suggestions save search and typing or add context and verification overhead. At Level 2, specifications and executable acceptance checks can make structured brownfield changes unusually tractable. At Level 3, the strongest cases come from greenfield systems or repetitive migrations whose environments, tests and work units were redesigned for autonomous execution.

Project contextRepresentative evidenceObserved effectInterpretation
Bounded greenfield taskControlled ChatGPT coding experimentHelped on coding puzzles; no efficiency or quality gain on typical development workTask realism changes the measured effect even within one experiment
Production greenfield systemOpenAI agent-built productEstimated 10× reduction in development timeLarge system claim, but based on a counterfactual rather than a control group
Unfamiliar brownfield codeTwo controlled student studies35% faster in one study; more tests passed in bothImplementation improves without corresponding evidence of deeper code comprehension
Familiar, mature brownfield codeMETR experienced-maintainer trial19% slowerTacit knowledge and verification overhead can outweigh generation speed
Structured brownfield changeCodeFuse, PicoScenes and migration casesModerate gains to multi-fold accelerationExisting code can be highly tractable when changes repeat and tests provide an oracle
Repository-wide agent adoptionCursor matched longitudinal studyVelocity increase was large but transientEarly output gains may be offset by persistent complexity and review costs

Two small controlled brownfield experiments make the distinction clearer. Ten undergraduates modifying an unfamiliar legacy web application completed tasks 35% faster and made 50% more solution progress with Copilot. A replication with 18 graduate students also found significantly lower task time and more tests passed, but no improvement in code comprehension. These results differ from METR’s slowdown among expert maintainers of large repositories they knew well.

Review finding: lifecycle status alone is a poor predictor. The strongest positive brownfield results occur when the task is deterministic, the relevant context can be supplied and correctness is executable. Brownfield work becomes less favourable when success depends on tacit architecture knowledge, broad side effects or expensive human verification. Basis: E-04, E-18E-22, E-39E-42.

Downstream conversion: from code activity to releases and use

This conversion is a cross-level outcome rather than a fourth operating model. At Level 1, local task gains must survive integration and review. Levels 2 and 3 can produce much larger increases in implementation activity, but place correspondingly greater demand on validation, release and adoption capacity.

The largest direct study of that conversion followed more than 100,000 GitHub developers using matched event studies and internal usage telemetry. Autocomplete, interactive agents and autonomous agents were associated with cumulative commit gains of 40%, 140% and 180%, respectively. The final effect was much smaller: the 180% commit increase became 50% more projects and 30% more releases, while analysis across four application marketplaces found no increase in total usage. This is the clearest evidence that code activity, shipped software and realised value are different outcomes.

Two unusually large studies find smaller average effects and strong experience differences. A peer-reviewed analysis in Science detected AI-supported Python functions across more than 30 million contributions from 160,097 developers. Within-developer models estimated a 3.6% increase in quarterly output at observed adoption levels, concentrated among experienced developers; early-career developers showed no significant gain. A regression-discontinuity natural experiment covering 187,489 developers found that Copilot access increased the relative share of coding activity by 12.37% and reduced project-management activity by 24.93%. The latter measures work allocation rather than total delivered output, but provides strong evidence that AI changes the composition of work.

Other quasi-experiments reinforce the same pattern. A generalized synthetic-control study estimated 6.5% higher project productivity and no code-quality change after Copilot adoption, but integration time increased by 41.6%. A 1,350-developer difference-in-differences study found approximately 12% more commits and 6% more active repositories, alongside a 2% rise in copy-pasted code. A workplace study using a time discontinuity estimated 6.9–8.4% higher coding productivity and 3.1–4.3% fewer working hours without a measured quality decline. Finally, a Python-versus-R natural experiment estimated 17.82% more package releases and 51% more commits in the Copilot-supported ecosystem, with gains weighted toward maintenance; differences between language communities remain an important identification caveat.

Matched and longitudinal workplace studies are less uniformly positive. Uplevel compared approximately 800 developers with and without Copilot access and found no significant change in pull-request cycle time or throughput, alongside a 41% higher bug rate. A two-year NAV study covering 26,317 commits across 703 repositories similarly found no statistically significant post-adoption activity change. By contrast, a 13-month study of three agile teams found one stable team increased delivered story points by 59.1%, with the other teams and quality measures more mixed.

Vendor telemetry shows that more code can coexist with more downstream work. Faros’s 2025 dataset of 10,000 developers and 1,255 teams associated high AI adoption with 21% more completed tasks and 98% more merged pull requests, but also 91% longer review time and no organisation-level DORA improvement. Its later two-year study of 22,000 developers and 4,000 teams found 33.7% more completed tasks and 16.2% faster merges, alongside 861% more code churn and 242.7% more incidents per pull request. These studies disclose large samples but remain observational and are published by an engineering-intelligence vendor.

Other large datasets point in the same two-sided direction. DX reports, from 135,000 developers in 425 organisations, an average 3.6 hours saved per week and 60% more pull requests among daily users. Jellyfish reports that the highest-adoption cohort across more than 200,000 engineers shipped 56% more epics per 100 engineers and used 25% fewer person-months per epic. Neither comparison randomly assigned AI use, so selection, team maturity and task mix remain plausible explanations.

Developer surveys explain why perception and telemetry diverge. In Stack Overflow’s 2024 survey, 81% named productivity as a benefit. By 2025, 69% of agent users reported increased productivity, yet 46% distrusted AI accuracy and 45% said debugging AI-generated code was time-consuming. These are valuable population signals, not estimates of causal speed.

Large first-party surveys provide adoption context rather than causal estimates. GitHub’s 2024 international survey documents perceived effects on flow, cognitive effort and collaboration. JetBrains’ 2025 Developer Ecosystem survey and Atlassian’s 3,500-person developer-experience survey broaden that picture. Anthropic’s analysis of 500,000 coding interactions found that Claude Code use was predominantly automation rather than augmentation. None of these designs establishes an output multiplier.

The same distinction matters outside engineering telemetry: an organisation can report faster local work without obtaining adoption or financial value. The Enterprise AI Reality Check examines that wider gap between tool deployment, workflow change, measurement and realised return.

Review finding: the controlled evidence centres near 1.25× output, with a range that includes negative effects. Large comparative studies then show substantial attenuation between code activity and shipped value. In the largest matched event study, release effects were approximately 1.1× for autocomplete, 1.2× for interactive agents and 1.3× for the cumulative autonomous-agent stack. For budgeting and time-to-market analysis, 1.1–1.3× release output, centred on 1.2×, is the more defensible base range; 2–4.5× remains implementation-stage upside requiring workflow redesign and independent verification. Task fit, repository maturity, developer familiarity and downstream capacity explain more of the variation than tool availability alone. Basis: E-01E-20, E-23E-28, E-32E-37.

Programming language is a moderator, but “strongly typed is faster” is not yet established

The language hypothesis is plausible: a compiler or type checker gives an agent a cheap, deterministic signal before a human reads the patch. Google’s production completion system supports the mechanism. In Go, approximately 8% of raw suggestions failed to compile; semantic checks filtered 80% of those failures, and acceptance grew by 1.9× after the checks were introduced, compared with 1.3× in languages without them. This is evidence that compiler-like feedback improves an AI workflow—not that static typing alone increases end-to-end developer productivity.

Its role grows with automation. At Level 1, compiler and linter feedback filters suggestions before the developer accepts them. At Level 2, those checks become executable acceptance constraints for a multi-file change. At Level 3, they form part of the agent’s build–test–repair loop and reduce how often a human must inspect obviously invalid states.

Cross-language evaluations show genuine variation, but not a simple static-versus-dynamic ranking. An evaluation across 19 programming languages found above-average results for several statically typed, lower-abstraction languages. A study of all 2,033 LeetCode problems found substantial correctness differences across C, Java, JavaScript and Python. The newer SWE-bench Multilingual, Multi-SWE-bench and OmniCode evaluations likewise show materially different resolution rates by language and task. In OmniCode, for example, SWE-agent’s best Java test-generation result was only 20.9%.

Language familiarity and training-data density can move in the opposite direction to type-system strength. A controlled cross-language trajectory study found large differences in token consumption across Python, Java, Rust and OCaml, with agents repeatedly producing non-compiling solutions in less familiar languages. A Java unit-test study found that type-rich context helped, but generated tests still suffered from compilation and oracle errors (EASE 2024 evaluation). TypePilot provides more direct evidence that a Scala type system can constrain insecure generations, although it evaluates generated-code robustness rather than developer time.

Review finding: treat the language as part of the verification harness. Static types, exhaustive compilers and mature linters should reduce the cost of rejecting a bad change, while popular ecosystems may improve generation quality through better training coverage. No identified field experiment isolates static typing as a causal productivity multiplier.

Developer experience changes where AI helps

“Experience” is not one variable. General coding tenure, familiarity with a particular repository and knowledge of the problem domain can produce different effects. In assisted-development trials, less-experienced developers often gain more. The three randomised field experiments covering 4,867 developers found higher adoption and larger productivity gains among less-experienced developers. In the CodeFuse field experiment, the measured output increase was concentrated among junior programmers, despite minimal junior–senior differences in suggestion acceptance.

Repository familiarity can reverse that pattern. METR’s experienced maintainers were 19% slower on repositories they had worked in for an average of five years, while two small studies of developers modifying unfamiliar brownfield applications found faster completion and more tests passed. This does not show that novices outperform experts. It suggests that assistance is more valuable when the tool supplies missing context, but can impose prompting and verification overhead when a developer already holds a strong mental model of the code.

At higher degrees of automation, domain expertise becomes more—not less—important. In Anthropic’s analysis of approximately 400,000 Claude Code sessions, novice-rated sessions reached verified success 15% of the time, compared with 28–33% for intermediate-to-expert sessions. When sessions encountered trouble, verified recovery rose from 4% for novices to 15% for experts. The authors found that task-specific expertise predicted success more strongly than working in a software occupation, although the outcomes were inferred from transcripts and telemetry rather than production value.

The relationship is therefore not a simple equalising effect. A study of more than 160,000 open-source developers found that observed contribution gains were concentrated among experienced developers, with no significant gain among early-career developers. Different tasks, experience measures and workflows explain part of the apparent contradiction.

Qualitative and longitudinal studies describe the same shift from implementation toward supervision. Experienced professionals retain control over architecture and quality rather than relying on unconstrained “vibe coding” (Huang et al.). A six-month panel found 84% reporting productivity improvement at both waves, while worsened developer experience nearly doubled in the matched cohort (Vella and Blincoe). Microsoft’s SPACE study of more than 500 developers found benefits concentrated in routine work and weaker evidence for collaboration, while repository traces show use extending well beyond code generation (self-admitted open-source usage study).

Review finding: less-experienced developers may gain more from Level 1 assistance on bounded tasks, while experts may gain little—or slow down—when assistance disrupts work in a familiar repository. At Levels 2 and 3, domain knowledge improves specification, steering, exception handling and verification. Experience should therefore be measured as coding tenure, repository familiarity and domain expertise, not a single seniority label.

Automated review is necessary—but not a 4× cause

Review responsibility changes with the degree of automation. At Level 1, the developer inspects suggestions while implementing the change. At Level 2, a human reviews and owns a completed agent-authored change. At Level 3, automated review becomes a first-pass control before accountable human approval. The third-level increase cannot be attributed to automated review alone, and the available evidence does not isolate review as a causal factor.

OpenAI’s guide to AI-native engineering says automated review gives engineers more confidence that major bugs are caught, but does not necessarily make the pull-request process faster. Humans still need to understand the implications and own the merge.

A 2026 study of 31,073 CodeRabbit review comments across 10,191 pull requests found that only 36.4% were accepted, while 56.3% were rejected. Common reasons included false positives, redundant comments and suggestions outside the intended scope.

Atlassian reports a more favourable result from a tightly integrated review system. Its year-long evaluation across more than 1,900 repositories found that RovoDev reduced pull-request cycle time by 30.8% and human-written review comments by 35.6%; 38.7% of its comments triggered a subsequent code change. The paper was accepted to the ICSE 2026 Software Engineering in Practice track, but the evaluation was observational and conducted by the tool’s developer.

The review bottleneck is real. A GitLab survey of 1,528 developers and technology buyers found that 85% agreed AI had shifted the bottleneck from writing code to review and validation. But survey agreement is not proof that an automated reviewer removes it.

Large-scale repository studies show why review results depend on task and governance. A study of 40,214 human- and agent-authored pull requests found materially different collaboration patterns when agents authored the change. A separate analysis of 7,156 agent-authored pull requests found that acceptance varied more by task type than by agent: documentation reached 82.1%, compared with 66.1% for new features. Early-adoption research covering 25,264 agentic pull requests found that most repositories still produced only one or two agent-authored pull requests over three months and relied predominantly on a single human reviewer.

The broader repository record confirms that acceptance is not equivalent to wholesale integration. AIDev catalogues 932,791 agent-authored pull requests across 116,211 repositories. A study of 9,427 agentic pull requests found that core developers more consistently required passing CI before acceptance (Cynthia et al.), while a longitudinal comparison tracks changing mergeability, complexity and defect proneness in agent- and human-authored pull requests (Morovati et al.). PatchTrack found a median of only 25% of suggested code integrated across 338 pull requests; developers usually extracted, adapted or rejected parts of a patch. These studies do not estimate causal speed, but show why generated or accepted code volume can overstate delivered value.

Interventions at narrower review stages show both useful automation and displacement. Across 18,256 pull requests, AI-generated descriptions were associated with shorter review time and a higher merge probability. At Google, ML-generated edits now resolve 7.5% of reviewer comments, an effect estimated to save hundreds of thousands of engineer hours annually at that scale. AutoCommenter was deployed across four languages to tens of thousands of Google developers and produced a measurable workflow benefit, although the published study does not translate it into an end-to-end speed multiplier (industrial evaluation).

Negative and null results are equally instructive. In Beko’s ten-project deployment, developers resolved 73.8% of automated comments, but median pull-request closure increased from 5 hours 52 minutes to 8 hours 20 minutes (Automated Code Review in Practice). Meta’s review-comment patch system initially made reviewers more than 5% slower; moving suggestions from reviewers to authors removed that regression, and the final system achieved a 19.7% actionable-to-applied rate (randomised safety trials and production experiment).

The high-performing cases combine review with deterministic tests, static analysis, isolated environments, repository rules, short feedback loops and human ownership. Review is one control in a harness; How to Sandbox AI Agent Code compares practical isolation boundaries for the execution layer.

Review finding: automated review is a scaling control whose importance increases with the degree of automation—not a demonstrated cause of a 4× gain. At Level 1, narrow review automation can remove routine work but can also interrupt or slow the developer. At Level 2, it should reject deterministic defects before accountable human review. At Level 3, it is necessary to prevent generated volume overwhelming review capacity, but current systems still reject or fail to action many comments and do not remove the need for human approval. No identified study isolates automated review as an end-to-end productivity multiplier. Basis: ST-41ST-47, ST-51ST-52.

Quality effects are mixed and may emerge later

Speed and throughput do not fully describe the productivity effect, and the quality risk changes with the degree of automation. At Level 1, a developer can reject a bad suggestion before it becomes a change. At Level 2, larger agent-authored patches increase the amount of behaviour a human must verify. At Level 3, generation and initial review both scale, making independent tests, security controls, monitoring and ownership part of the operating model rather than optional safeguards. A GitHub randomised controlled trial with 202 experienced developers found that Copilot users were 53.2% more likely to pass all ten unit tests. Blind review also found small but statistically significant improvements in readability, reliability, maintainability and conciseness.

Longitudinal evidence is less uniformly positive. A matched difference-in-differences study of 806 repositories found that adopting Cursor produced a large but transient increase in development velocity, while static-analysis warnings and code complexity increased persistently. Another open-source study found that, after Copilot adoption, core developers reviewed 6.5% more code while producing 19% less original code, consistent with maintenance work being redistributed rather than removed.

A controlled maintenance experiment provides a useful boundary on that interpretation. When 75 participants evolved code written by someone else, the speed advantage for maintaining AI-assisted code was small and statistically unreliable. A separate analysis of 304,362 verified AI-authored commits found that more than 15% introduced at least one detectable issue, and 24.2% of the tracked issues remained at the latest repository revision.

Security evaluations consistently reject the assumption that compilable means safe. Copilot produced vulnerable output in approximately 40% of 1,689 programs in the peer-reviewed “Asleep at the Keyboard” study. A 2024 replication found a lower but still substantial 27.25% vulnerable share. In real repositories, security weaknesses affected 29.5% of observed Python snippets and 24.2% of JavaScript snippets (Fu et al.), while a 44-developer controlled study found no significant improvement in secure API use (Mousavi et al.).

The surrounding risks create verification work. Across 16 coding models, package hallucinations averaged 19.6%. A 2024 USENIX evaluation found that LLMs could assist code analysis but remained unreliable across tasks (Fang et al.). Newer security-by-default work improved secure generation by up to 2.5× across six models and four security-critical languages, but this is a model-level benchmark rather than evidence that deployed teams experience fewer incidents (GoodVibe). The strongest conclusion is therefore not that AI necessarily increases production bugs, but that generated code retains material defect and vulnerability risk and requires an explicit verification budget.

Automated tests can recover part of this cost, but generated tests need their own oracle. TestPilot generated JavaScript tests with median statement coverage of 70.2% in its 2024 journal publication. ChatTESTER increased compilable tests by 34.3% and tests with correct assertions by 18.7% relative to direct ChatGPT generation (FSE 2024 study). A broader evaluation across 17 Java projects found that prompt content and model choice materially changed outcomes (Yang et al.). The productive mechanism is therefore the executable feedback loop, not test text generation alone.

Repository telemetry indicates that these costs may emerge slowly. GitClear’s analysis of 153 million changed lines found rising churn and copy-paste activity as AI adoption grew, although it could not identify AI authorship causally. DORA’s 2024 cross-organisation analysis associated a 25% increase in AI adoption with better documentation, code quality and review speed, but also 1.5% lower delivery throughput and 7.2% lower stability. These are associations, but they illustrate why local speed and delivery performance must be measured separately.

Google’s 2025 DORA research develops that finding: AI appears to amplify existing organisational strengths and weaknesses, with higher adoption associated with both greater throughput and greater instability. The appropriate comparison is therefore not simply with and without AI, but the same degree of automation under stronger and weaker delivery systems.

Collectively, these studies indicate that short-task quality can improve while repository-level complexity, persistent defects and expert review load still rise over time. This distinction is especially consequential in brownfield systems, where each accepted change becomes part of the next task’s context. AI-Assisted Web Development: Where It Helps and Where It Fails applies the same boundary to pattern-heavy work that is easy to verify and judgement-heavy work where plausible output can hide defects.

Review finding: faster generation does not inevitably reduce quality, but the evidence does not support treating quality as constant across degrees of automation. Level 1 trials can improve immediate correctness on bounded tasks while longitudinal studies still detect complexity, maintenance and security costs. At Levels 2 and 3, larger agent-authored changes expand the verification surface; maintaining quality requires independent tests, security controls, monitoring and accountable ownership to scale with generation. No identified study demonstrates that quality remains stable at the article’s 2× or 4.5× implementation scenarios. Basis: ST-22, ST-27, ST-48ST-50, ST-53ST-55.

Local coding speed can be lost at downstream delivery stages A pipeline runs from intent through generation, verification, review and deployment. Generation is accelerated first, shifting queues and risk into downstream stages unless the whole system is redesigned. A faster generator does not guarantee faster delivery intentscope + constraints generateaccelerates first verifytests + state reviewqueue moves here deployvalue + incidents unverified volume creates downstream rework The relevant outcome spans accepted intent through production operation.
Productivity gains survive only when verification, review and deployment scale with generation.

Productivity effect register and confidence assessment

IDsStudy and outcomeDesignEffectConfidenceMain limitation
ST-01
E-01
Google enterprise RCT — task timeRandomised field experiment21% less timeHigh96 engineers; wide confidence interval
ST-02
E-02
Microsoft/Accenture/F100 — completed tasksControlled field experiments+26.08%HighEffects varied between companies
ST-03
E-03
Rocks Coding, Not Development — task performanceControlled experimentCoding gain; no development-task gainHigh109 participants; assigned tasks
ST-04
E-04
METR — mature-repository task timeRandomised experiment19% slowerHigh16 expert maintainers; early-2025 tools
ST-05
E-05
PACIS assistant–agent trial — completion timeControlled experiment61.7% less timeHigh24 developers; assistant comparator
ST-06
E-06
ANZ Bank — Python-task completionControlled experiment42.3% fasterHighExercises followed tool training
ST-07
E-07
Weber et al. — requirements per minuteWithin-subject experimentSignificant increaseHigh24 participants; mostly junior
ST-08
E-08
More Code, Less Understanding — task completenessControlled experiment>2× median completenessHighNo strong overall time effect
ST-09
E-09
GILT — code-understanding progressCrossover trial+0.47 subtasks; no time gainModerate32 participants; high risk of bias
ST-10
E-10
ClueBot — JavaScript time and errorsCrossover experimentLess time; fewer errorsModerate18 participants; student-heavy sample
ST-11
E-11
Writing Code vs. Shipping Code — commits, releases and useMatched event study+180% commits; +30% releases; no usage gainModerateAdoption remains non-random
ST-12
E-12
Global diffusion and impact — quarterly contributionsWithin-developer fixed effects+3.6%ModerateAI authorship inferred on public GitHub
ST-13
E-13
Generative AI and the Nature of Work — work allocationRegression discontinuity+12.37% coding shareModerateTask mix, not final output
ST-14
E-14
Collaborative open source — project output and integrationGeneralised synthetic control+6.5%; integration time +41.6%ModerateAdoption inferred
ST-15
E-15
Copilot X — commits and repositoriesMatched difference-in-differences+12% commits; +6% repositoriesModerateCopy-paste also rose
ST-16
E-16
Workplace AI-coder deployment — productivity indexRegression discontinuity in time+6.9–8.4%ModerateTime discontinuities permit confounding
ST-17
E-17
Python–R natural experiment — releases and commitsDifference-in-differences+17.82% releases; +51% commitsModerateLanguage ecosystems may differ
ST-18
E-18
Ant Group CodeFuse — code and workflowMatched field experiment+55% code; +13–22% workflowModerateCode volume was primary
ST-19
E-19
Brownfield undergraduate study — task timeControlled experiment35% fasterModerateTen students; narrow application
ST-20
E-20
Brownfield replication — time and testsWithin-subject experimentFaster; more tests passedModerateNo comprehension gain
ST-21
E-21
PicoScenes — implementation timeComparative industrial case68.3% less timeModerateOne system; bundled method
ST-22
E-22
Cursor repositories — velocityMatched longitudinal studyLarge, transient increaseModerateObservable adopters only
ST-23
E-23
Meta multi-line completion — keystrokes savedLarge-scale online experiment9–17%ModerateLocal activity proxy
ST-24
E-24
Uplevel — PR cycle and throughputMatched observational studyNo significant changeModerateObservational assignment
ST-25
E-25
NAV — commit activityLongitudinal repository studyNo significant changeModerate25 users; self-selection
ST-26
E-26
Three agile teams — story pointsLongitudinal multi-case study+59.1% in one teamModerateOne stable comparison
ST-27
E-27
Echoes of AI — task timeComparative experiment30.7% lower median timeModerateParticipants chose workflow
ST-28
E-28
GovTech Singapore — task speedStructured field study21–28% fasterLowLimited comparative detail
ST-29
E-29
Salesforce agentic SDLC — effective outputBefore-and-after telemetry+151.3%LowMany simultaneous changes
ST-30
E-30
Anthropic adoption — merged PRsBefore-and-after telemetry+67%LowTelemetry plus self-report
ST-31
E-31
AWS 25 teams — deployment velocityMulti-team organisational telemetryMedian 4.5×LowNot randomised; limited method
ST-32
E-32
UK Government — reported time savedVoluntary workplace trialAbout one hour per dayLowSelf-reported time
ST-33
E-33
Meta CodeCompose — accepted-code shareDeployment telemetry8% of codeLowActivity proxy
ST-34
E-34
Faros 2025 — tasks, PRs and reviewObservational telemetry+21% tasks; +98% PRsLowReview time rose 91%
ST-35
E-35
Faros 2026 — delivery telemetryLongitudinal telemetry+33.7% tasks; 16.2% faster mergesLowChurn and incidents rose
ST-36
E-36
DX Q4 2025 — time and PR outputSurvey plus telemetry3.6 hours/week; +60% PRsLowNon-random usage
ST-37
E-37
Jellyfish — epics and effortAdoption-cohort analysis+56% epics; −25% effortLowHigh-adoption cohort may differ
ST-31
E-38
AWS Prime Video — forecast durationDisclosed case estimate3.75×LowComparator was a forecast
ST-38
E-39
Cisco — tech-debt workflowDisclosed production caseUp to 3×LowDeterministic issue class
ST-39
E-40
OpenAI — estimated project timeCounterfactual case estimate10×LowNo control
ST-29
E-41
Salesforce migration — task durationCounterfactual case estimate18×LowRepetitive, parallelisable work
ST-40
E-42
Amazon Q migration — estimated effortCounterfactual portfolio estimate4,500 developer-yearsLowDependency-based estimate

Limitations

This rapid systematic review was not protocol-registered and did not include a formal risk-of-bias instrument. Searches prioritised English-language, web-accessible studies and first-party disclosures; relevant unpublished evaluations may therefore be absent. The corpus is broad rather than exhaustively database-indexed. Publication bias is likely to be strongest in the agent-native tier, where organisations have an incentive to publish exceptional successes.

The evidence base is too heterogeneous for a pooled effect estimate. Controlled studies mostly evaluate individual assistance, while agent-native reports measure teams, workflows or individual migrations. Historical estimates can also change after a project begins, and commit or line-volume measures may reward activity rather than value. Greenfield and brownfield labels were assigned from the disclosed project setting and were not always prespecified by study authors. Experience measures also differ: job tenure, coding tenure, repository familiarity and task-specific expertise are not interchangeable. Programming-language studies predominantly measure correctness, compilation or benchmark success rather than developer time, so the language analysis is mechanistic and exploratory. The three-level autonomy taxonomy was applied for this review and has not itself been independently validated.

Benchmark quality adds a separate constraint. OpenAI found contamination and test-design problems severe enough to stop using SWE-bench Verified, then estimated that approximately 30% of SWE-Bench Pro tasks were broken in a separate agent-assisted audit and human annotation study. Benchmark progress shows that agents can perform longer tasks under controlled conditions; it cannot be converted directly into hours saved without a human workflow comparator.

Implications for future measurement

The literature would benefit from staged comparisons of representative work rather than further surveys of perceived speed. A robust organisational study would:

  1. Select repeated task classes with sufficient history for a baseline.
  2. Measure accepted intent through production operation, rather than coding time alone.
  3. Separate active human time, elapsed time and agent runtime.
  4. Include review rounds, escaped defects, rollback and rework.
  5. Compare similar tasks at assisted, spec-driven and agent-native levels.
  6. Report the model, repository characteristics, team composition, coding tenure, repository familiarity and domain expertise.

Relevant outcome measures include lead time, work items deployed, change failure rate, reviewer minutes, post-merge rework and production incidents. Commits can diagnose activity, but are not an adequate primary productivity outcome. For a worked example of holding variables constant, exposing uncertainty and separating raw results from interpretation, see Building the Ed-o-meter: Notes on Writing My Own LLM Benchmark.

Discussion and conclusion

The central finding is a conversion problem. AI can substantially increase implementation output, but the gain is often attenuated as work passes through understanding, integration, review, testing, release and adoption. The largest matched study makes that loss visible: 180% more commits became 30% more releases and no increase in application usage (E-11). The independent meta-analysis reaches the same broad boundary from another direction: average productivity improves, but estimated effects are far smaller in enterprise and open-source settings than in controlled laboratories.

For end-to-end planning, the largest matched delivery study observed approximately 1.1×, 1.2× and 1.3× release output across successive tool generations. This review therefore uses 1.2× as an evidence-anchored central planning case, driven principally by E-11 and triangulated against E-14E-17; it is not a pooled or universal effect. The larger figures describe narrower implementation stages: approximately 1.25× for conventional assistant use, as a conditional hypothesis for bounded spec-driven work and 4.5× as a low-confidence agent-native upside scenario. They should not be applied to the whole delivery system. Increasing autonomy does not itself establish causation; the largest outcomes combine a model with redesigned repositories, specifications, environments, tests, review and parallel execution.

The evidence also rejects a simple speed-versus-quality trade-off. Short, bounded tasks can become both faster and more correct. The recurrent risk is asymmetry: generation scales before verification does, shifting cost into integration time, review queues, churn, complexity, security findings and weaker code ownership. Speed without scaled verification can come at the expense of quality. Maintaining quality at higher generation rates requires robust review and verification, but the evidence does not show that these controls guarantee it—particularly at the 2× and 4.5× implementation scenarios.

Project lifecycle alone is a poor predictor. Greenfield work has less legacy coupling but often lacks historical baselines and inherited test suites. Brownfield work has more constraints, yet repetitive migrations and well-instrumented changes can be highly tractable. Experience does not move the effect in one direction: bounded Level 1 trials often favour less-experienced developers, familiar-repository experts can slow down, and domain expertise becomes more valuable as automation increases. The more useful distinctions are tacit versus explicit context, subjective versus executable acceptance and generated versus independently verified output.

For investment decisions, the unit of analysis should be the verified outcome in production. A credible business case should estimate model and platform costs, integration effort, reviewer capacity, rework, incidents and the proportion of coding-time savings that reaches release and customer use. Time-to-market claims should be based on lead time through production operation, not commits, lines of code or self-reported hours saved.

As automation increases, AI coding becomes a systems-engineering problem. The limiting factor is increasingly not how quickly a model can generate code, but how effectively an organisation can specify, verify, integrate and release correct work.


Evidence and document registers

The registers preserve the review chain from included study to source document. Confidence is assigned in the study registers above; supporting and contextual documents are not graded as productivity evidence.

A. Included empirical studies (59)

Show the 59 studies ordered by study ID

B. Systematic reviews and meta-analyses (2)

Show the 2 systematic reviews and meta-analyses

C. Supporting and contextual documents (54)

Show the 54 supporting documents alphabetically