Do AI coding assistants make teams faster?

Developers often feel faster with AI coding tools. Measured results are less certain, and the team’s delivery system decides much of what the tools are worth.

By the Webair engineering team7 min read

A code editor with a short coral suggestion at the cursor, about to extend into place

Not automatically. Developers often feel faster with AI coding tools, but feeling faster and delivering faster are different things, and controlled measurements have been much less certain than developers’ impressions. Google Cloud’s DORA research finds that much of what the tools are worth depends on the team’s delivery system: how it tests, reviews and releases code. So the useful question isn’t how fast the tool feels. It’s whether your team now ships more, with no more rework, and how you’d know.

What people report and what gets measured

Belief in the tools is high. In DORA’s 2025 survey of nearly 5,000 technology professionals, 90% said they use AI at work, and more than 80% believe it has made them more productive.1 Trust is lower: 30% have little or no trust in code written by AI.

Reports like these tell you what developers believe, and that matters, because it shapes which tools they use and how. They don’t tell you what changed in the work.

The two can be far apart. In a trial run by METR, a research non-profit, 16 experienced open-source developers expected AI to speed them up by 24% before they started. When they’d finished, they believed it had sped them up by 20%. Measured, their tasks had taken 19% longer with AI allowed than without.2 METR concludes that developers’ own estimates of their speedup can be very inaccurate.

That doesn’t make asking developers pointless. METR still thinks carefully chosen survey questions, alongside studies of how people spend their time, can give a useful signal.3 But a reported speedup is where a question starts, not the answer to it.

The study that found a slowdown, and what changed since

METR’s early-2025 study was a randomised controlled trial. Each developer listed real issues from a repository they had worked on for years: bug fixes, features and refactors averaging about two hours each. Each of the 246 issues was then randomly assigned to allow AI or not, so the comparison wasn’t skewed by which tasks developers chose to use AI on. The repositories were large and mature, averaging more than 22,000 stars and a million lines of code. The AI tools were mainly Cursor Pro with Claude 3.5 and 3.7 Sonnet, frontier models at the time.2

METR is careful about what the 19% shows. It doesn’t claim that AI fails to speed up most developers, and it thinks AI tools may well help in other settings, such as for less experienced developers or in unfamiliar codebases. Its results also suggest AI may help less where quality standards are very high and many requirements are implicit: the documentation, tests and formatting a project expects.2

In August 2025 METR started a second, larger study: 57 developers, 143 repositories and more than 800 tasks, with the latest tools. Ten of the developers came from the first study.3 Its estimates point the other way.

Change in task time with AI allowed, in METR’s two studies. METR calls the late-2025 estimates unreliable, and thinks the true speedup is probably larger.
StudyDevelopersChange in task timeConfidence interval
Early 20251619% longer2% to 39% longer
Late 202510 from the first study18% shorter38% shorter to 9% longer
Late 202547 new to the study4% shorter15% shorter to 9% longer

Neither late-2025 range rules out no effect at all, and METR doesn’t present these numbers as a finding. It calls the data an unreliable signal, for reasons that are informative in themselves:

  • more developers declined to take part because they didn’t want to work without AI
  • 30% to 50% of developers said they had held back tasks they didn’t want to do without AI
  • pay was lower: $50 an hour, against $150 in the first study
  • time records were unreliable for developers running several AI agents at once, who often worked on something else while an agent ran

The developers who expected most from AI, and the tasks they expected it to help with most, were the ones dropping out. So METR thinks its estimate is probably a lower bound on the true effect. From conversations with participants, it believes developers are likely more sped up in early 2026 than in early 2025, but says its data is only very weak evidence of how much.3 It now marks its 2025 results as out of date, and is changing the design of its study.

So a careful attempt to measure this had to change its design, because wider use of AI changed who would take part and how developers work. The 19% is a snapshot of early-2025 tools in one demanding setting, not a verdict, and the late-2025 numbers aren’t one either.

Why the team’s system matters more than the tool

DORA, the research programme on software delivery that Google Cloud runs, sums up its 2025 findings in one line:1

AI doesn’t fix a team; it amplifies what’s already there.

DORA, announcing its 2025 report

In DORA’s data, unlike in 2024, more AI use now goes with higher software delivery throughput and better product performance. It still goes with lower delivery stability. DORA’s explanation is that AI speeds up development, and that the extra speed can expose weaknesses further down the line. Without strong automated testing, mature version control and fast feedback, more change leads to instability. Teams with loosely coupled systems and fast feedback see gains. Teams held back by tightly coupled systems and slow processes see little or no benefit.1

These are associations in survey data, not controlled measurements. But the logic is straightforward. A task finished faster only reaches users sooner if review, testing and release can take the extra work. If they can’t, the time saved turns into a longer review queue, or into fixes after release.

DORA also finds a direct correlation between a high-quality internal platform and an organisation’s ability to get value from AI, one reason to be deliberate about what a platform team builds first.1 And if the constraint is a tightly coupled system, a better assistant won’t remove it. Loosening the system will, and that can be done piece by piece rather than in one rewrite.

Rolling out an assistant is the easy part. The hard part is making sure the rest of the path to production can use what it produces.

How to measure it in your team

Start from the outcome you want, not from the tool. If that’s getting changes to users sooner without more of them going wrong, DORA’s five software delivery metrics measure it. The first three measure throughput, and the last two instability.4

DORA’s five software delivery metrics
MetricWhat it measures
Change lead timeTime from a code commit to its successful deployment in production
Deployment frequencyHow often application changes are deployed
Failed deployment recovery timeTime to recover from a failed deployment
Change fail rateThe share of deployments that cause a failure in production
Deployment rework rateThe share of deployments that are unplanned work to fix bugs

Then:

  • Take a baseline. Record the metrics before anything changes, or now if the tools are already in use, and keep the definitions fixed: what counts as a deployment, a failure and rework.
  • Read throughput and instability together. DORA’s data links more AI use with higher throughput and lower stability. More deployments mean little if the rework rate rises with them.
  • Watch where the wait moves. If code gets written faster, the queue can move to review and testing. Track how long changes wait for review, and how often they come back.
  • Check quality, not only time. Some developers in METR’s later study said the quality of their work, and how much documentation and testing they did, differed with and without AI.3 Time on task misses that.
  • Ask developers, and use the answers well. Surveys tell you where AI helps and where it gets in the way, which is how you learn what to point it at. Just don’t treat an estimate of time saved as the measurement.
  • Give it an owner and a date. One person owns the measurement, and the result is reviewed on a set date.

In practice

Keep these metrics at team level. They describe how a delivery system performs. Used to rank individuals, they tend to change behaviour instead of measuring it.

The same discipline applies when AI is part of what you ship, not just how you build it: decide what good looks like before launch, and test against it.

Using AI coding tools isn’t the milestone anymore. Knowing what they’ve changed in your delivery is.

What to ask your team

The research so far doesn’t settle whether AI coding assistants make teams faster, and your team’s answer may differ from any study’s. Three answers show where you stand:

  • what your delivery metrics were before the tools arrived, and how they’ve moved since
  • whether throughput and rework moved together
  • where the time saved went: into more changes, into review queues, or into fixing what shipped

Then the question that matters: since we adopted AI coding tools, are we shipping more with no more rework, and what’s our evidence that the tools are the reason?

Sources

  1. Google Cloud, Announcing the 2025 DORA Report: State of AI-assisted Software Development, 23 September 2025. A survey of nearly 5,000 technology professionals worldwide, with more than 100 hours of qualitative data. DORA is a research programme run by Google Cloud, which sells AI tools. The full report requires registration; the figures here are from the announcement.
  2. METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, 10 July 2025. A randomised controlled trial: 16 experienced open-source developers, 246 tasks, February to June 2025. METR, a non-profit funded by donations, now marks these results as out of date.
  3. METR, We are Changing our Developer Productivity Experiment Design, 24 February 2026. 57 developers, 143 repositories and more than 800 tasks, from August 2025. METR considers these estimates unreliable because of selection effects.
  4. DORA, A history of DORA’s software delivery metrics, 2 January 2026.
  1. AI product engineering

    How to test an AI feature before it launches

    A demo shows that an AI feature can work. Before launch, you need evidence of how often it does, on the cases your users will bring.

    8 min read

    A balance scale with its coral pan raised, about to settle level
  2. Platform engineering

    Internal developer platforms: what to build first

    Nine in ten organisations in DORA’s 2025 research use an internal developer platform. DORA’s advice on what to build first: one journey developers repeat often, made clearly better, with speed and stability measured together.

    8 min read

    A signpost with its coral arm tilted up, about to swing level and point the way
  3. Legacy modernisation

    The strangler fig pattern: modernise without a rewrite

    Replace an old system one piece at a time while it keeps running. Moving the first piece is the easy part. Switching the old one off is the milestone.

    9 min read

    Railway points with the coral switch rail standing open, about to move across and send the track to the new line

Working on something similar?

Tell us what you’re building or fixing, and we’ll share how we’d approach it.