Do AI coding assistants make teams faster?
Developers often feel faster with AI coding tools. Measured results are less certain, and the team’s delivery system decides much of what the tools are worth.
By the Webair engineering team7 min read

Not automatically. Developers often feel faster with AI coding tools, but feeling faster and delivering faster are different things, and controlled measurements have been much less certain than developers’ impressions. Google Cloud’s DORA research finds that much of what the tools are worth depends on the team’s delivery system: how it tests, reviews and releases code. So the useful question isn’t how fast the tool feels. It’s whether your team now ships more, with no more rework, and how you’d know.
What people report and what gets measured
Belief in the tools is high. In DORA’s 2025 survey of nearly 5,000 technology professionals, 90% said they use AI at work, and more than 80% believe it has made them more productive.1 Trust is lower: 30% have little or no trust in code written by AI.
Reports like these tell you what developers believe, and that matters, because it shapes which tools they use and how. They don’t tell you what changed in the work.
The two can be far apart. In a trial run by METR, a research non-profit, 16 experienced open-source developers expected AI to speed them up by 24% before they started. When they’d finished, they believed it had sped them up by 20%. Measured, their tasks had taken 19% longer with AI allowed than without.2 METR concludes that developers’ own estimates of their speedup can be very inaccurate.
That doesn’t make asking developers pointless. METR still thinks carefully chosen survey questions, alongside studies of how people spend their time, can give a useful signal.3 But a reported speedup is where a question starts, not the answer to it.
The study that found a slowdown, and what changed since
METR’s early-2025 study was a randomised controlled trial. Each developer listed real issues from a repository they had worked on for years: bug fixes, features and refactors averaging about two hours each. Each of the 246 issues was then randomly assigned to allow AI or not, so the comparison wasn’t skewed by which tasks developers chose to use AI on. The repositories were large and mature, averaging more than 22,000 stars and a million lines of code. The AI tools were mainly Cursor Pro with Claude 3.5 and 3.7 Sonnet, frontier models at the time.2
METR is careful about what the 19% shows. It doesn’t claim that AI fails to speed up most developers, and it thinks AI tools may well help in other settings, such as for less experienced developers or in unfamiliar codebases. Its results also suggest AI may help less where quality standards are very high and many requirements are implicit: the documentation, tests and formatting a project expects.2
In August 2025 METR started a second, larger study: 57 developers, 143 repositories and more than 800 tasks, with the latest tools. Ten of the developers came from the first study.3 Its estimates point the other way.
| Study | Developers | Change in task time | Confidence interval |
|---|---|---|---|
| Early 2025 | 16 | 19% longer | 2% to 39% longer |
| Late 2025 | 10 from the first study | 18% shorter | 38% shorter to 9% longer |
| Late 2025 | 47 new to the study | 4% shorter | 15% shorter to 9% longer |
Neither late-2025 range rules out no effect at all, and METR doesn’t present these numbers as a finding. It calls the data an unreliable signal, for reasons that are informative in themselves:
- more developers declined to take part because they didn’t want to work without AI
- 30% to 50% of developers said they had held back tasks they didn’t want to do without AI
- pay was lower: $50 an hour, against $150 in the first study
- time records were unreliable for developers running several AI agents at once, who often worked on something else while an agent ran
The developers who expected most from AI, and the tasks they expected it to help with most, were the ones dropping out. So METR thinks its estimate is probably a lower bound on the true effect. From conversations with participants, it believes developers are likely more sped up in early 2026 than in early 2025, but says its data is only very weak evidence of how much.3 It now marks its 2025 results as out of date, and is changing the design of its study.
So a careful attempt to measure this had to change its design, because wider use of AI changed who would take part and how developers work. The 19% is a snapshot of early-2025 tools in one demanding setting, not a verdict, and the late-2025 numbers aren’t one either.
Why the team’s system matters more than the tool
DORA, the research programme on software delivery that Google Cloud runs, sums up its 2025 findings in one line:1
AI doesn’t fix a team; it amplifies what’s already there.
In DORA’s data, unlike in 2024, more AI use now goes with higher software delivery throughput and better product performance. It still goes with lower delivery stability. DORA’s explanation is that AI speeds up development, and that the extra speed can expose weaknesses further down the line. Without strong automated testing, mature version control and fast feedback, more change leads to instability. Teams with loosely coupled systems and fast feedback see gains. Teams held back by tightly coupled systems and slow processes see little or no benefit.1
These are associations in survey data, not controlled measurements. But the logic is straightforward. A task finished faster only reaches users sooner if review, testing and release can take the extra work. If they can’t, the time saved turns into a longer review queue, or into fixes after release.
DORA also finds a direct correlation between a high-quality internal platform and an organisation’s ability to get value from AI, one reason to be deliberate about what a platform team builds first.1 And if the constraint is a tightly coupled system, a better assistant won’t remove it. Loosening the system will, and that can be done piece by piece rather than in one rewrite.
Rolling out an assistant is the easy part. The hard part is making sure the rest of the path to production can use what it produces.
How to measure it in your team
Start from the outcome you want, not from the tool. If that’s getting changes to users sooner without more of them going wrong, DORA’s five software delivery metrics measure it. The first three measure throughput, and the last two instability.4
| Metric | What it measures |
|---|---|
| Change lead time | Time from a code commit to its successful deployment in production |
| Deployment frequency | How often application changes are deployed |
| Failed deployment recovery time | Time to recover from a failed deployment |
| Change fail rate | The share of deployments that cause a failure in production |
| Deployment rework rate | The share of deployments that are unplanned work to fix bugs |
Then:
- Take a baseline. Record the metrics before anything changes, or now if the tools are already in use, and keep the definitions fixed: what counts as a deployment, a failure and rework.
- Read throughput and instability together. DORA’s data links more AI use with higher throughput and lower stability. More deployments mean little if the rework rate rises with them.
- Watch where the wait moves. If code gets written faster, the queue can move to review and testing. Track how long changes wait for review, and how often they come back.
- Check quality, not only time. Some developers in METR’s later study said the quality of their work, and how much documentation and testing they did, differed with and without AI.3 Time on task misses that.
- Ask developers, and use the answers well. Surveys tell you where AI helps and where it gets in the way, which is how you learn what to point it at. Just don’t treat an estimate of time saved as the measurement.
- Give it an owner and a date. One person owns the measurement, and the result is reviewed on a set date.
In practice
Keep these metrics at team level. They describe how a delivery system performs. Used to rank individuals, they tend to change behaviour instead of measuring it.
The same discipline applies when AI is part of what you ship, not just how you build it: decide what good looks like before launch, and test against it.
Using AI coding tools isn’t the milestone anymore. Knowing what they’ve changed in your delivery is.
What to ask your team
The research so far doesn’t settle whether AI coding assistants make teams faster, and your team’s answer may differ from any study’s. Three answers show where you stand:
- what your delivery metrics were before the tools arrived, and how they’ve moved since
- whether throughput and rework moved together
- where the time saved went: into more changes, into review queues, or into fixing what shipped
Then the question that matters: since we adopted AI coding tools, are we shipping more with no more rework, and what’s our evidence that the tools are the reason?
Sources
- Google Cloud, Announcing the 2025 DORA Report: State of AI-assisted Software Development, 23 September 2025. A survey of nearly 5,000 technology professionals worldwide, with more than 100 hours of qualitative data. DORA is a research programme run by Google Cloud, which sells AI tools. The full report requires registration; the figures here are from the announcement.
- METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, 10 July 2025. A randomised controlled trial: 16 experienced open-source developers, 246 tasks, February to June 2025. METR, a non-profit funded by donations, now marks these results as out of date.
- METR, We are Changing our Developer Productivity Experiment Design, 24 February 2026. 57 developers, 143 repositories and more than 800 tasks, from August 2025. METR considers these estimates unreliable because of selection effects.
- DORA, A history of DORA’s software delivery metrics, 2 January 2026.


