Crafting Excellence in Software
Let’s build something extraordinary together.
Rely on Lasting Dynamics for unparalleled software quality.
Mohamed Boukrim
Sep 10, 2026 • 17 min read

The AI productivity paradox is not that the tools fail. It is that the most careful experiment in the fieldwas retired before it could give anybody a second number, because the developers in it would not work without the tool long enough to be measured. Effort did not disappear when generation got cheap. It moved into specification, review, integration and rework, and three of those four are invisible to the person paying for them.
Every METR and DORA figure below was read at the primary source on 9 September 2026.
The AI productivity paradox has a headline number, and the number is no longer the story. In July 2025 METR published a randomized controlled trial in which experienced open-source developers, working on repositories they already knew, took 19 percent longer with AI tools than without. That figure travelled for a year. In February 2026 the same team published something more useful and almost entirely unquoted: they are changing the experiment design, because they can no longer run it.
The follow-up ran from August 2025. Fifty-seven developers, 143 repositories, more than 800 tasks. The returning developers came out 18 percent faster, the newly recruited 4 percent. Both intervals cross zero, so neither is a finding. The finding, and the reason the AI productivity paradox is worth an article at all: between 30 and 50 percent of participants told METR they had stopped submitting the tasks they did not want to do without AI.
That mechanism inverts the usual reading of the AI productivity paradox. Developers being paid to work on their own tickets declined the ones where giving up the tool would have hurt most. METR's conclusion is not that the effect is small. It is that their estimate is a lower bound, and that the real speedup is probably larger among exactly the developers and tasks which selected themselves out.
So the honest version of the AI productivity paradox is not that AI makes engineers slower. It is that completion time stopped being measurable once the tool became load-bearing, and completion time was the only number anybody was tracking. Effort did not disappear. It moved into work nobody counts, and the rest of this article is about where.
Every number in the AI productivity paradox comes from a study measuring a different population, and none of them measured the cost that lands later.
METR's early-2025 result: the use of AI caused tasks to take 19 percent longer, with a confidence interval from plus 2 to plus 39 percent. Sixteen experienced contributors, their own repositories, early-2025 tooling - the most careful design anybody had run, on conditions narrow enough that generalizing from it was always a stretch.
METR's February 2026 update reported an 18 percent speedup for the returning subset, interval from minus 38 to plus 9 percent, and 4 percent for the newly recruited, interval from minus 15 to plus 9 percent. All intervals here are on task completion time, so a negative number means faster. The team is explicit that this is only very weak evidence for the size of the change. They do not withdraw the original result. They retire the method.
Why did METR stop running the experiment?
Because task-level randomization stopped working. Recruitment fell as developers refused to do half their work without AI, and 30 to 50 percent of those who stayed declined to submit the tasks where the tool helped most. The design systematically loses the highest-uplift work, so its central estimate is a lower bound rather than a measurement.
The AI productivity paradox also has a time-horizon shape, and DORA is the software-specific source for it. The ROI of AI-assisted Software Development report describes an initial dip before the gain, calls that dip the tuition cost of transformation, and states that positive return is not guaranteed. Its authors attach the caveat that outranks every figure in the document: treat these calculations as a high-uncertainty estimate meant to spark a conversation, rather than a rigid mathematical formula.
We are not repeating that report's modelled dollar return here. It illustrates a hypothetical 500-person organization, its own authors say to treat it as a conversation starter, and the coverage of it puts a return, an investment and a percentage side by side, and the three do not reconcile. In an article about numbers that fail under a spreadsheet, we are not going to publish one that does the same thing.
Let’s build something extraordinary together.
Rely on Lasting Dynamics for unparalleled software quality.
The citation trail supplies a sharper illustration of the AI productivity paradox than any figure inside it. One widely repeated claim holds that AI delivers a 35 to 40 percent gain on simple greenfield tasks and 10 percent or less on complex legacy code, attributed to Stanford.
Follow it back: press coverage, citing the DORA report, citing Stanford's programme. Stanford's own AI Impact page publishes its method and its scale, 600 or more organizations and 120,000 or more engineers since 2022 - but no effect sizes at all. The figure may be real. It is not checkable, and half the commentary on the AI productivity paradox rests on it.
Adoption surveys are the weakest evidence in the set, near-universal uptake against a self-reported gain of around 10 percent, and they get stacked beside randomized trials as though the two were comparable. Nothing above measured the maintenance cost of generated code six months later either, which is the line item a supplier actually carries.
The range therefore runs from a 19 percent slowdown to a large speedup, and every serious paper in it is measuring a different population under different tooling. A CTO who wants a number for their own team will not find it in the research. That is the working content of the AI productivity paradox, and it is why everything below is about where the effort goes rather than how much of it there is.
Generation does not remove work from software engineering. It moves work downstream into four places.

The asymmetry is what turns a transfer into an apparent saving, and it is the whole mechanism of the AI productivity paradox. Typing is visible and it is felt. Reading is invisible and it is not. Three of those four destinations are reading costs, so the ledger looks like a win at the moment of the transaction and settles later, against a different budget, often in a different quarter.
This is also why the studies disagree rather than converge. A completion-time measurement catches specification and review, because both happen inside the task window. It structurally cannot catch integration and rework, because both happen after the timer stops. Every headline figure is a measurement of the first two destinations only.
So an engineering organization that has not moved review capacity has not adopted AI, it has deferred a bill. The capacity is the adoption. The tooling is just the tooling.
One quantity sets the price of that transfer in both directions: the gap between what you can write and what you can read. Hold that gap in view and the AI productivity paradox becomes a single sentence rather than a pile of anecdotes.
From idea to launch, we craft scalable software tailored to your business needs.
Partner with us to accelerate your growth.
Start with the expert on familiar ground, because that is where the AI productivity paradox was first measured. An engineer who knows the codebase already holds, unwritten, the context a model has to be told. The specification cost is therefore high relative to the writing cost it replaces, and the review is fast but unforgiving, because they can see every place the output diverges from how this system does things.
It is tempting to read the cohort split as proof of that mechanism, and it does not support it. The returning developers came out faster than the new recruits, 18 percent against 4, which is the opposite of what a pure expertise story predicts. It is also not readable as a cohort comparison. The returning subset is ten people - the ones who agreed to come back after a study which found they had been slowed down. Anybody citing that split is citing a self-selected sample of ten.
Does AI help experienced developers less?
On the current evidence nobody can say. The mechanism is plausible, because an expert already holds the context a model must be told, so specification costs more and saves less. But the only trial that tested it has been retired for selection bias and its own follow-up points the other way. Treat it as a reason to measure your team, not as a finding.
Now the other end, where the AI productivity paradox reverses. For a novice, or for anyone working in an unfamiliar language, the write-ability gap is enormous and the apparent gain is correspondingly huge. The problem is that the gain is largest exactly where the ability to evaluate the output is smallest. You cannot review what you cannot read, so what looks like acceleration is partly just acceptance.
That is the reviewability ceiling, and it is the real limit on AI-assisted delivery rather than model capability. Above the line of what the operator can actually evaluate, output is being accepted rather than reviewed. Improving the model does not move that line by a millimeter, because it is a property of the person. It is also the part of this that no dashboard will ever show you.
Two consequences are uncomfortable to write down. Expertise has historically been built by writing the code the model now writes, and nobody has a settled answer for how it gets built instead. And the counter-argument a good reader raises immediately is correct: an expert on an unfamiliar codebase is closer to the novice case, which is where the tool genuinely pays for a senior engineer.
So the rule of thumb: the gain is proportional to the distance between what you could not have written yourself and what you can still reliably judge once it is written. Which produces the inversion at the center of the AI productivity paradox. The tool pays best for the person who needs it least, and it is sold to everybody as though the opposite were true.
One ratio predicts both lists below, and it is the only practical tool the AI productivity paradox leaves you with: the cost of checking the answer against the cost of producing it. Task category is the proxy people reach for, because the ratio is harder to see than the label. Where AI-assisted effort pays:
Two of those we can defend from our own delivery record rather than from the literature. Specification-backed test generation and mechanical migration are where AI-assisted work has reliably paid on client engagements, for the same reason in both cases: somebody else already wrote down what correct means, so the AI productivity paradox never gets a chance to start.
Where it costs more than it saves, the ratio is simply inverted:
We design and build high-quality digital products that stand out.
Reliability, performance, and innovation at every step.
The loss we have actually paid for is the last one, and it is the AI productivity paradox in miniature. A small, well-understood change in a system somebody knew cold, specified and reviewed instead of typed, is slower and produces nothing the team did not already have. It reads as diligence and it costs throughput.
The common thread, and the most useful sentence in the whole AI productivity paradox: every win is a case where checking the answer is cheaper than producing it, and every loss is a case where it is not. That ratio, not the task type, is the predictor. Where generation should be switched off entirely is a governance question rather than an economic one, and it belongs in its own article.
Model-side effort is now a budgeted resource with a price, and most teams leave it at the default in both directions. This is the half of the AI productivity paradox that has an actual dial attached to it. The controls are request parameters rather than properties of the model, which makes them configuration you own and can get wrong.
Anthropic exposes effort with the levels low, medium, high, xhigh and max, defaults to high, and states the point that decides your cost model: effort is a behavioral signal, not a strict token budget. OpenAI exposes reasoning.effort and documents that both the accepted values and the default are model-dependent rather than universal. Google exposes thinking_level and documents that its models reason dynamically by default. Read all three for the model you are actually calling, because none of these vocabularies translate.
The consequence people get wrong: a high setting on a simple request can consume roughly what a low setting would, because the model declines the permission it was given. Token projections built on the assumption that level equals consumption are wrong in both directions. If you need a hard ceiling, the control for that is the maximum output tokens, and it is a different instrument.
The decision rule worth arguing for is the same ratio as the section above, applied to the other side of the exchange. Raise model effort when verification is expensive, lower it when checking the answer is cheap. Not when the task looks hard, which is the input everybody uses and the one input the METR result gives you least reason to trust. High effort is theatre on work whose failure mode is obvious on the first read, and low effort is a false economy on anything you will merge without a second pair of eyes.
We have written up separetely how we set reasoning effort in practice and how we isolate two Claude Code environments on one machine, and neither belongs in a summary here. What belongs here is the commercial half of the AI productivity paradox: per-token spend is an operating expense that lands in a client's run rate, so it belongs in the estimate rather than in a footnote.
The highest-leverage effort available to an operator is supplying the right context, not phrasing the request better. Most prompt engineering advice is a workaround for context the model was never given, which is why context engineering outperforms it and why the two are worth keeping apart when you are trying to spend your way out of the AI productivity paradox.
Go back to the expert case. Their specification cost is high because the context lives in their head. Writing it down once is capital, and it earns for as long as the file survives. Re-prompting is an expense, and it earns for exactly one turn. Context engineering is the name for choosing the first one on purpose, and it is the cheapest move available.
Chain-of-thought prompting and its descendants are now largely internal to reasoning models, which is a more useful thing to say than repeating the technique. If you learned to coax step-by-step reasoning out of a model in 2023, that skill has partly been absorbed into the product you are calling.
What actually moves the outcome is unglamorous and specific:
The durable-artifact test is the whole discipline in one line: if you are typing the same context for the third time, it belongs in a file the model reads and not in a prompt. Context engineering is the one place where user effort reduces total effort rather than relocating it, which makes it the only unambiguous win on this page. We keep a conventions file in every client repository for exactly this reason.
We treat AI-assisted effort as a budget with named line items rather than as a discount, and that single choice is what makes the AI productivity paradox survivable commercially. The line items are specification, generation, review, integration, and a reserve for rework. Five numbers, not one percentage.
Review capacity is planned before generation capacity, because review is the binding constraint. If review is the bottleneck, adding generation adds queue rather than throughput, which is how the AI productivity paradox arrives inside a delivery plan. DORA reaches the same conclusion from the opposite direction when it says the returns come from clearing bottlenecks rather than from the code the tool writes.
That rule costs us velocity, and it is worth saying so rather than presenting it as a virtue. Capping generation at what the available reviewer can actually read means work sits waiting for a person on days when the model could have produced more. It is a real cost, we pay it deliberately, and it is the reason the dates hold.
The budget is written down and auditable, which is a process requirement rather than a marketing one. We are certified to ISO 9001 for our development and process management, and the practical effect is that a number in an estimate has to be traceable to something other than a feeling.
What we do not claim: we do not publish an internal speedup figure, because we have not measured one to a standard we would defend. Given everything above about the AI productivity paradox, a supplier quoting a precise percentage for its own AI-assisted productivity is telling you about its marketing rather than its measurement.
The AI productivity paradox stops being an interesting research question the moment somebody signs for a date. Bidding the felt speedup rather than the measured one is how an AI-assisted project loses money without anyone making a single bad engineering decision.
Put the range from the evidence section to work. If the literature spans a 19 percent slowdown to a large speedup, and the best-designed trial in it has been retired for selection bias, then any single number in a fixed-scope bid is a bet dressed as an estimate. We answer that by bidding the process instead: the five line items, the review capacity, and the reserve.

The constraint that shapes the work most is set before the first line of code, and no tool choice relaxes it. On client engagements we deliver inside their own tentant, their managed settings and their NDA, so identity and isolation are configuration committed to the repository rather than a habit. That is the least impressive-sounding thing on this page and the most load-bearing, and it bounds every answer a supplier can honestly give.
Then there is the review floor. In a regulated environment the review effort that the AI productivity paradox relocates onto the engineer is not discretionary, because it produces evidence an external assessor will ask for on a schedule. That argument is carried in full in our piece on what a PCI DSS assessment actually asks for. The consequence for this one is short: the part of an estimate AI cannot compress is the part somebody else has to be able to verify.
DORA, Regulation 2022/2554, and the EU AI Act, Regulation 2024/1689, attach obligations to the process rather than to the output. A faster process with weaker evidence is therefore a net loss under both.
The transferable claim: the discipline of proving a control to a hostile external assessor is the same discipline that makes an AI-assisted estimate honest. Both require that a number be traceable to something. The client is not buying our speedup. They are buying a date we can hold.
The metrics that vendors supply measure adoption rather than benefit, which is why so many teams have a dashboard and no answer to the AI productivity paradox. Suggestion acceptance rate, lines generated and percentage of code AI-authored all rise when the tool is being used badly. They are usage counters in the costume of outcomes.
Self-report is the specific thing the evidence discredits, and the AI productivity paradox now supplies a sharper version of that point than any survey could. METR could not hold a controlled experiment together because participants would not give the tool up. If a research team paying by the hour cannot get clean self-denial data, an internal engagement survey has no chance.
What is worth measuring against the AI productivity paradox is slower and less flattering:
The DORA delivery metrics are the pre-existing frame here, and they answer the question better than any purpose-built AI dashboard because they were not designed to flatter the tool. The current model has five: change lead time, deployment frequency, failed deployment recovery time, change fail rate, and deployment rework rate. That last one is worth the whole read. Rework is the fourth destination of relocated effort, and DORA now measures it by name.
How long before AI adoption shows up in delivery metrics?
Longer than a quarter, and the first movement is usually in the wrong direction. DORA describes an initial dip before the gain and calls it the tuition cost of transformation. A team that reads that dip as failure and pulls the funding during it has bought the cost and cancelled the return.
You will not run a randomized trial, so state the confounders instead of pretending they are absent. The tooling changed, the team changed, the work changed, and the model changed twice while you were measuring. Which leads to the conclusion most engineering leaders resist, and the one the AI productivity paradox should leave you with: if you cannot tell whether it helped, the correct claim is that you do not know. Not that it helped.
Asking a supplier whether they use AI no longer discriminates between suppliers, because the answer is yes everywhere. These four questions do, because each one asks how they handle the AI productivity paradox rather than whether they have heard of it. None can be answered with marketing language.
The fourth is the one to ask first. It is the only question here where a comprehensive answer is the good sign and a reassuring one is not.
We do AI software development and bespoke software development for clients who mostly arrive holding an estimate they no longer trust. We are happy to be measured against the four questions above.
Trying to work out whether AI is moving your delivery dates, or just moving the effort? Tell us what you are trying to estimate ->
This article was written by Mohamed Boukrim at Lasting Dynamics, an EU-based custom software development company. We deliver agent-assisted work on fixed-scope engagements, which is why the gap between the speedup we feel and the speedup we get is a number we have to defend rather than estimate.
The AI productivity paradox is the gap between near-universal adoption of AI coding tools and measured gains that are small, inconsistent and sometimes negative. The resolution is relocation rather than illusion. Effort moved out of writing code and into specifying, reviewing, integrating and reworking it, and three of those four are reading costs that nobody was counting.
Sometimes, and the best evidence no longer says how often. METR's 19 percent slowdown was followed, in a differently recruited run, by an 18 percent speedup for returning developers and 4 percent for new recruits, every interval crossing zero, and the design was then retired for selection bias. The condition that predicts a slowdown is an expert making a small, precise change in a codebase they already know.
Not with acceptance rate, lines generated or percentage of code AI-authored, all of which rise when the tool is used badly. Measure time from merge to first defect, review turnaround, deployment rework rate on AI-assisted changes against everything else, and change fail rate. DORA's five delivery metrics already cover most of it, including deployment rework rate.
Raise it when verification is expensive and lower it when checking the answer is cheap. Do not set it from how hard the task looks, which is the input people judge worst. And do not model cost as level times tokens, because vendors document effort as a behavioral signal rather than a budget, so a high setting on an easy request may cost what a low one would.
No, and the evidence does not support that reading. What it supports is moving review capacity before you move generation capacity, and switching the tool off for the narrow case it reliably loses: small, precise changes inside a system the engineer already knows cold. Everything else is a question of whether checking the answer costs less than producing it.
Transform bold ideas into powerful applications.
Let’s create software that makes an impact together.
Mohamed Boukrim
I am a Software Engineer and Backend Developer with a relentless focus on software quality and robust architecture. I don't believe in shortcuts; outstanding results are the direct product of daily commitment and hard work. Alongside my team, I leverage Agile methodologies and continuous process evaluation to push boundaries and optimize performance. I approach every challenge with a highly competitive mindset: I work tirelessly to reach the top, and once there, I keep pushing to maintain the lead.