This post looks at both sides of the evidence on AI in software delivery: what it is measurably delivering, what it is measurably costing, and why most organisations end up closer to the second. Every figure comes from a named 2025 or 2026 source, linked where it appears.
What AI is measurably delivering

But let’s start with Amazon CEO Andy Jassy’s 2025 Letter to Shareholders where he shared a story that should give every technology executive something to consider:
A Signal Too Big to Ignore
“Six engineers rebuilt the entire Amazon Bedrock inference engine in 76 days using Kiro, Amazon’s agentic coding service. The original estimate was 40 engineers and a full year. This new engine, called Mantle, became the backbone of Amazon Bedrock’s rapid scaling. That compression ratio, from 40 person-years to 6 people in 76 days, is not an incremental improvement. It is a category shift in how software gets built.”
This is not a proof-of-concept. It is production infrastructure powering one of the world’s largest AI services.
The wider data points the same way. DX, surveying more than 500 organisations, found that AI wrote 52.7% of all code in the second quarter of 2026, up from 24% two quarters earlier. Median throughput rose from 1.42 to 1.94 pull requests per engineer per week.
Adoption is close to universal. In DORA's 2025 report, drawn from around 5,000 respondents, adoption stands at 90%, and more than 80% of respondents report efficiency gains.
Why do most AI projects stall before production?

The gains at the keyboard rarely reach the business. MIT NANDA's 2025 study of 150 enterprise leaders found that 95% of generative AI pilots produced no measurable profit impact. Gartner expects more than 40% of agentic AI projects to be cancelled by the end of 2027.
In our engagements, the same four causes come up again and again. Legacy codebases are the first: policy constraints and ten-year-old systems stop AI at the pilot stage. The second is that adoption stops at the prompt. Individuals use AI, and the delivery process around them stays as it was. Third, training has not kept pace with the tools. Fourth, nothing is auditable, because AI-written code arrives with no trail of who approved what.
What is AI costing teams that move fast without a process?

Speed without a process shows up later as risk. Veracode tested more than 100 models in 2026 and found that 44% of AI code-generation tasks introduce a risky security vulnerability. GitClear's analysis of 623 million code changes shows duplication up 81% since 2023 and refactoring down 70%, the pattern of code being added faster than it is maintained.
The effects reach production. Faros AI's 2026 report, covering 22,000 developers, found that incidents per pull request roughly tripled and that pull requests merged with no review rose 31.3%. DX recorded change confidence falling by 6.1% over the same period.
Why don't the productivity gains add up?
The clearest sign is in the spending. DX found that AI spend rose roughly 28 times over four quarters, while the share of engineering effort going to new product work stayed flat at 57–58%. Teams are writing more code, and the same share of their time still goes to new features.
Perception makes the gap harder to see. In METR's 2025 randomised trial, experienced developers using AI took 19% longer to finish their tasks, and believed they had been 20% faster. A team that judges AI by how it feels can miss what it is costing.
What do the teams that get results do differently?
What stands out in the Bedrock story is the shape of the team: a small, separate group, starting over, building on an agentic coding service from the first day. The tool mattered, and so did the way the work was organised around it.
We build delivery on the same idea, with four rules. A person approves every stage. AI goes only on the steps where it adds value. Every model call is costed in a per-call ledger from day one. And some places stay free of AI by design.
That is the G.U.I.D.E. method we deliver with. You can see it applied in our engineering work and in the agentic accelerators we build for clients.
Sources
| Source | Figures | Date |
|---|---|---|
| Amazon, 2025 Letter to Shareholders | 6 engineers, 76 days, against an estimate of 40 engineers and a year | 2026 |
| DX, State of AI Impact in Engineering Q2 2026 | 52.7%, 57–58%, 28×, 1.42 to 1.94, −6.1% | Q2 2026 |
| DORA, 2025 State of AI-assisted Software Development | 90%, 80%+ | 2025 |
| MIT NANDA, The GenAI Divide, 2025 preprint | 95% | 2025 |
| Gartner, June 2025 | over 40% | 2025 |
| Veracode, 2026 GenAI Code Security Report | 44% | 2026 |
| GitClear, The Maintainability Gap, 2026 | +81%, −70% | 2026 |
| Faros AI, The AI Engineering Report 2026 | incidents 3×, +31.3% | 2026 |
| METR, 2025 randomised trial | 19% slower, believed 20% faster | 2025 |
Related on DeployProAI
About the author

Ashish Tripathi is the founder of DeployProAI. He has spent more than twenty years building enterprise-scale systems, from wealth platforms at Morgan Stanley to technology strategy for a $300M+ ARR portfolio at Amazon Web Services. He is 4X AWS certified, and has mentored more than 200 engineers across 100+ enterprises and software companies. Before founding DeployProAI he delivered AI-native software development cohorts for AWS enterprise customers, including Deloitte, Axis Bank, HDFC Bank and Care Health Insurance.