The enterprise conversation about AI productivity often collapses into two positions: transformation is inevitable, or the gains are hype. The evidence supports a more useful conclusion. AI can produce meaningful improvements, but the size and even the direction of the effect depend on the task, the worker and the operating system around the tool.
That is good news for leaders. It turns productivity from an article of faith into a design problem.
Large gains can be real
In a field study of 5,179 customer-support agents, the authors of Generative AI at Work found that access to an AI assistant increased resolved issues per hour by 14% on average. The gains were much larger—34%—for novice and lower-skilled workers, with little measured improvement for the most experienced group.
The study took place in one support setting, so the percentages should not be projected across an enterprise. Its deeper finding is more portable. The system appeared to diffuse patterns associated with experienced workers, helping newer colleagues reach competence faster.
That suggests a powerful application for enterprise AI: making proven organisational knowledge available in the flow of work. The value is not only faster generation. It is the transfer of context, procedure and judgement.
Time saved does not always change the job
Another field experiment, Shifting Work Patterns with Generative AI, covered 7,137 knowledge workers across 66 firms. Among active users, access to generative AI was associated with roughly two fewer hours spent on email each week and less work outside normal hours. The researchers did not detect significant changes in the overall quantity or composition of tasks from individual-level access alone.
The result illustrates the difference between local efficiency and organisational redesign. A person may recover time without the surrounding process changing. To convert that time into a business outcome, teams need to decide what work should expand, disappear or move.
AI can also slow experts down
Counter-evidence matters. In a randomised study involving 16 experienced open-source developers and 246 tasks, METR found that early-2025 AI tools increased completion time by 19%. Participants had expected the tools to make them faster.
The sample was small, the developers were highly familiar with their codebases and the tools have since changed. It would be wrong to conclude that AI slows software development in general. It would be equally wrong to ignore the result because it conflicts with the prevailing narrative.
Experts may spend time steering the system, checking plausible errors or adapting generated work to complex local constraints. In some tasks, doing the work directly is faster.
Averages hide the deployment decision
An enterprise does not deploy an average. It deploys a specific system into a specific workflow for a varied group of people.
Measure results by task type, experience level, risk and source quality. Include the time spent prompting, reviewing, correcting and recovering from errors. Track quality and error severity alongside speed.
The right question is not, “How much productivity does AI create?” It is, “For which work, under which conditions, for whom, and with what downstream effects?”
Design for learning, not licence utilisation
Start with workflows where the system has access to relevant evidence and where outputs can be evaluated. Establish a baseline, introduce the tool to a defined group and compare complete outcomes. Interview users, but do not substitute confidence or satisfaction for measured performance.
Where novices benefit more, treat AI as part of capability development. Where experts see little gain, investigate whether the system lacks context or whether the task simply does not need assistance. Where time is saved, redesign capacity deliberately instead of assuming it will automatically become value.
Productivity is a portfolio
The enterprise opportunity is likely to be a portfolio of effects: dramatic improvement in some workflows, modest convenience in many and negative value in a few. Mature organisations will stop looking for one universal number and build the capability to identify each category quickly.
That is a more demanding approach than announcing adoption. It is also how AI becomes durable: through measured improvements in real work, with enough honesty to stop or redesign the applications that do not deliver.
