hirevolution

Where AI actually saves a small team time today (and where it wastes it)

AI is in every headline, every product update, and most sales pitches, and the claims are large. On a small team the useful question is narrower than any of that. It is not whether AI is powerful. It is which of this week’s hours it actually gives back, and which it quietly takes. The gap between the demo that saves forty per cent and the working week that feels much the same is where a lot of small businesses get stuck. The honest answer is that AI does save real time today, but only on certain kinds of work, and on the wrong kinds it can cost more than it returns. Knowing the difference is most of the value.

Where it genuinely saves time today

The clearest wins are well evidenced, and they cluster around drafting, summarising, and triage. In a randomised trial of 453 professionals, giving people a chatbot for occupation-specific writing tasks such as press releases, short reports, and sensitive emails cut the average time taken by 40% and raised the quality of the output by 18% 3. That is the strongest single piece of evidence for the most common use: a first draft a person then edits, rather than a finished article the machine hands over.

The same pattern shows up in customer enquiries. When a gen-AI assistant was rolled out to 5,179 customer-support agents, it lifted the number of issues resolved per hour by 14% on average 4. And contained, well-specified coding has it too: in a trial where 95 developers were asked to write a standard web server, the group given an AI coding assistant finished the task 55.8% faster 5.

There is a thread running through all three worth pulling out, because it matters most to a small team with mixed experience. The gains land hardest on the least-experienced people. In the support study, the 14% average hid a 34% improvement for novice and lower-skilled staff and very little change for the most experienced 4. In the writing trial, weaker-skilled participants benefited most, and the gap between workers narrowed 3. For an owner, that points somewhere practical: AI is often most useful handed to the newer member of the team on the routine drafting and triage, not reserved for the person who already does that work fast.

Why the whole-week saving is smaller than the demo suggests

A 40%-faster task is not a 40%-shorter week, and this is where most of the disappointment comes from. When researchers measured actual hours saved across real jobs rather than one ideal task, the numbers are modest. A large Danish study of around 25,000 workers across some 7,000 workplaces found users reported average time savings of just 2.8% of their work hours 1. The OECD’s survey, which includes the United Kingdom and covers over 5,000 SMEs, puts the same figure between 2.8% and 5.4% depending on the country, and notes plainly that this may be too small to change how many staff a business needs 2.

The reason is not that the trials were wrong. It is that most of a real job is not the one task AI is good at. The OECD survey found that of the SMEs using gen-AI, 91.6% use it to generate text, but only 28.7% use it in their core, revenue-producing work 2. In practice it gets pointed at simple, one-off, peripheral tasks rather than the complex, recurring work that actually fills the week. The Danish study adds a detail that is easy to miss: AI also created new tasks for 8.4% of workers, so it does not only remove work, it adds some 1.

The honest read is that the gains are real but bounded. They land on the slice of the week that fits, which is genuinely useful, and they do not scale up to the whole business the way the big task-level percentages imply. Expecting the demo number to apply to everything is the quickest way to feel let down by a tool that is, in fact, helping.

Where it costs more than it saves

The more important finding for a busy owner is that AI does not just fail to help on the wrong tasks. It can actively make the work worse, and it does so most when people trust it most. In a field experiment with 758 management consultants, AI was a clear help on tasks inside its competence. On tasks just outside it, the consultants using AI were about 19 percentage points less likely to reach the correct answer than those working without it 6. The researchers called this the jagged frontier: the line between what AI does well and what it does badly is uneven and hard to see from the outside, and people over-rely right where it is weakest.

The trap is the feeling. In a 2025 study, 16 experienced developers worked on 246 real tasks in large codebases they knew well. When allowed to use AI, they took 19% longer to finish. The striking part is that they expected AI to speed them up by 24%, and even after living through the slowdown they still believed it had sped them up by 20% 7. People felt faster while being slower. That is worth holding onto, because it means a team’s own sense that AI is saving time is not reliable evidence that it is.

Two caveats keep this honest. That developer study deliberately looked at experts working in their own large, complex projects, and the authors are clear it does not represent all software work 7. It does not contradict the earlier finding that a contained, well-specified coding task got much faster 5. The two together are the jagged frontier made concrete: a small, common, well-defined job speeds up, while expert work in a big unfamiliar context slows down. The lesson generalises beyond code. High-stakes, expert, or context-heavy work is exactly where the cost of checking the output eats the time the draft saved, and where a confident wrong answer is worse than no answer at all.

A simple test for handing a task over

Putting the evidence together gives a short test that works for most tasks an owner is weighing up. Three questions, in order.

First, is the task mostly text, and is it well specified? Drafting an email, summarising a long thread, turning notes into a first proposal, sorting incoming enquiries by what they need. These are the jobs the trials show real gains on 3.

Second, is it low-stakes, or a first draft that a person checks before anything ships? The saving is real when a human stays in the loop and the cost of a mistake is small. It evaporates when nobody verifies the output and a wrong answer goes out the door 6.

Third, does it sit on common, well-documented ground rather than your specialist niche or a large, unfamiliar context? AI is strong on the ordinary and weak on the particular, and it is weakest precisely where it sounds most confident 7.

Three yeses and it is a good candidate: hand it over, then check it. Any no and the task is better kept human, or automated a different way that does not depend on the model getting your specifics right.

A simple test for handing a task to AI: text-heavy and well-specified, low-stakes or a first draft a human checks, and on common ground rather than your specialist niche. Three yeses and it is a good candidate; any no and it stays human or gets automated another way. Synthesis of the jagged-frontier and developer-productivity studies and the OECD SME survey.

A simple test for handing a task to AI: text-heavy and well-specified, low-stakes or a first draft a human checks, and on common ground rather than your specialist niche. Three yeses and it is a good candidate; any no and it stays human or gets automated another way. Synthesis of the jagged-frontier and developer-productivity studies and the OECD SME survey.

One more factor decides how much a team actually gets from all this, and it is not the tool. The Danish study found the benefits, in time saved, quality, and satisfaction, were 10% to 40% greater where employers actively encouraged and supported people in using AI 1. In other words, “we tried AI and it did not help” is often a sign of how it was adopted rather than proof the tool was useless. Getting value is partly about picking the right tasks, and partly about building the small habits and workflows around them so the saving actually shows up.

Worth a closer look

The gains from AI are real, but they come from choosing tasks deliberately and building the small workflows around them, which is the unglamorous part most coverage skips. That is the part hirevolution helps a small team get right.

See how hirevolution helps a small team get real value from AI

Sources

  1. Anders Humlum & Emilie Vestergaard, ‘Large Language Models, Small Labor Market Effects’ (working paper), 15 April 2025; ~25,000 workers across ~7,000 Danish workplaces. [link]
  2. OECD, ‘Generative AI and the SME Workforce: New Survey Evidence’, November 2025; survey of over 5,000 SMEs across seven countries including the United Kingdom, fielded end-2024. [link]
  3. Shakked Noy & Whitney Zhang, ‘Experimental evidence on the productivity effects of generative artificial intelligence’, Science, 13 July 2023; randomised trial of 453 professionals. [link]
  4. Erik Brynjolfsson, Danielle Li & Lindsey R. Raymond, ‘Generative AI at Work’, NBER Working Paper 31161, 2023; staggered rollout to 5,179 customer-support agents. [link]
  5. Sida Peng, Eirini Kalliamvakou, Peter Cihon & Mert Demirer, ‘The Impact of AI on Developer Productivity: Evidence from GitHub Copilot’, arXiv:2302.06590, 2023; randomised trial of 95 developers on a self-contained task. [link]
  6. Dell’Acqua et al., ‘Navigating the Jagged Technological Frontier’, Harvard Business School / BCG working paper, 2023; field experiment with 758 consultants. [link]
  7. METR, ‘Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity’, 10 July 2025; 16 experienced developers, 246 real tasks in large codebases they maintain. [link]