Article
AI Is Here. Where Are the Productivity Gains?
A CEO pulled me aside after a session. He had bought the licenses, his team was using AI, and he wanted to know where the payoff was. "I am not seeing it in the numbers yet."
In 1987, economist Robert Solow wrote: "You can see the computer age everywhere but in the productivity statistics." The line appeared in his review, We'd Better Watch Out, in The New York Times Book Review.
If you've bought AI tools and you're still waiting for the payoff, you can understand the frustration. You should expect an answer that goes beyond how many people have logged in or how impressive the latest demonstration looked.
Start with a recurring workflow and measure it from the request through to a finished result somebody can use. Include the time spent preparing inputs, checking the answer and fixing mistakes. Then decide what the business will do with any capacity released. That gives you something you can manage while the broader argument about AI and productivity continues.
Keep investing in useful experiments, and put a decision date on each one. Ask the team to show what changed, what it cost and what they recommend next. That gives you a reason to expand, change course or stop.
What the evidence actually shows
When you read an AI productivity study, ask how closely the work resembles yours. Pay attention to who used the tool, what they did and how the researchers judged the result.
In Generative AI at Work, published in 2025, Erik Brynjolfsson, Danielle Li and Lindsey Raymond studied the staggered introduction of an assistant to 5,172 customer-support agents at a single software company and found an average 15% increase in issues resolved per hour, with larger benefits for less experienced and lower-skilled agents.
The assistant suggested responses during customer conversations, and the agents could edit or ignore them. Experienced, highly skilled agents gained little speed and showed small declines in quality. Read the published study.
A different result came from METR's randomized study of 16 experienced open-source developers completing 246 tasks in familiar repositories: access to early-2025 AI tools made completion take 19% longer in that specific setting, despite participants believing the tools helped.
In February 2026, METR reported that its follow-up suggested improvement but that participant selection and time-measurement problems prevented a reliable estimate of the current effect. Read the original study account and the follow-up.
Test the work your team actually does. Ask enthusiastic users to show you a finished example and the time it took, including the checking. Use that evidence to choose where you'll invest next.
Why investment and results can arrive at different times
The gap between spending and results also depends on the work you do around the technology. In their 2021 paper on the productivity J-curve, Brynjolfsson, Daniel Rock and Chad Syverson describe how technologies such as AI require complementary investments, including intangible investments that are poorly captured in national accounts. Their model explains how measured productivity can initially understate the benefit of those investments. Read the paper's abstract and publication details.
Look at what you've funded beyond the subscriptions. Have you given people time to practice? Have you agreed on the inputs? Has somebody changed the handoff to the next person? Is anyone responsible for keeping the process useful?
Some uses deserve to be dropped. If checking every output costs more than the work saves, look for a better approach. A template or a change to your existing software may solve the problem more cheaply.
Set a time and spending limit for learning, and decide what you need to see at the end. Give people room to practice and change the process. Then hold the owner to the review date.
Decide which result you're trying to improve
Before you count hours, write down the business problem. A sales team may need proposals to reach customers sooner. A finance team may need more time to investigate exceptions before its monthly review. A service team may need to reduce the queue without increasing mistakes.
Those problems require different measures. Choose a main outcome and a quality condition that must hold for the improvement to count.
- If proposals arrive too late, track time from an agreed brief to a proposal ready to send. Check scope, pricing and customer-specific commitments before counting the result.
- If reporting leaves too little time for analysis, track preparation time and whether the agreed review happens earlier. Reconcile figures and check the explanation against source records.
- If service work is accumulating, track accepted resolutions and the age of the queue. Check repeat contacts, corrections and escalations.
- If senior staff spend too much time rewriting, track their total review and revision time. Check whether the final work meets the same standard.
Choose the example closest to your problem and make it specific enough that two people would count it the same way.
Keep labor time separate from elapsed time. A document may require little hands-on work and still spend days waiting for approval. If waiting is the problem, speeding up the draft alone won't achieve the outcome you chose. Look at who approves the work, what they need to see and how long it sits with them before you pay for more automation.
Find out what the work costs today
Ask the person who does the work to walk through several recent examples with you. Include an ordinary case and a difficult one. Record how the request arrived, what information was missing, who touched it, how much time each person spent and what caused revisions.
Do this before the pilot. Reconstructing the old process after everyone is excited about the new one invites a favorable comparison. Where historical time records are weak, label the baseline an estimate and collect a short period of direct observations before relying on it.
Start with a batch of comparable cases that your team can review properly. Include enough variation to expose the normal exceptions. Choose how much to test based on how often the work happens, how much it varies and what a mistake would cost.
Keep a simple record for each case:
- Record the case type and complexity so you can compare similar work.
- Record preparation, production, review and correction time separately.
- Record when the work was requested and when the next person accepted it.
- Record material errors, missing information and any need to repeat the work.
- Record which tool and process version were used.
If possible, alternate comparable cases between the existing method and the proposed method. Where that's impractical, record other changes that could explain the result, such as a quieter period, different staff or a simpler mix of requests. You want enough detail to tell whether the new method helped and where it still needs work.
Count the whole workflow
Here is a hypothetical example.
Suppose a team completes 20 proposals a month. Under its existing process, each proposal requires 90 minutes of drafting and 30 minutes of review. That's 40 hours of work a month.
In the proposed AI-assisted process, each proposal requires 15 minutes to prepare the inputs, 25 minutes to produce and revise the draft, and 35 minutes for review and correction. That's 25 hours a month, releasing 15 hours if the case mix and accepted quality remain comparable.
At a loaded labor rate of $80 an hour, those 15 hours represent $1,200 of capacity. If recurring software and support cost $300 a month, the capacity value after that cost is $900. If the same people stay on the same payroll, your payroll spending stays the same. You also need to recover the one-time setup and training costs.
Now decide what happens to the released time. Does a senior person spend more time developing qualified opportunities? Can the team handle additional work without increasing overtime? Does the organization deliver existing commitments more reliably? Name the use and check whether it occurs.
If there's no demand for extra proposals, don't assume that producing more of them creates value. If customers need a conversation before they can decide, the better use of capacity may be better preparation and follow-up. The operating goal should determine where the time goes.
Protect the quality that customers pay for
A faster draft counts only when somebody can use the finished work. Set the acceptance criteria before you see the AI output, and have a person who understands the work apply them.
For a proposal, I'd check whether it accurately describes the customer's problem, reflects the agreed scope, uses approved pricing and avoids unsupported promises. I'd also ask the next person in the process whether it made their job easier. A salesperson's faster draft may create more work for delivery if the commitments are unclear.
For a research brief, require the reviewer to open the sources that support material claims. For reporting, reconcile the numbers against the records. For customer communication, check whether the message answers the actual question and whether the proposed next step is authorized.
Record serious errors separately from minor editing. A typo and an incorrect contractual commitment should not disappear into the same average correction count. Decide in advance which failures pause the pilot, even if the average completion time improves.
Give the reviewer enough time to check the work and authority to reject it. Ask them to show you the errors they caught and how they decided the final version was ready. If reliable checking takes longer than doing the task directly, narrow the task or use another method.
Decide what changes on Monday
Once you've evidence that the workflow helps, document what will change on Monday. Who prepares the inputs? Where is the approved template? Who reviews the output? What happens when the information is incomplete? Which old step can actually be removed?
Avoid carrying both processes indefinitely without a reason. If the pilot succeeds but everybody must still recreate the work manually for approval, include that duplication in the result and resolve it with the accountable manager.
Ask the owner to keep a short note covering when to use the workflow, how to check it and what to do when it fails. Make it usable by a colleague who missed the demonstration. Then test whether that colleague can obtain an acceptable result without the original enthusiast sitting beside them.
I'd do that second-person test before expanding. Watch where your colleague gets stuck and improve the instructions there. Record how much help they need so you can budget the time to bring another team on.
Decide whether to continue, revise or stop
Put the decision date in the calendar before the pilot starts. At that review, inspect completed work alongside the time and cost record. Ask the people who produce and receive the output what changed.
Continue when comparable cases meet the quality standard, the whole workflow improves and the released capacity has a useful destination. Name the person who will keep it running and choose the next group to try it.
Revise when there's a specific, testable reason the result fell short. Perhaps gathering the inputs takes too long, the task is too broad or reviewers lack a common standard. Change that part and agree on another limited test.
Stop when checking costs consume the benefit, important failures persist or a simpler solution serves the business better. Record what you learned so the next team doesn't repeat the same experiment without understanding the problem.
Judge the pilot against the outcome you approved. If you wanted faster responses, ask whether customers are getting them. If you wanted more senior capacity, ask what those people can now spend time on. Use the answers to decide whether to renew the investment.
What I'd ask for at the next leadership meeting
Ask each team running an AI experiment to bring a completed example, its before-and-after workflow record and a recommendation. Keep the discussion focused on what changed for the organization, including the cost of making it work.
If nobody has a baseline yet, choose a workflow and measure how it works this week. If there's a measured benefit but no plan for the capacity, assign that decision to the manager. If the evidence says to stop, stop and redirect the effort.
For leaders who still need personal experience with the tools, The CEO's Guide to Getting Started with AI gives you a plan for your first week. For a broader rollout, The Leader's Guide to Implementing AI covers ownership, practice and ongoing management.
Bring a completed example to your next leadership meeting. Show what it took to produce, what you had to fix and how you used the time you saved. Then recommend what you want to do next.
Research checked September 26, 2026. The publication date above is the original article date.
Sources
We'd Better Watch Out
Robert M. Solow, The New York Times Book Review. July 12, 1987, page 36. Scan of the original printed page, hosted by The Standup Economist.
Generative AI at Work
Erik Brynjolfsson, Danielle Li and Lindsey R. Raymond, The Quarterly Journal of Economics, Oxford University Press. Published online February 4, 2025; May 2025 issue. Author-hosted copy of the published paper.
Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity
METR. July 10, 2025.
We are Changing our Developer Productivity Experiment Design
METR. February 24, 2026.
The Productivity J-Curve: How Intangibles Complement General Purpose Technologies
Erik Brynjolfsson, Daniel Rock and Chad Syverson, American Economic Journal: Macroeconomics, American Economic Association. January 2021.
Frequently asked
- How do I find out whether AI is improving productivity?
- Choose a recurring workflow and measure it from the request to a finished result your team can use. Include preparation, review and corrections, then compare the time and quality with your current method. Decide where you'll use any capacity you free up.
- How should I count the value of time saved with AI?
- Put a value on the time saved across the whole workflow, then include software, support and setup costs in your calculation. If payroll spending stays the same, track what your team does with the extra capacity, such as serving more customers or spending more time on valuable work. Count cash savings when spending actually falls.
- When should I expand an AI pilot?
- Expand when comparable work meets your quality standard, the full workflow improves and you have a useful plan for the time saved. Ask another colleague to try the documented method before bringing in a larger group. Change or stop uses where checking costs consume the benefit.
Where to go next
Written by Colin Cox
Colin Cox is Co-Founder of The Work Smarter Company, where he leads every training engagement and client relationship personally. A former operator and COO turned executive coach, he is a member of the Alan Weiss Million Dollar Consulting Hall of Fame.