The impossibility of measuring AI productivity
On this page
This is a long one, sorry. I’ve written before that I tend to write when I’m trying to figure something out. This particular question has been bothering me for a while, and every time I thought I’d answered it, I found something else wrong with the answer.
At Flowstate, we’re regularly asked how to measure the return on AI investment. Sometimes people want to know whether their teams are more productive. Sometimes they just need something credible to put in front of the board when it asks about the bill. Fair enough.
I’m interested in this for reasons beyond having a company that needs to answer the question, though. I believe AI can make people much more capable, and I’d like that to be what we build towards. Let agents deal with the repetitive shite. Give people more time for the work they actually want to do, including things they couldn’t previously afford to attempt.
That’s a large part of the thinking behind Flowstate. But I don’t get to assume it’s happening just because I like the idea.
I’ve been reading The Humanist Review, and holy shit, there’s some good writing there. Seriously, go and have a look. Daron Acemoglu’s essay on AI and work is worth reading in full, but this line about the corporate sector stuck with me:1
It will demand more pro-worker tools when it focuses on increasing productivity and innovation, rather than only labor cost-saving.
I think he’s right about that. If I make the case for AI entirely in terms of salaries it could replace, I can hardly be surprised when the conversation turns to cutting jobs. I’d rather be able to explain what the team can now do that it couldn’t before. Or whether getting rid of some tedious process has actually made anyone’s working day better.
The problem is showing it. Someone saying they save an hour a day is interesting, but I still want to know what changed. Did they get more done? Did they finally have time to think about something properly? Perhaps the tool saved them an hour and gave somebody else an hour of checking to do.
So I started looking for a way to compare the work a company does with what it spends doing it. Ideally, something useful across different teams, whether the work was done by employees, contractors, a service provider or agents. I wasn’t expecting every department to have the same idea of success. I did hope we could agree on some of the accounting.
I’ve ended up with an approach I think is worth testing, rather than an answer I’d tell everyone to adopt. Quite a few of my earlier ideas were wrong. I’ve left the useful mistakes in, because explaining why they failed is probably more helpful than pretending I arrived at the current version through exceptional foresight.
The first problem is easy enough to demonstrate: give an agent the straightforward support questions, and the people left dealing with the difficult ones can suddenly look worse at their jobs.
Apparently, helping the team made it worse
Take a support team handling 800 easy cases and 200 difficult ones a month. An easy case needs six minutes of human handling. A difficult one needs thirty. All are resolved to the same standard.
These are made-up numbers.
Before AI, this hypothetical team needed 180 handling hours for its 1,000 cases. That’s 5.56 cases an hour.
Now an agent handles 600 of the easy cases. The people have 200 easy and 200 difficult cases left: 120 hours for 400 cases, or 3.33 an hour.
| What happened | Before | After |
|---|---|---|
| Easy cases handled by people | 800 | 200 |
| Difficult cases handled by people | 200 | 200 |
| Cases handled by the agent | 0 | 600 |
| Human handling hours | 180 | 120 |
| Cases per human handling hour | 5.56 | 3.33 |
| Total cases resolved | 1,000 | 1,000 |
The dashboard reports a 40% fall in the human team’s rate. Average handling time goes from 10.8 minutes to 18.
Nobody has become slower. The people have the harder cases. The easy ones used to dilute the average, and now they’re gone.
The company still gets all 1,000 cases resolved and needs sixty fewer hours of human handling. I’ll come back to what those hours are worth. For now, asking the team to recover its old average would be a fairly stupid response to a change we deliberately made.
Hand the easy cases to an agent
- Cases per person-hour
- 5.56 → 3.33
- Human handling hours
- 180 → 120
- Recognised output
- $7,200 → $7,200
- Allocated handling cost
- $7,200 → $5,280
- Actual spending, payroll unchanged
- $7,200 → $7,680
Cases per person-hour fall 40%. The same 1,000 cases still get resolved, and 60 handling hours are free for something else.
This happened before chatbots. The US Postal Service’s Remote Encoding Center deals with address images its machines can’t decipher. A 2022 account describes machines taking the easier images and leaving people with ones that took longer to interpret.2 Tom Scott’s video from the centre is still one of my favourite examples of an automated system changing a business process, years before anyone called it AI. Intercom makes the same observation about the human caseload left after support automation.3
There is also good research using the very rate I’ve just criticised. In Generative AI at Work, Erik Brynjolfsson, Danielle Li and Lindsey Raymond report an average 15% increase in issues resolved per hour. Their AI assistant helped the people answering customers. It didn’t take over a separate queue of easy cases. They also examined quality and the experience of the work.4
I don’t disagree with counting cases in that setting. I disagree with assuming they’re still comparable after changing which cases the people receive. “Cases an hour” can be useful. It needs a description of the cases.
And even in my neat example, a day with fewer easy questions might be a more demanding day. Nothing in the table tells us whether the job got better.
What are we actually asking?
Did we deliver more? Did it take fewer resources? Did the work achieve anything useful? Did the people doing it have a better day?
Those are different questions. I’ve been guilty of treating “productivity” as though it answers all four.
Counts tell us about volume. Costs tell us what we put in. Lead time tells us how long someone waited. Revenue and margin matter, but they won’t tell us what happened inside every team. Nor will a survey asking people whether they feel faster.
METR’s early-2025 developer experiment makes that last point rather painfully. Experienced developers working in familiar open-source repositories took 19% longer with the available AI tools. Afterwards, they still estimated that the tools had made them 20% faster.5 That is a result from those developers and those tools, not a verdict on coding agents in 2026. METR’s February 2026 update says newer tools probably help more, while explaining why selection effects make that improvement hard to estimate.6
There is, inevitably, a McKinsey survey. In its August 2026 results, 80% of respondents said AI improved their own productivity. Only 37% attributed some impact on company earnings to it.7 A useful gap for a consultancy to have found. Also two different self-reported questions, so we can’t subtract one percentage from the other and declare the missing benefit stolen.
The SPACE framework helped me put a boundary around the question. Nicole Forsgren, Margaret-Anne Storey and their co-authors argue that developer productivity can’t be captured by a single metric or dimension.8 I agree. A measure of delivery efficiency can still be useful. It just has to stop claiming it describes everything else.
I’ll call the proposal a standardised delivery index. It deals with the work delivered and the resources behind it. The business outcome and the experience of doing the job need their own evidence.
I tried the numbers we already had
Each looked reasonable until I gave it an awkward example.
| What I tried | Where it let me down | What I kept |
|---|---|---|
| Counting finished items | The agent takes the easy cases and the people look worse. Splitting one feature into ten tickets creates ten apparent outputs. | Counting completed work |
| Hours, headcount or cost as output | Treat inputs as output and efficiency gains disappear by definition. | The cost side |
| Revenue or margin | The company’s financial result doesn’t give every internal service an attributable selling price. | Business results, alongside delivery |
| Self-reported time saved | People can feel faster while taking longer, as METR’s participants did. | A reason to investigate |
| Vendor outcome units | A vendor’s billable unit isn’t automatically comparable with another vendor’s, or with work done elsewhere. | Evidence about the workflow |
| Story points | A larger estimate can become a better score without any change in delivery. | Relative sizing, if the basis holds up |
| Effort-weighted records | Better on the support example. Still vulnerable to drafts, split tickets, rework and steps that disappear. | Weights for different kinds of work |
The tables aren’t a reason to throw every existing measure away. I needed parts of several. What kept failing was the assumption that one of them could do the whole job.
I tried the candidates against eleven invented cases, including an agent taking the easy work and a team splitting into two. My working notes have the cases, the dummy data and the calculations. Treat it as a record of how the idea changed, not eleven tests establishing that the final version works in a company.
Work isn’t a ticket
My first definition was pleasingly neat: count work somebody asked for and accepted.
Unfortunately, the company I had described didn’t exist.
At Flowstate, our own Linear setup is almost always behind the work actually being done. I don’t need a historical study of issue trackers to recognise that problem.
An engineer notices something, discusses it, fixes it and creates a ticket somewhere along the way. A customer opens a conversation in Intercom. An invoice arrives. A lawyer gives advice on a call. The on-call engineer investigates because keeping the service running is already their responsibility.
We don’t make the work legitimate by putting it behind an approval button. The person who creates the record isn’t necessarily the person who needed the work, either.
What I’m trying to count is a business deliverable or service with a purpose, a scope and a completion condition we can explain. A request can establish those things. So can an agreed objective, a customer problem or a standing responsibility. An engineer shouldn’t need another department to commission a vulnerability fix before it counts.
The definitions will differ by function. I don’t see any way around that.
| Setting | A possible unit | Where we might check it |
|---|---|---|
| Support | A customer problem handled to an agreed standard | Conversation and relevant follow-up |
| Accounts payable | An invoice processed correctly | Invoice, payment record and exceptions |
| Engineering | A defined change or investigation | Discussion, code, tests and deployment |
| Legal | Advice or a matter handled to an agreed scope | Instructions, advice and the receiver’s response |
| Planning | A forecast cycle completed to a defined standard | Forecast, assumptions, review and delivery |
| Reliability | A defined service maintained over a period | Service scope, load and operating records |
These aren’t alternative names for rows in a database. A merged change can be incomplete. An invoice marked paid can be wrong. A silent customer can be satisfied or thoroughly fed up.
The evidence is spread around. One incident might leave a support conversation, an engineering issue, three pull requests and a follow-up email. Counting six records doesn’t establish six outputs. Object-centric process mining is useful here because it represents events connected to multiple business objects.9 It helps join the records. It doesn’t decide what the incident should count for.
Whether companies will share the relevant context, and whether AI can interpret it reliably, are separate problems. I’m leaving them outside this post. For the accounting, assume we can establish a good enough account of the work and its costs, through human review, automation or both.
That still leaves a difficult question: given the evidence, what do we count? Missing evidence should leave work unresolved, not declare it worthless. Otherwise I’m building a measure of who writes the most enthusiastic completion notes.
The accountants have been here before
The support example needs a way to distinguish six minutes of work from thirty. This part, at least, isn’t new.
ACCA describes standard hours as a way to combine different products into one production measure. It separates how much was produced, how much of the available capacity was used and how efficiently.10 The ONS uses cost-weighted activity indices for much public service output.11
I’m borrowing that structure: count outputs by type, then use stable reference weights to add them together. It gives us a scale without requiring every internal service to have a selling price.
The ONS also provides a useful check on my ambition. About a third of the public service output covered by its methodology still uses an inputs-equal-outputs convention, making productivity constant for that part.11 The existence of cost-weighted measurement doesn’t mean somebody has already worked out how to measure every service.
Back to support. At a reference labour rate of $40 an hour, an easy resolution carries a $4 weight and a difficult one carries $20. Multiply each count by its weight and add them up:
We get $7,200 before the agent and $7,200 afterwards. The business received the same 800 easy and 200 difficult resolutions. Making 600 of them cheaper doesn’t make them disappear from output.
Written generally, that’s:
D is delivered output in reference-cost dollars. n is the count of each type and s is its end-to-end reference cost. t identifies the period we’re measuring, and b identifies the reference period. The sum runs across work types.
Those dollars aren’t revenue or savings. They let us compare quantities of unlike work using one declared set of weights.
I initially tried to justify the weight by arguing that historical cost was a floor on value. Surely a company wouldn’t pay four hours for work worth less than four hours?
An extremely generous assessment of corporate decision-making. Companies buy things they shouldn’t, and sensible investments can fail. The old cost tells us something about production. It doesn’t prove what the result was worth.
For a real business, the standard would cover the relevant mix of people, tools, agents and purchased services. Where we have credible role-based effort estimates, fixed reference rates keep the same job from getting a bigger weight because a more expensive person did it. A bought service may instead need a comparable reference price.
Actual salaries and invoices belong on the cost side. The same scoped deliverable should have the same output weight in London and Lisbon. Changing the supplier shouldn’t change it either.
Standard hours can work where they make sense. But an index covering automated work needs more than the remaining human hours, or its weights can approach zero as automation succeeds. Reference resource costs give us a broader basis.
And the reference period needn’t be pre-AI. It needs usable evidence. “Before anyone used ChatGPT” won’t remain a useful date forever.
Isn’t this just story points in dollars?
It could be. First, a correction I’ve seen engineers get rightly annoyed about: story points aren’t hours. They were meant to size work relative to other work, by complexity and uncertainty, which is why so many teams use a Fibonacci-ish scale. The gaps widen as confidence drops. Converting points into hours was never the point.
The risk here is different. Taking the team’s estimates, turning them into dollars and calling the result objective would be the same guess with better finance branding.
Ron Jeffries’s apology for possibly inventing story points is funny, but it doesn’t answer this objection.12 Gergely Orosz describes a developer inflating estimates once the team started treating its sprint points as a test of success.13 A reference-cost system offers the same temptation. Put a larger weight on the work and the score improves.
What I want is a standard for a class of delivered work, not a running estimate of how difficult this particular attempt feels. A planning estimate can rise when we find a problem. The output weight shouldn’t rise just because we spent longer. If the scope changed, we need to say what changed.
I’d build the standard from reviewed examples of comparable work and test it on other examples. Fix the weights for the comparison. Record changes so neither the team nor its manager can improve a reported result by enlarging them afterwards.
There’s still judgement in that. Which jobs belong together? Is the expensive case a different kind of work, or an expensive attempt at the same thing? A reviewer needs to be able to challenge both the category and the weight.
Even choosing the average matters. A median describes a typical case. Multiply it by volume and it won’t generally reproduce total resource use in a workload with a long tail. For an expected resource-cost weight, I’d normally start with a mean for a well-defined class and show the spread. I wouldn’t select whichever statistic made the index look nicest.
AI could help suggest weights. Anthropic’s work on estimating task duration found useful ranking information, but Claude overestimated short software tasks and underestimated long ones.14 Getting jobs in roughly the right order isn’t enough when their relative weights determine the result.
A points system could adopt similar controls. I don’t get to claim victory by changing the name. The test is whether another team can apply the definitions and whether a reasonable disagreement changes the conclusion.
If that doesn’t hold, then yes, I’ve reinvented story points in dollars.
Did we finish it, or just stop talking about it?
Fin gives us a useful concrete example. Intercom counts confirmed resolutions and assumed ones, where the customer leaves after an answer without asking for more help. Its documentation also says the charge is reversed if the customer returns to the same conversation seeking further assistance.15
That’s a clear billing rule. It leaves a question about silent customers, but it’s unfair to discuss it without the reversal rule. And human-closed conversations deserve the same scrutiny.
For this index, I’d specify what counts as completion for each kind of work. Sometimes there is an explicit acceptance. Sometimes there is a test, a payment record or evidence of delivery. Sometimes we have only indirect signals and need to show that uncertainty. “Closed” can’t settle every case.
Rework isn’t simple either. A reopened conversation may be a failed answer or a new question. A code reversal may fix a defect or reflect a changed product decision. A later dispute doesn’t automatically invalidate the legal advice that preceded it.
We need a rule for which corrections reverse earlier output, how much they reverse and which belong to new scope. Use it for people and agents alike. Compare work that has had the same opportunity to reveal defects, and revise the original period when earlier credit no longer holds. A freshly closed queue shouldn’t look better merely because nobody has had time to complain.
The checking takes resources too. GDPval’s review-and-redo scenarios show how much the apparent cost advantage can change once expert review is included. They are modelled scenarios, not observed company savings, and the paper notes that equivalent review and failure costs aren’t included for its human baseline.16 I’d want the whole workflow costed on both sides.
Planning exposed a different mistake in my earlier version. I suggested giving a forecast cycle no credit if the planner’s adjustments made it worse. That confused completing the work with getting a favourable outcome.
Forecast value added is useful for evaluating adjustments over comparable observations.17 It isn’t a universal switch for whether planning happened. A sound forecast can meet its brief and still turn out wrong. Good legal advice can stop a deal. An experiment can be useful precisely because it tells us to abandon the project.
If the measure only recognises good news, it will miss some good work.
Forty drafts of what?
Now give the index some marketing work.
A team used to produce four ad variants a week. Each took two hours, giving each an $80 reference weight at our illustrative rate. The week’s output was $320.
With AI it produces forty. Apply the weight to every generated file and output becomes $3,200, while the campaign’s results barely move.
I initially treated that as proof the measure was broken. It might be, but not for the reason I thought. Forty distinct, comparable deliverables could be ten times the production without being ten times as useful. I’d already said I wasn’t measuring value, then expected the number to do it anyway.
The other question is whether those forty files were forty deliverables. They could be drafts used to choose what goes into one experiment.
For this example, suppose the service is a campaign experiment for a specified audience, with a defined test question and reporting requirements. I’d count that experiment. Generating four drafts or forty doesn’t change the unit. Their production and selection are part of its cost.
If the team runs another genuinely distinct experiment of comparable scope, that is more output. Calling ten variations on the same test ten experiments isn’t. We’d need the definition before seeing the result, not an inventive explanation afterwards.
In another workflow, producing ready-to-use variants could itself be the service. Then individual variants may be the right unit. That’s a different reporting boundary, and I wouldn’t switch between the two depending on which produced the larger number.
METR’s Tom Cunningham and Parker Whitfill helped clarify why the value question remains separate. They distinguish AI’s effect on the old task mix, the new task mix and the value produced. Making something cheap changes what people choose to attempt, as well as how quickly they finish yesterday’s workload.18
I agree with separating those questions. A production count doesn’t become a value measure because the work was expensive before. Nor does a manager’s approval make forty variants forty times more useful than one.
What happens when a step disappears?
Software breaks the count in the opposite direction.
Suppose a feature used to need a three-hour specification, ten hours of implementation and two hours of review. At $40 an hour, the reference cost is $600.
Now an agent can build the same feature from the context already available. Usage costs $25. Human review takes three hours, representing $120 of effort at our rate. The modelled resources used total $145. Assume the feature meets the same requirements and quality checks.
If I count the old process steps, I lose the specification. The surviving implementation and review have old weights totalling $480. But the business still received the whole feature.
| What we count | Output at reference cost | Resources allocated to delivery |
|---|---|---|
| Original process | $600 | $600 |
| New process, surviving steps only | $480 | $145 |
| New process, completed feature | $600 | $145 |
Counting steps loses 20% of the output because one document became unnecessary. Counting the whole feature keeps the comparison intact.
This only works if the step really was unnecessary. Removing a security check or dropping a requirement would change what got delivered. The same feature, to the same standard, is doing important work in that sentence.
So is “the whole feature”. It can’t mean whatever happens to be called a request. Split a feature into three tickets and it shouldn’t get three feature weights. Combine three genuine features into an epic and two shouldn’t disappear.
Then there are parents and children. If the end-to-end feature already includes design and review, I can’t add the feature’s weight to the same work counted again by those teams. We can attribute contributions within the total. We can’t create extra company output by moving it between departments.
Try reorganising the same work without changing what gets delivered. The company total should stay put. Try outsourcing part of it. Same test.
Long projects need care with timing as well. I would use cumulative comparisons or independently useful milestones whose weights add up to the agreed scope. Otherwise months of cost lead to one enormous completion spike, and a monthly dashboard spends most of the year being misleading.
Fine. What counts as the same feature?
I’ve been rather generous to myself with that fifteen-hour feature. A spelling correction and a billing system can both be called a feature. Putting them in one category would bring us straight back to the support problem.
Let’s pick something more specific: a CSV export of a customer’s audit history. An authorised account administrator needs to export eight specified fields for a chosen date range. The brief defines the volume and response-time limits, accuracy requirements and access controls. It must work in the live product.
I’d count one delivered export capability of that scope. Not the button, the endpoint, the tests and the documentation separately. Nor would I count it every time somebody downloaded the file. Here we’re measuring delivery of a change, not operation of the product afterwards.
In this fictional company, suppose we find five earlier report-export changes with comparable data scope, user behaviour, operating limits and assurance requirements. On the same reference-price basis, their end-to-end costs were $480, $560, $600, $640 and $720. The mean is $600.
That’s how we could construct the weight. Five convenient numbers don’t establish that it’s a good one. The reviewer has to check what made those jobs comparable, rather than simply searching for the word “export”. We’d then test the class and its weight on other work.
Now we have a definition to argue with:
| What changes | What I would count |
|---|---|
| One issue becomes nine | One $600 unit either way. |
| The agent removes the separate specification step | One $600 unit, provided the same requirements are met. |
| An existing library makes the build much easier | One $600 unit. Reuse changed the production cost, not the delivered scope. |
| The team spends thirty hours on an awkward implementation | One $600 unit. The extra effort goes on the cost side. |
| A provider delivers it for a fixed fee | One $600 unit. The fee and our oversight go on the cost side. |
| It fails the agreed access-control checks | No completed unit yet. It hasn’t met the brief. |
| The customer also needs a real-time external event stream | New scope. The export class doesn’t cover it. |
Notice that the class doesn’t depend on the chosen implementation, the number of pull requests or the salary of the engineer. Those can affect cost without changing what the business received.
The event stream is different. It changes the required behaviour. We’d need a suitable reference class or another explicit treatment for that scope, applied consistently to earlier and later work. We can’t invent a “very complex export” category after an expensive build and award it more output.
What about a new reporting platform with no credible precedent? I wouldn’t force it into this class. I’d describe its scope, costs and useful milestones, and keep it outside this comparable series until we have a basis for including it.
Its cost must remain visible. The report needs to reconcile the measured service with unmatched work and the rest of the spending. Otherwise I’d be selecting the easy-to-classify successes and leaving all the awkward investment elsewhere.
This is the hardest part of the proposal. Another reviewer might reject the export class or its $600 weight. At least we can locate the disagreement and see whether another reasonable choice changes the answer.
Some engineering work may support classes like this. Some may not. Finding that out would be a useful result, even if it ruins my hope of a company-wide total.
The sixty hours haven’t left the payroll
Back to support, and the saving I was tempted to report.
Originally, 180 handling hours at $40 gave us $7,200. Afterwards, 120 hours give $4,800. Add an illustrative agent charge of $0.80 for each of its 600 resolutions and the allocated handling cost is $5,280.
That’s $1,920 less. Except the people are still employed on the same terms.
Suppose payroll in this simplified example remains $7,200. Add the $480 agent bill and the company spends $7,680. The handling requirement fell. The bill went up.
| Which question are we answering? | Before | After |
|---|---|---|
| Cost allocated to handling required | $7,200 | $5,280 |
| Actual spending, payroll unchanged | $7,200 | $7,680 |
| Human handling capacity released | 0 hours | 60 hours |
Robert Kaplan and Steven Anderson make this distinction in their work on time-driven activity-based costing: unused capacity offers opportunities for savings or growth.19 I agree, and it changes the rule here. Released hours go on a capacity line. Savings need evidence of reduced spending or a credible account of spending avoided.
The sixty hours could support more customers without another hire. They could reduce overtime, absorb peaks or make the job less frantic. They could also remain unused. I don’t want the accounting to choose an outcome before the company has done anything with the time.
For the main cost-efficiency index, I’d use the cost of supplying the measured service, including unsuccessful work and capacity left unused. Compare delivery growth with cost growth:
D is the delivery total. C is cost within the same reporting boundary. E starts at 100, so 100 means no change in cost efficiency.
In the support example, delivery stays at $7,200 while spending rises from $7,200 to $7,680. The index is 93.8: cost efficiency fell by 6.25% that month.
Use the separate allocated-handling model and the ratio is 136.4. That’s the improvement in resources required for the workflow, not a realised financial gain. Both numbers are useful once they’re labelled. Putting “AI savings” above either one would save a lot of explanation and create a much larger problem.
For a real cost account, include payroll and on-costs, contractors, bought services, licences, agent usage and infrastructure. Review, maintenance and measurement also consume resources. Don’t add review labour twice if it’s already in the payroll total.
A provider’s fee belongs in the account once, alongside our own oversight. We don’t also invent its internal payroll and token costs. Outsourcing shouldn’t make the purchase cost disappear, or count it twice.
We also need a consistent basis over time. Period expenses aren’t cash payments. A build paid for this quarter may serve several years. Choose and disclose the treatment. Don’t expense the build in one comparison and spread it across years in another. The support example assumes expense and cash spending coincide.
Prices matter too. Cheaper tokens improve the economics even if nothing about the workflow changes. A pay rise does the reverse. Where input quantities and quality can be compared, a constant-price view helps separate price changes from resource use. Without that evidence, show the nominal cost result and explain its limits.
Finally, sixty fewer handling hours don’t establish that a whole position can go. Staffing still has to cover the work when it arrives. The next question is about demand and service commitments, not division by a convenient number of hours.
The error I thought would cancel
I’d written that a wrong weight would cancel when we compared a team with itself. The same mistake appears on both sides, so surely we’re fine?
No. Not when the mix moves.
Take two types of work. For this example, stipulate that their correct reference weights are $1 for simple work and $10 for complex work. In the first quarter, the team completes ninety simple units and one complex unit. In the second, ten simple and nine complex. Correctly weighted output is $100 in both quarters. Hold total resources unchanged too.
Now give complex work an incorrect $20 weight.
| First quarter | Second quarter | |
|---|---|---|
| Simple units | 90 | 10 |
| Complex units | 1 | 9 |
| Output at the correct weights in this example | $100 | $100 |
| Output with complex work weighted at $20 | $110 | $190 |
We report 72.7% growth where there was none in the correctly weighted measure. The error stayed fixed. We did more of the work we’d overvalued.
A wrong price tag can fake growth
Real growth 0.0%. The index says +72.7%. Same work, different mix.
The relationship is:
G is the growth ratio using the correct weights in the example. Ĝ uses our estimated weights. e is each type’s proportional weight error and w is its share of correctly weighted output. ē is the average error after weighting by those shares.
The errors cancel when that average is the same in both periods. An unchanged work mix would do it. So would getting every weight wrong by the same proportion. Other combinations can also cancel. Those two aren’t the only possibilities.
The practical check is to look at what gained share. If the apparent improvement comes from a category whose weight we barely trust, try plausible alternatives. Does the gain survive? The type-level results should sit beside the total, not require a special request from the person who doubts it.
The categories can conceal another shift. If the agent takes the easiest of the “easy cases”, even my two-category support model may adjust too little. We have to check what changed inside each category as well.
We can also get better at seeing the work
Even with perfect weights, the records can fool us.
Suppose real output grows by 10%. We capture 80% of the correctly weighted output in the first period and 84% in the second. The observed comparison becomes:
The dashboard reports 15.5% growth. Four percentage points of better coverage added 5.5 percentage points to apparent growth. With no real growth, that same change in coverage would manufacture 5%.
An agent may leave an exhaustive log of work a person would have finished through a conversation. Better records would be welcome. They wouldn’t, by themselves, be more production.
“Coverage” in that equation means the share of output captured, weighted on the same reference basis. It doesn’t mean the percentage of payroll attached to an integration, connected seats or visible spending.
Knowing where 80% of the money went doesn’t tell us that we found 80% of the work. That only follows with further assumptions about the observed and missing work. Keep spending coverage as a useful financial check. Don’t put it into the output equation because it’s easier to obtain.
This is why I kept the coverage example even though I’m leaving the evidence-collection system outside the post. A change in what we can see changes the measurement, whichever way we collect the records.
More records won’t automatically fix it. Volume can reduce random noise. A consistent bias can remain. Nor does four quarters of data automatically halve uncertainty. Errors shared across quarters don’t behave like independent mistakes.
I’d rather publish a narrower series we observe consistently than make a whole-company claim on coverage we can’t defend. It needs a clear boundary, the same follow-up windows and checks for changes in recording. Work outside that boundary stays visible as unmeasured work.
Earlier simulations helped me explore these errors. I haven’t retained their numerical bands as evidence of how accurately the method would perform. That needs the implementation and assumptions published, then checked against real work. The examples here establish the failure modes without pretending to know how often each will occur.
Eventually, even the baseline becomes a problem
We can’t freeze one set of weights forever and assume the business will obligingly stay comparable. Services change. Some disappear, and new ones don’t come with an old price.
One option is to update the weights periodically, comparing each pair of adjacent periods with the same weights. An annual link using the previous year’s weights is:
The numerator counts this year’s output using last year’s weights. The denominator counts last year’s output using those same weights. Multiply the links for a longer-running index.
That avoids calling a change of weights a change in production. It also has a catch: separately chained components generally don’t add up to the chained total. The BEA explicitly warns about treating chained-dollar components as additive.20
So build the company series from its own non-overlapping set of outputs. Don’t add independently designed team dashboards and hope they’re measuring compatible things. Some internal work may be useful in a function view but already included in the end-to-end company deliverable.
A changed classification also needs a bridge: an overlap period, a credible reconstruction of the earlier series or a disclosed break. Chaining won’t repair a bad service definition or missing work we only just found.
This may leave us with useful function-level comparisons and no defensible company-wide total. I’d accept that. It would be better than inventing a total because the original brief asked for one.
Sometimes good work makes the count smaller
Suppose engineering fixes the defect behind a thousand support conversations. Our support-case index reports fewer cases handled.
It should. There were fewer cases. The mistake would be deciding the company had become less productive.
At a wider boundary, we’re trying to help customers use the product successfully. Maintaining or improving that service with fewer avoidable contacts may be a substantial gain. The narrow case count needs the wider service definition or the outcome measures alongside it.
Reliability and prevention have the same problem. A quiet on-call week can be a good week. A useful compliance intervention may stop a case from existing at all.
For those functions, I’d test units based on the service maintained over a period: what was kept working, for whom, under what load and risk, to what standard. An entry on the rota isn’t enough. Neither would I turn a month’s output into zero because one threshold was missed. The quality adjustment needs to reflect the failure rather than make everything depend on one switch.
These units are harder to define. That’s part of the problem I’m trying to understand, not something a ticket count lets us avoid.
What about work we couldn’t do before?
An audit team that sampled transactions can now examine all of them. Put the old hypothetical human cost against every additional check and the output claim becomes enormous.
But nobody was necessarily going to buy that manual process. More checks also don’t mean a proportional increase in assurance.
If the new service is comparable with something we already measure, we can account for the changed scope or quality on that basis. If it isn’t, I would report the capability, its cost and what we expect it to achieve separately while we establish a useful definition.
I’m calling that a capability ledger. It doesn’t mean the spending qualifies for capitalisation. It doesn’t mean we can hide expensive work under “innovation” whenever the efficiency ratio disappoints. The account must still reconcile to the finances.
Once the service becomes repeatable and definable, it can join the delivery index under an explicit introduction rule. Being enabled by AI shouldn’t keep it outside forever.
This is where Acemoglu’s point about corporate demand comes back in. If companies only reward labour-cost savings, they’ll only buy tools that replace people.1 A capability ledger is one way to reward the other kind.
My contribution here is narrower. Give the company somewhere to explain what the investment makes possible, rather than asking every project to justify itself as cheaper old work. An owner, evidence and a review date would be a useful start. “It’s AI” isn’t an investment case.
Acemoglu also argues against tax distortions that favour capital over labour, drawing on work with Andrea Manera and Pascual Restrepo.121 I agree with removing an artificial preference for replacement. For this post, though, the decision is inside the budget: what does the existing service now cost, and what can we do that we couldn’t do before?
The ledger won’t make that decision for us. It should make it harder to avoid.
Would I put this in front of a CFO?
Yes, but alongside the rest of the account, not as a score for the company.
Kent Beck and Gergely Orosz make a strong case against effort and output targets in their response to McKinsey. Discussing its custom metrics, they write: “Customers don’t care. Executives don’t care. Investors don’t care.”22
I think that dismissal goes too far. A CFO can reasonably ask whether we can deliver the same payroll service, to the same standard, with fewer resources. We shouldn’t have to wait for a movement in company profit to investigate why it became more expensive.
Their argument is more nuanced than that sentence. They also recognise the use of effort and output when diagnosing problems, and the dangers of judging people only by outcomes.13 In his concluding section, Orosz recommends using effort and output to investigate problems rather than making them the public measures of success.
That’s the narrower disagreement worth having. I think a regularly reported comparison of a defined service’s output and cost might help, alongside outcomes. It has to justify its cost and its effect on behaviour. A beautiful dashboard that turns the team into full-time dashboard improvers has failed.
This is roughly what I’d want to see:
| Question | What belongs in the report |
|---|---|
| Did we deliver more for the resources supplied? | Comparable delivery, reconciled cost and the effect of uncertain weights |
| Was delivery faster and more reliable? | Lead time, queues, rework and quality measures |
| Did it matter to the business? | Relevant outcomes, such as adoption, service quality or financial benefit |
| What happened to released capacity? | More delivery, credible hiring avoidance, lower spending, resilience or time still available |
| Did the job get better? | Workload, autonomy, learning and the burden of review |
The outcome measures should differ by function. Bookings belong in a sales discussion. They aren’t a sensible acceptance test for legal advice. I’d rather show those differences than conceal them inside an “impact multiplier”.
At company level, the financial account must also reflect how we bought the work. Revenue per employee can rise after outsourcing even if total delivery cost rises. Add the relevant bought services and AI costs before deciding the business has become more efficient. Revenue still has other drivers, so that ratio won’t attribute the change to AI.
I wouldn’t use the delivery index to rank individuals. Its weights describe classes of work, not the full contribution of one person. Mentoring, interruptions and helping someone else finish a job don’t fit into an individual’s output count. Tie pay to the index and we give people a reason to argue for larger weights and more countable work.
Even comparing AI-assisted and human-only work needs the same care as the opening example. The assignments may differ within a category. Knowing that an agent was involved doesn’t prove it caused the improvement.
And a finished artefact doesn’t prove the person responsible understood it. That’s another problem, particularly if we claim the purpose is to make people more capable.
The test I haven’t done yet
I’ve shown how a few calculations behave with inputs I chose. I’ve borrowed accounting methods and learned from other people’s objections. None of that tells me whether two teams can use this proposal on real work and reach a useful result.
That’s the next test.
I’d start with three settings: a repeatable case or transaction workflow, a project workflow and an ongoing service. For each, write down the unit, scope and reporting boundary before seeing the results. Add the completion rule, quality treatment, reference weights and cost basis. Say what happens to unmatched work and later corrections.
Then give another competent team the same evidence. Can they reproduce the calculation? Where do they disagree about what belongs in it?
The arithmetic should be easy to agree on. The export example shows where the real argument will be. If two reasonable choices of category reverse the headline result, the report needs to show that rather than settle the disagreement with a decimal place.
Choose reference classes and estimate their weights using one set of work, then test them on another. Include excluded records and work outside the main tracking system in the review. Don’t redraw the categories until the history looks convincing.
Repeat the split-ticket, merged-ticket, reorganisation and outsourcing tests. Check for changes in recording. The company total shouldn’t move when all we’ve changed is how we describe or purchase the same work.
First test this accounting with an independently established account of the work. Whether a machine can recover that account from the available systems is a later implementation test. If humans with adequate evidence can’t agree on the units, a better classifier won’t rescue the definition.
Testing whether AI caused an improvement is another job. Randomised access can help where feasible, as can a carefully designed staged rollout with a credible comparison group. Learning, spillovers, task selection and equivalent quality follow-up all matter. Enthusiastic adopters and reluctant ones aren’t automatically comparable because they share a job title.
How accurate does the measure need to be? It depends on the decision. A rough signal for where to investigate has a different burden from evidence used to cut staffing or report savings. The acceptable error should follow that decision, not a universal audit sample size.
Publish the definitions, calculation and revision policy with the result. My bar is that another team can apply the method, challenge its choices and see whether the conclusion survives. Recognition from an accountant or an impressive-looking equation wouldn’t be enough.
I have a better question now
I started with “did AI make us more productive?” I now want to know which service changed, whether we’re counting the same thing and where the released capacity went. Those questions are less convenient on a slide. They’re much more useful when someone asks what we should do next.
This matters to Flowstate because connecting work with the resources behind it is the problem we’re trying to solve. I don’t want the product to depend on defending a formula I became fond of while writing a blog post. If real work breaks the model, the model has to change.
I still want agents to take the repetitive work out of people’s days. I want a company to recognise the benefit when someone gets time to think, helps a colleague or attempts something new, rather than demanding more tickets to prove the software purchase worked.
I’m not sure how much of that fits in one index. Perhaps less than I hoped. A useful result might be a set of comparisons we trust, with the gaps left plainly visible.
When the agent takes the easy cases, the remaining people shouldn’t need to manufacture activity to defend themselves. When somebody claims a saving, we should be able to ask where it went. And when the proposed number doesn’t answer the question, we should say so before it becomes somebody’s target.
I went looking for a measure. I’ve ended up with a method I’d test and a fairly long list of reasons to be careful. Given where I started, I think that’s progress.
Technical notes
These are the workings behind the examples. The assumptions matter: an identity can be exact without telling us how accurately we can estimate its inputs in a company. The ratios below assume positive reference weights and non-zero denominators.
The weight-error identity
Let the stipulated correct reference weight for type k be s, and let the estimated weight be:
For accurately counted output, define the correct weighted total and each type’s share as:
Then:
Taking the ratio between periods gives the identity used above. It assumes fixed type-level proportional weight errors, accurate quantities and a common output definition. It doesn’t model omitted work, changing classifications or uncertain deliverable boundaries.
For small errors, the relative error in the growth ratio is approximately:
If, additionally, the type errors are independent, mean-zero random variables with a common standard deviation, the first-order standard deviation of that growth-ratio error is:
Under those specific assumptions, transferring twenty percentage points of output share between two types and setting the error standard deviation to 20% gives approximately 5.66% relative error in the growth ratio, one standard deviation. It isn’t an empirically established error band, and it isn’t generally the same number of percentage points of reported growth. Correlated or systematically biased standards require a different calculation.
Coverage is an output concept in the identity
In the coverage-only example, let c be the share of correctly weighted output visible in the records, with no false positives and no other errors. Then observed output is c times complete output, so:
The equation is exact under those definitions. Estimating c is the difficult part. A ratio of connected spending to total spending is a different statistic and can’t be substituted without an explicit model linking resource coverage to production coverage.
The accounting boundary matters here too. Observing part of the output while using the entire cost base doesn’t estimate the same object as measuring both output and cost for a deliberately restricted, consistently observed service. Neither should be relabelled as whole-company efficiency without justification.
Why overall classifier accuracy is insufficient
For a simple binary recognition problem with correctly assigned, fixed weights, define true-positive, false-positive and false-negative totals using those weights rather than item counts. Weighted precision is true-positive weight divided by all recognised weight. Weighted recall is true-positive weight divided by all eligible weight. Where the denominators are non-zero:
Ordinary count-based precision and recall don’t give that identity for a cost-weighted index. Misclassification that changes the weight of a recognised item requires additional treatment. So do candidates that were never found and uncertain relationships between parent and child records. Combining these errors into one multiplicative equation requires compatible definitions. It isn’t justified simply because each term has a plausible name.
Probabilistic recognition is possible, but a proposed estimator must keep classifications mutually exclusive and avoid counting overlapping deliverables. A collection of independently scored tickets doesn’t satisfy those conditions automatically. Calibrated probabilities address uncertainty in observed candidates, not work missing from the candidate set entirely.
What an audit sample can establish
The familiar estimate of 385 observations comes from a particular calculation: a simple random sample of an independent binary proportion, a 95% normal-approximation interval, a worst-case proportion of one-half and a margin of five percentage points. The unrounded sample size is:
It doesn’t establish that 385 records will calibrate reference costs, validate every work type or resolve a small change between periods. Stratification, clustering, unequal weights, rare high-cost cases and reviewer disagreement change the required design. Audit planning should start from the error that could change the business decision, not from a familiar round number.
Reproducibility and interpretation
The worked examples use the stated inputs. They aren’t estimates of typical company performance. Numerical Monte Carlo error bands would also need the implementation, parameter distributions, dependence assumptions and random seeds published. We would then need to establish which assumptions fit observed work before interpreting the bands as expected accuracy.
The delivery sum has units of reference-cost currency. A presentation index can normalise that sum to 100 in the base period. The cost-efficiency index already uses that normalisation. Neither measures economic value, causal AI impact or employee worth.
Will Hackett is co-founder and CTO of Flowstate.
Footnotes
-
Daron Acemoglu, Will AI Replace Workers? Not If We Build It Right., The Humanist Review of AI, 15 July 2026. Argues for tools that complement workers and for corporate demand centred on productivity and innovation rather than labour-cost savings alone. The connection to the reporting design here is my inference, not a method proposed in that essay. Source. ↩ ↩2 ↩3
-
National Association of Letter Carriers, The Remote Encoding Center: Where bad addresses go to get better, The Postal Record, July 2022, pp. 34 and 35. The account describes machines taking easier address images and leaving harder ones for human keyers. Source. ↩
-
Intercom, How Fin AI Agent and Copilot Cut Handle Time and Boost Agent Productivity. The discussion of the remaining human caseload is vendor commentary supporting the work-mix mechanism, not independent evidence for the size of a productivity gain. Source. ↩
-
Erik Brynjolfsson, Danielle Li and Lindsey Raymond, Generative AI at Work, The Quarterly Journal of Economics 140(2), 2025, pp. 889 to 942. The published study reports an average 15% increase in issues resolved per hour in its customer-support setting, with differing speed and quality effects across workers. Published article. Authors’ summary. ↩
-
Joel Becker, Nate Rush, Beth Barnes and David Rein, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, METR, 10 July 2025. The result concerns the participating experienced developers, their repositories and the tools available in that experiment. Source. ↩
-
Joel Becker and colleagues, We are Changing our Developer Productivity Experiment Design, METR, 24 February 2026. The follow-up discusses participant selection, task selection and difficulties measuring time with concurrent agents. Source. ↩
-
McKinsey, The state of AI in 2026: On the road to ROI, 25 August 2026. Survey responses were collected from 1,719 participants between 4 May and 8 June 2026. The reported 80% individual-productivity and 37% EBIT-impact figures answer different questions. Source. ↩
-
Nicole Forsgren, Margaret-Anne Storey, Chandra Maddila, Thomas Zimmermann, Brian Houck and Jenna Butler, The SPACE of Developer Productivity: There’s more to it than you think, ACM Queue 19(1), 2021. Argues that developer productivity can’t be captured by a single metric or dimension. The bounded delivery index proposed here is not presented as a replacement for that framework. Source. ↩
-
Alessandro Berti and colleagues, OCEL (Object-Centric Event Log) 2.0 Specification, arXiv:2403.01975, submitted 4 March 2024. The specification supports events and relationships involving multiple business objects. It is a representation standard, not a productivity metric. Source. ↩
-
ACCA, The standard hour in performance measurement. Standard hours provide a common activity measure for heterogeneous products and support separate volume, utilisation and efficiency ratios. Source. ↩
-
Office for National Statistics, Public service productivity estimates: sources and methods, revised 1 May 2026, especially section 1 on output, inputs and index numbers. The methodology uses cost-weighted activity measures for much, but not all, public service output. Source. ↩ ↩2
-
Ron Jeffries, Story Points Revisited, 23 May 2019. His qualified apology for possibly inventing story points is the reference here, not evidence against every form of relative estimation. Source. ↩
-
Gergely Orosz and Kent Beck, Measuring developer productivity? A response to McKinsey, Part 2, The Pragmatic Engineer, 31 August 2023. Their joint discussion addresses outcome-only incentives. The separately identified closing section is Orosz’s and includes the estimate-inflation example and recommendation to use effort and output to diagnose problems rather than as public success measures. Source. ↩ ↩2
-
Alex Tamkin and Peter McCrory, Estimating AI productivity gains from Claude conversations, Anthropic, 25 November 2025. The software-task check reports Spearman correlations of 0.44 for Claude and 0.50 for developers, alongside compressed model estimates. Ranking performance does not establish calibration in hours. Source. ↩
-
Intercom, Fin AI Agent outcomes, documentation checked 7 October 2026. See “Resolution definition” and the explanation that later requests for further assistance in the same conversation reverse the resolution charge. The worked examples in this post do not use Fin’s actual pricing. Source. ↩
-
Tejal Patwardhan and colleagues, GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks, 2025, appendix A.2.1 and table 2. The review-and-redo scenarios use specified assumptions and omit comparable review and failure treatment for the human baseline. They should not be read as general workplace savings estimates. Source. ↩
-
Ralf Seifert, Richard Markoff and Matthew Spooner, How a new approach to demand planning can redefine success, I by IMD, 5 August 2024. Discusses forecast value added and the possibility that a better baseline reduces the incremental contribution of human adjustments. Source. ↩
-
Tom Cunningham and Parker Whitfill, Task Substitution and Uplift, METR, 8 May 2026. Distinguishes uplift on old tasks, new tasks and value, with relationships derived under explicit assumptions. Source. ↩
-
Robert S. Kaplan and Steven R. Anderson, Rethinking Activity-Based Costing, Harvard Business School Working Knowledge, 24 January 2005. Distinguishes supplied capacity from capacity consumed and discusses opportunities arising from unused capacity. Source. ↩
-
US Bureau of Economic Analysis, Chained-dollar estimates. Notes non-additivity outside the reference year and cautions against using components as though they were ordinary additive dollar values. Source. ↩
-
Daron Acemoglu, Andrea Manera and Pascual Restrepo, Does the U.S. Tax Code Favor Automation?, Brookings Papers on Economic Activity, Spring 2020. Analyses tax treatment that favours investment in equipment and software over labour. Historical estimates in that paper are not a calculation of any particular company’s current tax treatment. Source. ↩
-
Gergely Orosz and Kent Beck, Measuring developer productivity? A response to McKinsey, The Pragmatic Engineer, 29 August 2023, especially section 4. The quoted dismissal concerns McKinsey’s custom effort and output metrics, not every possible form of operational measurement. Source. ↩