Working notes · 7 October 2026
Workings: trying to measure work done by people and AI
These are my notes from working on the post. They’re not tidied up. Most of it was written while I was arguing with myself, so it repeats, changes its mind and occasionally contradicts the post, because the post is where I ended up and this is how I got there.
If you only want the argument, read the post. If you want to check my working, or steal the test cases, carry on.
What I was actually trying to do
- customers keep asking “is AI making us more productive”. finance wants a number they can defend to a board
- every number they have was built for a workforce of people. tickets, story points, hours, headcount
- wanted ONE unit that works across functions. support, eng, finance, legal, marketing, whatever
- had to be neutral: same work earns the same whether a person, a person + AI, or an agent did it
- had to be computable from systems a company already has. no new timesheets. nobody wants to narrate their day to a machine
- assumed (and still assume) LLMs + company context do the messy interpretation. the post leaves that out on purpose, these notes don’t
Pass criteria I set before testing anything
- one definition across functions, no custom model per team
- computable from existing systems, with AI filling gaps
- neutral between people and agents
- answers what leaders actually ask: faster? cheaper? fewer people? different mix? what do we need next?
- hard to game, at least against the known tricks in each function
- has precedent an accountant or auditor would recognise
- stable over time, including once pre-AI history runs out
Nothing passed all seven. The final version gets closest and still fails 5 and 7 in places (see the mix and coverage cases below).
The functions I tested against
I wanted functions that are genuinely different, not different job titles doing the same thing. Went through 18. Most teams turned out to run three shapes of work at once. Five shapes kept coming back: transactions, cases, cycles, projects, advice. Two more broke volume counts outright: ongoing duties (nothing to count when it goes well) and judgement calls (doing more can make it worse).
| Function | Shape | Where it’s recorded | What arrives | What says it was good | Where value shows up | Constraints that bit me |
|---|---|---|---|---|---|---|
| Accounts payable | transactions + exceptions | ERP, AP tools | invoices | paid right, no duplicate | discounts kept | most invoices are touchless now, the people get the exceptions. residue again |
| Customer support | cases | Zendesk, Intercom | tickets | not reopened | retention, months later | vendor “resolution” units count customers who just leave |
| Address keying (USPS) | transactions | image queue | unreadable addresses | delivered | none really | the original residue example, machines took the easy ones |
| Insurance claims | cases, years long | claims systems | claims | not reopened, not overpaid | loss ratio, years later | complexity is assessed at intake, not at the end |
| Software engineering | projects + tickets | Jira, Linear, GitHub | plan items, bugs | merged, not reverted | adoption | tickets lag the work, split arbitrarily, closed late. see below |
| Sales | deals | CRM | pipeline | won or lost | bookings | only function with money on the record. timing games around commission |
| SDRs and AI SDRs | activity | outreach tools | leads | meeting held | pipeline | activity is free now, meetings aren’t |
| Demand planning | cycles + judgement | planning tools | monthly cycle | forecast value added | stock and service | planner changes often make it worse |
| FP&A | cycles + advice | Excel, email, planning tools | close, ad hoc asks | used in a decision | almost never recorded | most of the work is a question in Slack |
| Legal | cases + advice | contract tools, matter systems | contracts, questions | signed, no dispute | something bad that didn’t happen | good advice can kill a deal |
| Brand marketing | duty + creative | Asana, Workfront | plan | ran, reached people | years later | cheap drafts, see the drafts case |
| Product management | judgement | roadmap tools | plan | shipped and adopted | months later | output is decisions |
| SRE and on-call | duty | PagerDuty | rota periods | SLO met | incidents that didn’t happen | a quiet week is a good week |
| Internal audit | projects, assurance | audit tools | audit plan | report issued, findings closed | counterfactual | AI lets you test everything, see capability case |
| Recruiting | projects | ATS | requisitions | hire still in post | quality of hire | quality of hire is rarely measured |
| Clinicians with AI scribes | cases | EHR | encounters | coded, not denied | outcomes | scribes push billing codes up |
| Team leads | overhead | nothing | nothing | the team’s results | the team’s results | no record at all |
| Pharma R&D | stage gates | notebooks, portfolio tools | portfolio plan | gate passed | a decade later | a failed gate can be good work |
Every one had some record with a created date, closed date, owner, status and type. That’s the only data shape you can count on almost everywhere. Agents are starting to show up in the owner field. Only sales had money on the record.
Edge cases list (the long version)
- residue: machines take the easy work, people look slower
- the cruiser: uses AI, delivers the same, costs more
- the slow person: 4 hours on a 1-hour investor update. the fast one: 20 minutes
- new starters, no history
- new types of work, no baseline
- cheap-work inflation: 40 ad variants instead of 4
- brand-new capability: testing 100% of transactions
- judgement work where doing more makes it worse
- duties and prevention: good work produces nothing countable
- step elimination: AI removes the spec step entirely
- agents opening tickets for other agents
- splitting and reclassifying (DRG creep is the famous one)
- AI inflating the evidence (scribes and billing codes)
- finished but wrong
- lag: value turns up 2 years later
- unlogged work
- roll-up: team to company without double counting
- ageing baselines: “pre-AI” stops meaning anything
The rounds
Roughly ten rounds of propose a measure, throw the cases at it, write down what broke, keep the useful bit.
- count finished items. broke on residue straight away (−40% tickets per hour while the system got better). splitting inflates it. ignores size. kept: count at completion
- inputs (hours, headcount, cost). AI gains invisible by definition. the ONS does this for about a third of public service output and says outright that productivity is constant there.1 kept: the cost side
- money (revenue, throughput). only sales has a price. attribution is a swamp. Google gave up on Shapley-style attribution in GA4 and went to experiment-trained models, which says a lot.2 kept: value shown alongside, never instead
- time saved, self-reported. METR: experienced devs 19% slower, believed 20% faster.3 Danish chatbot users: 2.8% of hours saved, nothing in earnings or recorded hours.4 kept: AI time estimates as a starting point for weights, nothing more
- vendor outcome units. every vendor bills its own unit. Fin counts an assumed resolution when the customer leaves (and reverses it if they come back).5 Zendesk split resolutions into tiers and only bills verified ones.6 Devin bills “agent compute units”. kept: vendor counts as evidence, normalised
- unsized flow items (flow velocity etc). splitting inflates it, residue still breaks it.7 kept: lead time and WIP alongside output
- standard effort × impact weighting. this was the first spec I actually liked. handled residue, cruiser, slow person, new starter. then: drafts 10x’d output, planners got paid for changes that made forecasts worse, duties earned nothing, step elimination showed as lost output, impact weights were pure opinion. Medicare’s RVU committee is the cautionary tale for opinion weights: CMS agreed with 87.4% of its recommendations.8 kept: standard effort per type
- standard effort + demand gate + acceptance gate, cost weighted. demand gate = “only counts if someone outside the team asked for it or it’s in an approved plan”. felt clever. step elimination still undercounted by 20%. prevention, agent-made demand, splitting, new capability all broke it. kept: acceptance, cost weighting
- round 8 + count at the request level, end-to-end standards, service periods for duties, a new-capability ledger, chain-linking. nothing in my list broke it. which mostly told me my list was too kind
- review round. got torn apart, mostly fairly. the big ones:
- historic cost ISN’T a floor on value. companies buy rubbish
- capacity released isn’t cash
- weight errors don’t cancel when the mix moves (I’d claimed they did. they don’t, see mix case)
- parent/child double counting survives “count at request level”
- the demand gate is a permission system that doesn’t exist. engineers fix things nobody asked for and that’s often the best work they do. dropped it for “defensible purpose, scope and completion condition”
- spending coverage ≠ output coverage. I’d used the share of payroll I could see as if it were the share of output I could see. wrong
- my Monte Carlo error bands looked lovely and proved nothing about a real company. pulled them from the post
The test cases
Each of these is invented, with dummy numbers picked to make the mechanism obvious. The structures below are generated from research/unit-of-work/trials.py, so they can’t drift from the maths. “WRONG” means the measure would mislead you, “meh” means half right, “ok” means it tells you what actually happened.
The agent takes the easy work
the one that started all this. does the measure blame the people when the machine takes the easy stuff? anything that counts items per hour fails. weighting by type fixes it.
case: residue
setup: A support team resolves 800 easy cases (6 minutes each) and 200 difficult ones (30 minutes) a month. An agent takes 600 of the easy ones at $0.80 each.
truth: Every case still gets resolved. Nobody got slower. The people need fewer hours.
data:
Cases resolved by people 1000 -> 400
Cases resolved by the agent 0 -> 600
Human handling hours 180 -> 120
Allocated handling cost $7,200 -> $5,280
results:
Items finished −40% WRONG # per person-hour: 1000 ÷ 180 = 5.56, then 400 ÷ 120 = 3.33
Hours as output −33.3% WRONG # 180 hours, then 120
Effort-weighted records no change ok # 120 people-hours of standard + 600 × 6 min = 180 standard hours, as before
Standardised delivery index no change ok # 800 × $4 + 200 × $20 = $7,200 in both monthsAI makes drafts nearly free
what happens when AI makes something nearly free and people just make loads of it. first version of this test assumed a manager-approved plan of 10. dropped that, nobody works like that. the real question is what the deliverable is: a draft or the thing that gets used.
case: drafts
setup: A marketing team used to make 4 ad variants a week, 2 hours each. With AI it generates 40, and 10 go into independently defined tests.
truth: 10 things were delivered. The other 30 are drafts, part of the cost of making them.
data:
Variants generated 4 -> 40
Variants used in tests 4 -> 10
Reference weight each $80 -> $80
results:
Items finished +900% WRONG # 4 files, then 40
Hours as output −25% WRONG # 8 hours a week, then 6 (illustrative)
Effort-weighted records +900% WRONG # 40 × 2h = 80h against 8h
Standardised delivery index +150% ok # 10 × $80 = $800 against $320AI removes a step
AI deletes a whole step (the spec). if you count steps you lose output for doing the job better. had to go end to end.
case: step
setup: A feature used to take a 3-hour spec, a 10-hour build and a 2-hour review. Now an agent builds it from existing context for $25 of usage, and review takes 3 hours.
truth: The business gets the same feature for a quarter of the resources.
data:
Records closed 3 -> 2
Human hours 15 -> 3
Resources consumed $600 -> $145
results:
Items finished −33.3% WRONG # 3 records, then 2
Hours as output −80% WRONG # 15 hours, then 3
Effort-weighted records −20% WRONG # 10 + 2 = 12 standard hours against 15
Standardised delivery index no change ok # efficiency = 100 × 1 ÷ ($145 ÷ $600) = 413.8The same work, split into more tickets
ticket hygiene. same work, 1 ticket vs 10. every count-based thing explodes. per-ticket effort standards are worse than I expected because sub-task standards are small but there are loads of them.
case: split
setup: One feature, the same work, logged once as a single ticket and once as ten sub-tasks.
truth: One feature, either way.
data:
Tickets 1 -> 10
Hours 15 -> 15
results:
Items finished +900% WRONG # 1 ticket, then 10
Hours as output no change ok # 15 hours both times
Effort-weighted records +100% WRONG # 10 × 3h = 30h against 1 × 15h
Standardised delivery index no change ok # 1 feature × $600 both timesWork that has to be redone
agent closes stuff that comes back. does the measure count the redo as extra output? needed acceptance to mean 'held up', not 'closed'.
case: rework
setup: People used to close 100 cases with none coming back. Now an agent closes 100, and 25 come back and get redone by a person.
truth: 100 customers helped, with extra cost for the rework.
data:
Closures recorded 100 -> 125
Cases reopened and redone 0 -> 25
results:
Items finished +25% WRONG # 100 closures, then 125
Hours as output −75% WRONG # 30 people-hours, then 7.5
Effort-weighted records +25% WRONG # 125 × 0.3h against 100 × 0.3h
Standardised delivery index no change ok # 100 accepted cases both timesMore AI spend, the same work
same output, more AI spend. people coasting with expensive tools. only something with a cost side catches it, and even then it's a reason to look, not a verdict.
case: cruiser
setup: A six-person team costing $54,000 a month raises its AI spend from $2,800 to $4,600. It delivers the same work.
truth: Same output, more cost. Worth investigating, not a verdict on anyone.
data:
Payroll $54,000 -> $54,000
AI spend $2,800 -> $4,600
Work delivered same -> same
results:
Items finished no change meh # same records both months
Hours as output no change meh # same hours both months
Effort-weighted records no change meh # same standard hours both months
Standardised delivery index no change ok # 100 × 1 ÷ ($58,600 ÷ $56,800) = 96.9One weight is wrong and the mix moves
the one I got wrong. I assumed a wrong weight cancels out over time. it doesn't when the mix moves. the index fails this too, nothing I've got fixes it except checking type-level results and trying other weights.
case: mix
setup: Simple work's correct weight is $1 and complex work's $10, by stipulation. Quarter 1 is 90 simple and 1 complex. Quarter 2 is 10 simple and 9 complex. The complex weight was set at $20 by mistake.
truth: $100 of correctly weighted output both quarters. No growth.
data:
Simple units 90 -> 10
Complex units 1 -> 9
Correctly weighted output $100 -> $100
results:
Items finished −79.1% WRONG # 91 units, then 19
Hours as output no change meh # same hours both quarters
Effort-weighted records +72.7% WRONG # $110, then $190
Standardised delivery index +72.7% WRONG # $110, then $190The records get more complete
agents log everything, people don't. does better logging look like more work? yes, for every measure that reads the records. index fails too unless you measure output coverage separately.
case: coverage
setup: True output grows 10%. The records capture 80% of the output in period 1 and 84% in period 2, because the agent logs everything it does.
truth: 10% growth.
data:
True output 100 -> 110
Share recorded 80% -> 84%
Recorded output 80 -> 92.4
results:
Items finished +15.5% WRONG # 1.10 × 0.84 ÷ 0.80 − 1 = +15.5%
Hours as output no change WRONG # same payroll hours
Effort-weighted records +15.5% WRONG # 1.10 × 0.84 ÷ 0.80 − 1 = +15.5%
Standardised delivery index +15.5% WRONG # 1.10 × 0.84 ÷ 0.80 − 1 = +15.5%One team becomes two
split a team in two and they start raising tickets at each other. company output shouldn't move. only works if you count at the company boundary.
case: reorg
setup: A team handling 500 requests a quarter splits in two. The halves now pass 50 pieces of work to each other through tickets.
truth: The company delivers the same work.
data:
Customer requests 500 -> 500
Internal hand-off tickets 0 -> 50
results:
Items finished +10% WRONG # 500 records, then 550
Hours as output no change ok # same hours
Effort-weighted records +2.5% WRONG # 500 × 2h + 50 × 0.5h against 500 × 2h
Standardised delivery index no change ok # 500 requests both quartersEngineering fixes the cause of the cases
engineering fixes the bug, support gets fewer cases. every count says support did less. needs a service-level unit or outcomes alongside. index only half gets it.
case: prevention
setup: Support handles 5,000 contacts a quarter. Engineering fixes the defect behind 1,000 of them.
truth: Customers are better served with less work. Output measured as cases falls, and that's fine.
data:
Contacts handled 5,000 -> 4,000
results:
Items finished −20% WRONG # 5,000, then 4,000
Hours as output −20% WRONG # hours fall with the contacts
Effort-weighted records −20% WRONG # standard hours fall with the contacts
Standardised delivery index −20% meh # 5,000 cases, then 4,000Work nobody could afford before
audit goes from sampling to testing everything. valued at human cost it's a 100x 'gain' nobody would ever have paid for. ended up as a separate ledger until it's comparable.
case: capability
setup: An audit team used to test a 1% sample, 100 transactions. With AI it tests all 10,000.
truth: A new capability. Its value isn't proportional to the number of checks.
data:
Transactions tested 100 -> 10,000
results:
Items finished +9,900% WRONG # 100, then 10,000
Hours as output −50% WRONG # illustrative: human hours halve
Effort-weighted records +9,900% WRONG # 10,000 × the old manual standard
Standardised delivery index ledger meh # reported separately
Data quality notes
Engineering project tools
From years of running engineering teams, not a study. Linear and Jira are almost always behind the work. Tickets get created after the fact, or never. One piece of work gets split into whatever shape suits the sprint board. Types are wrong half the time because nobody changes the default. Time fields stay empty. Work gets closed when the sprint ends, not when it’s done. Links between tickets, PRs and incidents exist when someone was feeling diligent. Our own Linear at Flowstate is a decent example of all of the above. Anything that reads these systems as ground truth is measuring how tidy a team is.
Everything else
- process mining people rate ERP and CRM logs as “trustworthy but not necessarily complete” and rate document and work management systems lower.9 they list 27 classes of log quality problem10
- OCEL 2.0 exists because one-to-one relationships between records and business objects are the exception.11 one incident = a support convo + an issue + 3 PRs + an email. that’s not six outputs
- IT service desks are no better. TIGTA found 88% of the IRS’s top-priority incident tickets came from a monitoring process that should have been switched off.12
- LLMs as labellers: GPT-4 median accuracy about 0.85 across 27 labelling tasks, and the authors say to validate on 250 to 1,250 samples.13 as judges they favour the first answer and longer answers, possibly their own (couldn’t be confirmed).14 agreement with humans on hard judgements can be awful, kappa 0.10 to 0.28 in one benchmark.15
- correcting classifier counts: Rogan–Gladen is the classic fix for a known sensitivity and specificity.16 nobody has a method robust to every kind of dataset shift.17
- privacy: Microsoft pulled user names out of Productivity Score after a backlash.18 Viva won’t report groups under five.19 ICO expects a DPIA and the least intrusive option.20
Finance and accounting notes
- standard hours: ACCA, “a common measure for combining heterogeneous (dissimilar) products”.21 the core idea I borrowed
- TDABC: Kaplan and Anderson. practical capacity is usually 80 to 85% of theoretical. unused capacity is “opportunities for savings or growth”, not savings. aim for “approximately right”.22
- ACCA on activity-based management: asks the awkward question of whether freed-up staff have actually been redeployed.23
- earned value: budgets at standard labour rates, rate changes sit in the cost variance, not in output.24
- chain-linking: BEA says chained-dollar components aren’t additive.25 the CPI manual covers linking with an overlap period and admits the data for it is seldom there.26
- national accounts weights: ESA 2010 builds volumes from quantities weighted by previous-year unit costs.27 ab Iorwerth: unit costs as weights, not social valuations, which “introduce a degree of subjectivity”.28
- hours aren’t recorded: UK Statistics Authority says direct hours measurement “is rarely possible”.29
- keep it out of statutory numbers: capitalisation and R&D relief all want actual cost by time actually spent. HMRC: “only that proportion of their staffing costs can qualify”.30 US regs say time “actually spent”.31 Deloitte: a token charge isn’t capitalisable just because it’s measurable.32 FASB’s ASU 2025-06 moves the internal-use software rules from 2028.33
- pay rises: a 4% pay rise with nothing else changing knocks nominal efficiency down to 96.2. need a constant-price view too
Stuff from other fields that shaped it
- healthcare RVUs: work = time, skill and effort, judgement, stress from risk.34 RUC time estimates off by 19.8% on average, no consistent direction.35 CMS applied a −2.5% efficiency adjustment for 2026 and called time assumptions for many services “very likely overinflated”.36 AI scribes pushed the share of higher billing codes up 7 to 12 points at six health systems (vendor data).37
- Goodhart via Strathern: “When a measure becomes a target, it ceases to be a good measure.”38 call centre agents hanging up on customers to get rest time.39 SSA disability determinations per workyear fell from 302.8 to 240.2 with no complexity weighting.40
- lines of code: Dijkstra calls them “lines spent”.41 Bill Atkinson’s −2000 lines.42
- story points: Ron Jeffries’ apology.43 SPACE: productivity can’t be reduced to one dimension.44
- the residue in the wild: USPS keyers left with harder images, 1.2 billion keyed by hand a year against 19 billion in 1997.4546 Intercom says the remaining human caseload “becomes harder by definition”.47 Brynjolfsson, Li and Raymond saw 15% more issues resolved per hour with an assistant that helped people rather than taking a queue.48
- AI measured in human time: METR time horizons,49 GDPval (1,320 tasks, 44 occupations, expert time × wage, cost advantage collapses once review and redo are in),50 Anthropic’s estimates (Claude ranks task length about as well as devs, 0.44 vs 0.50, but squashes the range).51
- cheap work: HBR “workslop”, 40% got some last month, nearly two hours each to deal with.52 CMI: 87% of marketers say AI made them more productive, 39% say content performs better.53
- planning: IMD, planner intervention makes accuracy worse on often 80%+ of items.54
- value vs uplift: METR’s task substitution piece, “Cadillac tasks” you only do because AI made them cheap.55
- macro sanity check: Acemoglu’s own estimate is no more than 0.66% TFP over ten years.56 his essay on pro-worker AI is where the capability ledger idea came from.57
Things I still don’t know
- can two competent teams apply the definitions to the same evidence and get the same answer?
- how often does the mix move towards types with dodgy weights in a real company? no idea yet
- can output coverage actually be measured, rather than spending coverage?
- will companies share enough context for any of this to work? (Flowstate problem, not a blog problem)
- what the right observation window is for “it held up” in each function
- whether any of this survives contact with a works council
Footnotes
-
ONS, Public service productivity estimates: sources and methods. Link ↩
-
METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, 10 July 2025. Link ↩
-
Humlum and Vestergaard, Large Language Models, Small Labor Market Effects, BFI working paper 2025-56. Link ↩
-
Bose, Mans and van der Aalst, Wanna improve process mining results?, 2013. ↩
-
TIGTA, Incident and Service Ticket Management Needs Improvement, report 2026-200-020. Link ↩
-
Pangakis, Wolken and Fasching, Automated annotation with generative AI requires validation. Link ↩
-
Zheng et al., Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. Link ↩
-
Bavaresco et al., LLMs instead of Human Judges? (JUDGE-BENCH). Link ↩
-
Rogan and Gladen, Estimating prevalence from the results of a screening test, 1978. Link ↩
-
Microsoft, Our commitment to privacy in Microsoft Productivity Score, December 2020. Link ↩
-
Microsoft Learn, Viva Insights privacy considerations. Link ↩
-
Kaplan and Anderson, Rethinking Activity-Based Costing, HBS Working Knowledge, 2005. Link ↩
-
US DOE, Earned Value Management System Interpretation Handbook. Link ↩
-
ILO and others, Consumer Price Index Manual, 2004, chapter 1. Link ↩
-
Aled ab Iorwerth, To Capture Production or Well-being?, International Productivity Monitor 23, 2012. Link ↩
-
UK Statistics Authority, National Statistician’s independent review of the measurement of public services productivity, March 2025, Annex F. Link ↩
-
Deloitte DART, Accounting for AI costs associated with internal-use software development. Link ↩
-
CPA Practice Advisor, FASB issues standard to improve internal-use software guidance. Link ↩
-
CMS, CY 2026 Physician Fee Schedule final rule fact sheet. Link ↩
-
Trilliant Health, Increased outpatient coding intensity following hospital adoption of AI-enabled scribing. Vendor analysis. Link ↩
-
Marilyn Strathern, ‘Improving ratings’, European Review, 1997, p. 308. Link ↩
-
Brown et al., Statistical analysis of a telephone call center, 2005. Link ↩
-
Forsgren et al., The SPACE of Developer Productivity, ACM Queue, 2021. Link ↩
-
NALC, The Remote Encoding Center, The Postal Record, July 2022. Link ↩
-
Intercom, How Fin AI Agent and Copilot Cut Handle Time. Link ↩
-
Brynjolfsson, Li and Raymond, Generative AI at Work, QJE 2025. Link ↩
-
METR, Measuring AI Ability to Complete Long Tasks, March 2025. Link ↩
-
Tamkin and McCrory, Estimating AI productivity gains from Claude conversations, Anthropic, November 2025. Link ↩
-
Niederhoffer et al., AI-Generated “Workslop” Is Destroying Productivity, HBR, September 2025. Link ↩
-
Content Marketing Institute, B2B content marketing trends research. Link ↩
-
Cunningham and Whitfill, Task Substitution and Uplift, METR, May 2026. Link ↩
-
Daron Acemoglu, The Simple Macroeconomics of AI, NBER w32487. Link ↩
-
Daron Acemoglu, Will AI Replace Workers? Not If We Build It Right., The Humanist Review, July 2026. Link ↩