← The impossibility of measuring AI productivity

Working notes · 7 October 2026

Workings: trying to measure work done by people and AI

These are my notes from working on the post. They’re not tidied up. Most of it was written while I was arguing with myself, so it repeats, changes its mind and occasionally contradicts the post, because the post is where I ended up and this is how I got there.

If you only want the argument, read the post. If you want to check my working, or steal the test cases, carry on.

What I was actually trying to do

  • customers keep asking “is AI making us more productive”. finance wants a number they can defend to a board
  • every number they have was built for a workforce of people. tickets, story points, hours, headcount
  • wanted ONE unit that works across functions. support, eng, finance, legal, marketing, whatever
  • had to be neutral: same work earns the same whether a person, a person + AI, or an agent did it
  • had to be computable from systems a company already has. no new timesheets. nobody wants to narrate their day to a machine
  • assumed (and still assume) LLMs + company context do the messy interpretation. the post leaves that out on purpose, these notes don’t

Pass criteria I set before testing anything

  1. one definition across functions, no custom model per team
  2. computable from existing systems, with AI filling gaps
  3. neutral between people and agents
  4. answers what leaders actually ask: faster? cheaper? fewer people? different mix? what do we need next?
  5. hard to game, at least against the known tricks in each function
  6. has precedent an accountant or auditor would recognise
  7. stable over time, including once pre-AI history runs out

Nothing passed all seven. The final version gets closest and still fails 5 and 7 in places (see the mix and coverage cases below).

The functions I tested against

I wanted functions that are genuinely different, not different job titles doing the same thing. Went through 18. Most teams turned out to run three shapes of work at once. Five shapes kept coming back: transactions, cases, cycles, projects, advice. Two more broke volume counts outright: ongoing duties (nothing to count when it goes well) and judgement calls (doing more can make it worse).

FunctionShapeWhere it’s recordedWhat arrivesWhat says it was goodWhere value shows upConstraints that bit me
Accounts payabletransactions + exceptionsERP, AP toolsinvoicespaid right, no duplicatediscounts keptmost invoices are touchless now, the people get the exceptions. residue again
Customer supportcasesZendesk, Intercomticketsnot reopenedretention, months latervendor “resolution” units count customers who just leave
Address keying (USPS)transactionsimage queueunreadable addressesdeliverednone reallythe original residue example, machines took the easy ones
Insurance claimscases, years longclaims systemsclaimsnot reopened, not overpaidloss ratio, years latercomplexity is assessed at intake, not at the end
Software engineeringprojects + ticketsJira, Linear, GitHubplan items, bugsmerged, not revertedadoptiontickets lag the work, split arbitrarily, closed late. see below
SalesdealsCRMpipelinewon or lostbookingsonly function with money on the record. timing games around commission
SDRs and AI SDRsactivityoutreach toolsleadsmeeting heldpipelineactivity is free now, meetings aren’t
Demand planningcycles + judgementplanning toolsmonthly cycleforecast value addedstock and serviceplanner changes often make it worse
FP&Acycles + adviceExcel, email, planning toolsclose, ad hoc asksused in a decisionalmost never recordedmost of the work is a question in Slack
Legalcases + advicecontract tools, matter systemscontracts, questionssigned, no disputesomething bad that didn’t happengood advice can kill a deal
Brand marketingduty + creativeAsana, Workfrontplanran, reached peopleyears latercheap drafts, see the drafts case
Product managementjudgementroadmap toolsplanshipped and adoptedmonths lateroutput is decisions
SRE and on-calldutyPagerDutyrota periodsSLO metincidents that didn’t happena quiet week is a good week
Internal auditprojects, assuranceaudit toolsaudit planreport issued, findings closedcounterfactualAI lets you test everything, see capability case
RecruitingprojectsATSrequisitionshire still in postquality of hirequality of hire is rarely measured
Clinicians with AI scribescasesEHRencounterscoded, not deniedoutcomesscribes push billing codes up
Team leadsoverheadnothingnothingthe team’s resultsthe team’s resultsno record at all
Pharma R&Dstage gatesnotebooks, portfolio toolsportfolio plangate passeda decade latera failed gate can be good work

Every one had some record with a created date, closed date, owner, status and type. That’s the only data shape you can count on almost everywhere. Agents are starting to show up in the owner field. Only sales had money on the record.

Edge cases list (the long version)

  • residue: machines take the easy work, people look slower
  • the cruiser: uses AI, delivers the same, costs more
  • the slow person: 4 hours on a 1-hour investor update. the fast one: 20 minutes
  • new starters, no history
  • new types of work, no baseline
  • cheap-work inflation: 40 ad variants instead of 4
  • brand-new capability: testing 100% of transactions
  • judgement work where doing more makes it worse
  • duties and prevention: good work produces nothing countable
  • step elimination: AI removes the spec step entirely
  • agents opening tickets for other agents
  • splitting and reclassifying (DRG creep is the famous one)
  • AI inflating the evidence (scribes and billing codes)
  • finished but wrong
  • lag: value turns up 2 years later
  • unlogged work
  • roll-up: team to company without double counting
  • ageing baselines: “pre-AI” stops meaning anything

The rounds

Roughly ten rounds of propose a measure, throw the cases at it, write down what broke, keep the useful bit.

  1. count finished items. broke on residue straight away (−40% tickets per hour while the system got better). splitting inflates it. ignores size. kept: count at completion
  2. inputs (hours, headcount, cost). AI gains invisible by definition. the ONS does this for about a third of public service output and says outright that productivity is constant there.1 kept: the cost side
  3. money (revenue, throughput). only sales has a price. attribution is a swamp. Google gave up on Shapley-style attribution in GA4 and went to experiment-trained models, which says a lot.2 kept: value shown alongside, never instead
  4. time saved, self-reported. METR: experienced devs 19% slower, believed 20% faster.3 Danish chatbot users: 2.8% of hours saved, nothing in earnings or recorded hours.4 kept: AI time estimates as a starting point for weights, nothing more
  5. vendor outcome units. every vendor bills its own unit. Fin counts an assumed resolution when the customer leaves (and reverses it if they come back).5 Zendesk split resolutions into tiers and only bills verified ones.6 Devin bills “agent compute units”. kept: vendor counts as evidence, normalised
  6. unsized flow items (flow velocity etc). splitting inflates it, residue still breaks it.7 kept: lead time and WIP alongside output
  7. standard effort × impact weighting. this was the first spec I actually liked. handled residue, cruiser, slow person, new starter. then: drafts 10x’d output, planners got paid for changes that made forecasts worse, duties earned nothing, step elimination showed as lost output, impact weights were pure opinion. Medicare’s RVU committee is the cautionary tale for opinion weights: CMS agreed with 87.4% of its recommendations.8 kept: standard effort per type
  8. standard effort + demand gate + acceptance gate, cost weighted. demand gate = “only counts if someone outside the team asked for it or it’s in an approved plan”. felt clever. step elimination still undercounted by 20%. prevention, agent-made demand, splitting, new capability all broke it. kept: acceptance, cost weighting
  9. round 8 + count at the request level, end-to-end standards, service periods for duties, a new-capability ledger, chain-linking. nothing in my list broke it. which mostly told me my list was too kind
  10. review round. got torn apart, mostly fairly. the big ones:
    • historic cost ISN’T a floor on value. companies buy rubbish
    • capacity released isn’t cash
    • weight errors don’t cancel when the mix moves (I’d claimed they did. they don’t, see mix case)
    • parent/child double counting survives “count at request level”
    • the demand gate is a permission system that doesn’t exist. engineers fix things nobody asked for and that’s often the best work they do. dropped it for “defensible purpose, scope and completion condition”
    • spending coverage ≠ output coverage. I’d used the share of payroll I could see as if it were the share of output I could see. wrong
    • my Monte Carlo error bands looked lovely and proved nothing about a real company. pulled them from the post

The test cases

Each of these is invented, with dummy numbers picked to make the mechanism obvious. The structures below are generated from research/unit-of-work/trials.py, so they can’t drift from the maths. “WRONG” means the measure would mislead you, “meh” means half right, “ok” means it tells you what actually happened.

The agent takes the easy work

the one that started all this. does the measure blame the people when the machine takes the easy stuff? anything that counts items per hour fails. weighting by type fixes it.

case: residue
setup: A support team resolves 800 easy cases (6 minutes each) and 200 difficult ones (30 minutes) a month. An agent takes 600 of the easy ones at $0.80 each.
truth: Every case still gets resolved. Nobody got slower. The people need fewer hours.
data:
  Cases resolved by people              1000 -> 400
  Cases resolved by the agent              0 -> 600
  Human handling hours                   180 -> 120
  Allocated handling cost             $7,200 -> $5,280
results:
  Items finished               −40%       WRONG  # per person-hour: 1000 ÷ 180 = 5.56, then 400 ÷ 120 = 3.33
  Hours as output              −33.3%     WRONG  # 180 hours, then 120
  Effort-weighted records      no change  ok     # 120 people-hours of standard + 600 × 6 min = 180 standard hours, as before
  Standardised delivery index  no change  ok     # 800 × $4 + 200 × $20 = $7,200 in both months

AI makes drafts nearly free

what happens when AI makes something nearly free and people just make loads of it. first version of this test assumed a manager-approved plan of 10. dropped that, nobody works like that. the real question is what the deliverable is: a draft or the thing that gets used.

case: drafts
setup: A marketing team used to make 4 ad variants a week, 2 hours each. With AI it generates 40, and 10 go into independently defined tests.
truth: 10 things were delivered. The other 30 are drafts, part of the cost of making them.
data:
  Variants generated                       4 -> 40
  Variants used in tests                   4 -> 10
  Reference weight each                  $80 -> $80
results:
  Items finished               +900%      WRONG  # 4 files, then 40
  Hours as output              −25%       WRONG  # 8 hours a week, then 6 (illustrative)
  Effort-weighted records      +900%      WRONG  # 40 × 2h = 80h against 8h
  Standardised delivery index  +150%      ok     # 10 × $80 = $800 against $320

AI removes a step

AI deletes a whole step (the spec). if you count steps you lose output for doing the job better. had to go end to end.

case: step
setup: A feature used to take a 3-hour spec, a 10-hour build and a 2-hour review. Now an agent builds it from existing context for $25 of usage, and review takes 3 hours.
truth: The business gets the same feature for a quarter of the resources.
data:
  Records closed                           3 -> 2
  Human hours                             15 -> 3
  Resources consumed                    $600 -> $145
results:
  Items finished               −33.3%     WRONG  # 3 records, then 2
  Hours as output              −80%       WRONG  # 15 hours, then 3
  Effort-weighted records      −20%       WRONG  # 10 + 2 = 12 standard hours against 15
  Standardised delivery index  no change  ok     # efficiency = 100 × 1 ÷ ($145 ÷ $600) = 413.8

The same work, split into more tickets

ticket hygiene. same work, 1 ticket vs 10. every count-based thing explodes. per-ticket effort standards are worse than I expected because sub-task standards are small but there are loads of them.

case: split
setup: One feature, the same work, logged once as a single ticket and once as ten sub-tasks.
truth: One feature, either way.
data:
  Tickets                                  1 -> 10
  Hours                                   15 -> 15
results:
  Items finished               +900%      WRONG  # 1 ticket, then 10
  Hours as output              no change  ok     # 15 hours both times
  Effort-weighted records      +100%      WRONG  # 10 × 3h = 30h against 1 × 15h
  Standardised delivery index  no change  ok     # 1 feature × $600 both times

Work that has to be redone

agent closes stuff that comes back. does the measure count the redo as extra output? needed acceptance to mean 'held up', not 'closed'.

case: rework
setup: People used to close 100 cases with none coming back. Now an agent closes 100, and 25 come back and get redone by a person.
truth: 100 customers helped, with extra cost for the rework.
data:
  Closures recorded                      100 -> 125
  Cases reopened and redone                0 -> 25
results:
  Items finished               +25%       WRONG  # 100 closures, then 125
  Hours as output              −75%       WRONG  # 30 people-hours, then 7.5
  Effort-weighted records      +25%       WRONG  # 125 × 0.3h against 100 × 0.3h
  Standardised delivery index  no change  ok     # 100 accepted cases both times

More AI spend, the same work

same output, more AI spend. people coasting with expensive tools. only something with a cost side catches it, and even then it's a reason to look, not a verdict.

case: cruiser
setup: A six-person team costing $54,000 a month raises its AI spend from $2,800 to $4,600. It delivers the same work.
truth: Same output, more cost. Worth investigating, not a verdict on anyone.
data:
  Payroll                            $54,000 -> $54,000
  AI spend                            $2,800 -> $4,600
  Work delivered                        same -> same
results:
  Items finished               no change  meh    # same records both months
  Hours as output              no change  meh    # same hours both months
  Effort-weighted records      no change  meh    # same standard hours both months
  Standardised delivery index  no change  ok     # 100 × 1 ÷ ($58,600 ÷ $56,800) = 96.9

One weight is wrong and the mix moves

the one I got wrong. I assumed a wrong weight cancels out over time. it doesn't when the mix moves. the index fails this too, nothing I've got fixes it except checking type-level results and trying other weights.

case: mix
setup: Simple work's correct weight is $1 and complex work's $10, by stipulation. Quarter 1 is 90 simple and 1 complex. Quarter 2 is 10 simple and 9 complex. The complex weight was set at $20 by mistake.
truth: $100 of correctly weighted output both quarters. No growth.
data:
  Simple units                            90 -> 10
  Complex units                            1 -> 9
  Correctly weighted output             $100 -> $100
results:
  Items finished               −79.1%     WRONG  # 91 units, then 19
  Hours as output              no change  meh    # same hours both quarters
  Effort-weighted records      +72.7%     WRONG  # $110, then $190
  Standardised delivery index  +72.7%     WRONG  # $110, then $190

The records get more complete

agents log everything, people don't. does better logging look like more work? yes, for every measure that reads the records. index fails too unless you measure output coverage separately.

case: coverage
setup: True output grows 10%. The records capture 80% of the output in period 1 and 84% in period 2, because the agent logs everything it does.
truth: 10% growth.
data:
  True output                            100 -> 110
  Share recorded                         80% -> 84%
  Recorded output                         80 -> 92.4
results:
  Items finished               +15.5%     WRONG  # 1.10 × 0.84 ÷ 0.80 − 1 = +15.5%
  Hours as output              no change  WRONG  # same payroll hours
  Effort-weighted records      +15.5%     WRONG  # 1.10 × 0.84 ÷ 0.80 − 1 = +15.5%
  Standardised delivery index  +15.5%     WRONG  # 1.10 × 0.84 ÷ 0.80 − 1 = +15.5%

One team becomes two

split a team in two and they start raising tickets at each other. company output shouldn't move. only works if you count at the company boundary.

case: reorg
setup: A team handling 500 requests a quarter splits in two. The halves now pass 50 pieces of work to each other through tickets.
truth: The company delivers the same work.
data:
  Customer requests                      500 -> 500
  Internal hand-off tickets                0 -> 50
results:
  Items finished               +10%       WRONG  # 500 records, then 550
  Hours as output              no change  ok     # same hours
  Effort-weighted records      +2.5%      WRONG  # 500 × 2h + 50 × 0.5h against 500 × 2h
  Standardised delivery index  no change  ok     # 500 requests both quarters

Engineering fixes the cause of the cases

engineering fixes the bug, support gets fewer cases. every count says support did less. needs a service-level unit or outcomes alongside. index only half gets it.

case: prevention
setup: Support handles 5,000 contacts a quarter. Engineering fixes the defect behind 1,000 of them.
truth: Customers are better served with less work. Output measured as cases falls, and that's fine.
data:
  Contacts handled                     5,000 -> 4,000
results:
  Items finished               −20%       WRONG  # 5,000, then 4,000
  Hours as output              −20%       WRONG  # hours fall with the contacts
  Effort-weighted records      −20%       WRONG  # standard hours fall with the contacts
  Standardised delivery index  −20%       meh    # 5,000 cases, then 4,000

Work nobody could afford before

audit goes from sampling to testing everything. valued at human cost it's a 100x 'gain' nobody would ever have paid for. ended up as a separate ledger until it's comparable.

case: capability
setup: An audit team used to test a 1% sample, 100 transactions. With AI it tests all 10,000.
truth: A new capability. Its value isn't proportional to the number of checks.
data:
  Transactions tested                    100 -> 10,000
results:
  Items finished               +9,900%    WRONG  # 100, then 10,000
  Hours as output              −50%       WRONG  # illustrative: human hours halve
  Effort-weighted records      +9,900%    WRONG  # 10,000 × the old manual standard
  Standardised delivery index  ledger     meh    # reported separately

Data quality notes

Engineering project tools

From years of running engineering teams, not a study. Linear and Jira are almost always behind the work. Tickets get created after the fact, or never. One piece of work gets split into whatever shape suits the sprint board. Types are wrong half the time because nobody changes the default. Time fields stay empty. Work gets closed when the sprint ends, not when it’s done. Links between tickets, PRs and incidents exist when someone was feeling diligent. Our own Linear at Flowstate is a decent example of all of the above. Anything that reads these systems as ground truth is measuring how tidy a team is.

Everything else

  • process mining people rate ERP and CRM logs as “trustworthy but not necessarily complete” and rate document and work management systems lower.9 they list 27 classes of log quality problem10
  • OCEL 2.0 exists because one-to-one relationships between records and business objects are the exception.11 one incident = a support convo + an issue + 3 PRs + an email. that’s not six outputs
  • IT service desks are no better. TIGTA found 88% of the IRS’s top-priority incident tickets came from a monitoring process that should have been switched off.12
  • LLMs as labellers: GPT-4 median accuracy about 0.85 across 27 labelling tasks, and the authors say to validate on 250 to 1,250 samples.13 as judges they favour the first answer and longer answers, possibly their own (couldn’t be confirmed).14 agreement with humans on hard judgements can be awful, kappa 0.10 to 0.28 in one benchmark.15
  • correcting classifier counts: Rogan–Gladen is the classic fix for a known sensitivity and specificity.16 nobody has a method robust to every kind of dataset shift.17
  • privacy: Microsoft pulled user names out of Productivity Score after a backlash.18 Viva won’t report groups under five.19 ICO expects a DPIA and the least intrusive option.20

Finance and accounting notes

  • standard hours: ACCA, “a common measure for combining heterogeneous (dissimilar) products”.21 the core idea I borrowed
  • TDABC: Kaplan and Anderson. practical capacity is usually 80 to 85% of theoretical. unused capacity is “opportunities for savings or growth”, not savings. aim for “approximately right”.22
  • ACCA on activity-based management: asks the awkward question of whether freed-up staff have actually been redeployed.23
  • earned value: budgets at standard labour rates, rate changes sit in the cost variance, not in output.24
  • chain-linking: BEA says chained-dollar components aren’t additive.25 the CPI manual covers linking with an overlap period and admits the data for it is seldom there.26
  • national accounts weights: ESA 2010 builds volumes from quantities weighted by previous-year unit costs.27 ab Iorwerth: unit costs as weights, not social valuations, which “introduce a degree of subjectivity”.28
  • hours aren’t recorded: UK Statistics Authority says direct hours measurement “is rarely possible”.29
  • keep it out of statutory numbers: capitalisation and R&D relief all want actual cost by time actually spent. HMRC: “only that proportion of their staffing costs can qualify”.30 US regs say time “actually spent”.31 Deloitte: a token charge isn’t capitalisable just because it’s measurable.32 FASB’s ASU 2025-06 moves the internal-use software rules from 2028.33
  • pay rises: a 4% pay rise with nothing else changing knocks nominal efficiency down to 96.2. need a constant-price view too

Stuff from other fields that shaped it

  • healthcare RVUs: work = time, skill and effort, judgement, stress from risk.34 RUC time estimates off by 19.8% on average, no consistent direction.35 CMS applied a −2.5% efficiency adjustment for 2026 and called time assumptions for many services “very likely overinflated”.36 AI scribes pushed the share of higher billing codes up 7 to 12 points at six health systems (vendor data).37
  • Goodhart via Strathern: “When a measure becomes a target, it ceases to be a good measure.”38 call centre agents hanging up on customers to get rest time.39 SSA disability determinations per workyear fell from 302.8 to 240.2 with no complexity weighting.40
  • lines of code: Dijkstra calls them “lines spent”.41 Bill Atkinson’s −2000 lines.42
  • story points: Ron Jeffries’ apology.43 SPACE: productivity can’t be reduced to one dimension.44
  • the residue in the wild: USPS keyers left with harder images, 1.2 billion keyed by hand a year against 19 billion in 1997.4546 Intercom says the remaining human caseload “becomes harder by definition”.47 Brynjolfsson, Li and Raymond saw 15% more issues resolved per hour with an assistant that helped people rather than taking a queue.48
  • AI measured in human time: METR time horizons,49 GDPval (1,320 tasks, 44 occupations, expert time × wage, cost advantage collapses once review and redo are in),50 Anthropic’s estimates (Claude ranks task length about as well as devs, 0.44 vs 0.50, but squashes the range).51
  • cheap work: HBR “workslop”, 40% got some last month, nearly two hours each to deal with.52 CMI: 87% of marketers say AI made them more productive, 39% say content performs better.53
  • planning: IMD, planner intervention makes accuracy worse on often 80%+ of items.54
  • value vs uplift: METR’s task substitution piece, “Cadillac tasks” you only do because AI made them cheap.55
  • macro sanity check: Acemoglu’s own estimate is no more than 0.66% TFP over ten years.56 his essay on pro-worker AI is where the capability ledger idea came from.57

Things I still don’t know

  • can two competent teams apply the definitions to the same evidence and get the same answer?
  • how often does the mix move towards types with dodgy weights in a real company? no idea yet
  • can output coverage actually be measured, rather than spending coverage?
  • will companies share enough context for any of this to work? (Flowstate problem, not a blog problem)
  • what the right observation window is for “it held up” in each function
  • whether any of this survives contact with a works council

Footnotes

  1. ONS, Public service productivity estimates: sources and methods. Link ↩

  2. Google Analytics Help, attribution models in GA4. Link ↩

  3. METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, 10 July 2025. Link ↩

  4. Humlum and Vestergaard, Large Language Models, Small Labor Market Effects, BFI working paper 2025-56. Link ↩

  5. Intercom, Fin AI Agent outcomes. Link ↩

  6. Zendesk, About automated resolution tiers. Link ↩

  7. Planview, Flow metrics guide. Link ↩

  8. Laugesen, Wada and Chen, Health Affairs, May 2012. Link ↩

  9. van der Aalst et al., Process Mining Manifesto. Link ↩

  10. Bose, Mans and van der Aalst, Wanna improve process mining results?, 2013. ↩

  11. Berti et al., OCEL 2.0 Specification. Link ↩

  12. TIGTA, Incident and Service Ticket Management Needs Improvement, report 2026-200-020. Link ↩

  13. Pangakis, Wolken and Fasching, Automated annotation with generative AI requires validation. Link ↩

  14. Zheng et al., Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. Link ↩

  15. Bavaresco et al., LLMs instead of Human Judges? (JUDGE-BENCH). Link ↩

  16. Rogan and Gladen, Estimating prevalence from the results of a screening test, 1978. Link ↩

  17. González et al., quantification under dataset shift. Link ↩

  18. Microsoft, Our commitment to privacy in Microsoft Productivity Score, December 2020. Link ↩

  19. Microsoft Learn, Viva Insights privacy considerations. Link ↩

  20. ICO, Monitoring workers. Link ↩

  21. ACCA, The standard hour. Link ↩

  22. Kaplan and Anderson, Rethinking Activity-Based Costing, HBS Working Knowledge, 2005. Link ↩

  23. ACCA, Activity-based management. Link ↩

  24. US DOE, Earned Value Management System Interpretation Handbook. Link ↩

  25. BEA, Chained-dollar estimates. Link ↩

  26. ILO and others, Consumer Price Index Manual, 2004, chapter 1. Link ↩

  27. Eurostat, Guidance on non-market output. Link ↩

  28. Aled ab Iorwerth, To Capture Production or Well-being?, International Productivity Monitor 23, 2012. Link ↩

  29. UK Statistics Authority, National Statistician’s independent review of the measurement of public services productivity, March 2025, Annex F. Link ↩

  30. HMRC, CIRD133000. Link ↩

  31. 26 CFR §1.41-2. Link ↩

  32. Deloitte DART, Accounting for AI costs associated with internal-use software development. Link ↩

  33. CPA Practice Advisor, FASB issues standard to improve internal-use software guidance. Link ↩

  34. AMA, RBRVS overview. Link ↩

  35. Chan, Huynh and Studdert, NEJM, 2019. Link ↩

  36. CMS, CY 2026 Physician Fee Schedule final rule fact sheet. Link ↩

  37. Trilliant Health, Increased outpatient coding intensity following hospital adoption of AI-enabled scribing. Vendor analysis. Link ↩

  38. Marilyn Strathern, ‘Improving ratings’, European Review, 1997, p. 308. Link ↩

  39. Brown et al., Statistical analysis of a telephone call center, 2005. Link ↩

  40. SSA OIG report, July 2025. Link ↩

  41. Dijkstra, EWD1036. Link ↩

  42. Andy Hertzfeld, -2000 Lines Of Code. Link ↩

  43. Ron Jeffries, Story Points Revisited, 2019. Link ↩

  44. Forsgren et al., The SPACE of Developer Productivity, ACM Queue, 2021. Link ↩

  45. NALC, The Remote Encoding Center, The Postal Record, July 2022. Link ↩

  46. Intercom, How Fin AI Agent and Copilot Cut Handle Time. Link ↩

  47. Brynjolfsson, Li and Raymond, Generative AI at Work, QJE 2025. Link ↩

  48. METR, Measuring AI Ability to Complete Long Tasks, March 2025. Link ↩

  49. Patwardhan et al., GDPval, 2025. Link ↩

  50. Tamkin and McCrory, Estimating AI productivity gains from Claude conversations, Anthropic, November 2025. Link ↩

  51. Niederhoffer et al., AI-Generated “Workslop” Is Destroying Productivity, HBR, September 2025. Link ↩

  52. Content Marketing Institute, B2B content marketing trends research. Link ↩

  53. Seifert, Markoff and Spooner, I by IMD, August 2024. Link ↩

  54. Cunningham and Whitfill, Task Substitution and Uplift, METR, May 2026. Link ↩

  55. Daron Acemoglu, The Simple Macroeconomics of AI, NBER w32487. Link ↩

  56. Daron Acemoglu, Will AI Replace Workers? Not If We Build It Right., The Humanist Review, July 2026. Link ↩