The randomized controlled trial corpus on AI productivity is small, contradictory at first glance, and easy to misuse. The contradictions resolve when you account for task type, developer experience, model generation, and workflow integration. Below: the studies that matter, the era each tested, and how to read them together.


The Core RCTs (in date order)

Study Date n Population Model Era Headline Tier
Brynjolfsson, Li, Raymond (NBER → QJE) 2023 / pub 2025 5,172 Customer support agents GPT-3-era +15% avg (+34% novices) TIER 4
Cui et al. (Peng et al., GitHub) 2023 95 Developers, greenfield JS GPT-3.5 +55.8% task completion TIER 4
Dell’Acqua et al. (BCG/HBS) 2023 758 Consultants GPT-4 +12% on-frontier tasks; −19 pts off-frontier TIER 4
Google internal Oct 2024 96 Google engineers GPT-4-class +21% (CI includes near-zero) TIER 3
HBS Cybernetic Teammate (Sadun) 2025 776 P&G professionals GPT-4-class AI-individual ≈ human-team TIER 2
METR Jul 2025 16 / 246 tasks Experienced devs, mature OSS Claude 3.5/3.7, GPT-4o −19% (perceived +20%) TIER 2
AISI UK Government Feb 2026 500 Mixed workforce, 4 O*NET task types Early-2025 frontier LLM +25% quality; +102% PPM on info-interpretation; zero on planning TIER 1
Ju & Aral (MIT/Johns Hopkins) Feb 2026 2,234 Creative teams (advertising) GPT-4o Aug 2024 +50% output; text quality ↑; image quality ↓; diversity collapse; zero real-world CTR advantage TIER 1
Cui, Demirer, Jaffe et al. (Management Science) Exp 2022–2023, pub 2025 4,867 Enterprise developers (Microsoft, Accenture, Fortune 100) GitHub Copilot 2022–2023 +26.08% completed tasks; junior/short-tenure 21–39%; senior ~8%; no quality decline TIER 3 data, HIGH methodology
Suh & Oh (Bank of Korea) Feb 2026 5,512 Korean workers nationally representative GenAI (2025) 3.8% time savings; 0.008 correlation with output; gains captured as on-the-job leisure TIER 1
Dillon, Jaffe, Immorlica & Stanton (Microsoft Research / HBS) Apr 2025 / rev Nov 2025 7,137 Knowledge workers, 66 large enterprises Microsoft 365 Copilot (2023-era) −2 hrs/week email (−17%); no task composition shift; firm management explains adoption 2× more than individual behavior TIER 2
Daniotti, Wachs, Feng & Neffke (Science) Feb 2026 160,000 devs / 30M commits Developers globally (GitHub), 2019–2024 GitHub Copilot / AI coding tools (2022–2024) +3.6% output; experienced devs capture all gains; beginners: no significant benefit TIER 1
Ni, Wang, Feng, Lu et al. (Fudan/Zhejiang/Dartmouth Tuck/Alibaba) Nov 2025 5,940 agents / 2.56M chats Taobao e-commerce customer service Alibaba Qwen fine-tuned (2024) −8.2% ID time, +0.042 rating (avg); bottom Q1: −14.8% duration, +2.144 rating; top Q5: −1.514 rating, +4.1pp retrial rate — skill inversion TIER 1
Sun, Li, Foo, Zhou & Lu (Journal of Applied Psychology) Jun/Dec 2025 250 Knowledge workers, technology consulting firm ChatGPT (general LLM) Creativity gains for high-metacognition employees; zero gain for low-metacognition employees; dual-rating design (supervisor + external) TIER 1
Otis, Clarke, Delecourt, Holtz, Koning (MIT Sloan MR) Apr 2026 640 Small business owners, Kenya (food/bev/agriculture/car-wash) GPT-4 via WhatsApp Top 50%: +15% revenue/profit; Bottom 50%: −10% revenue/profit; Average: ~0% (not sig.) — judgment amplification TIER 1
Humlum & Vestergaard (NBER WP 33777) May 2025, rev Mar 2026 25,000 workers / 7,000 workplaces 11 AI-exposed occupations, Denmark (linked to admin payroll) All chatbots incl. ChatGPT (2022–2024) Null earnings/hours effect (rules out >2%); 93% adoption in high-investment workplaces; occupational switchers +12 pp earnings growth TIER 1
Bono (Microsoft) — Phishing Triage Agent Nov 2025 167 Professional security analysts MS Security Copilot 6.5× true positives/analyst-min (ground truth); 3.1× pessimistic; +77% F1 TIER 1 (vendor-authored)
Bono, Cheng, Lozano (Microsoft) — Conditional Access Agent Nov 2025 162 Identity administrators MS Security Copilot +48% accuracy; −43% task time; +204% on Zero-Trust gap detection TIER 1 (vendor-authored; simulated env.)
Lång, Gommers et al. (MASAI) — Mammography Screening RCT Jan 2026 105,934 Women ages 40–74, Sweden national program Transpara Detection −12% interval cancers; 73.8%→80.5% sensitivity; −27% aggressive cancers; −44% workload; 0 false positive increase HIGH / TIER 1
Tabachnyk et al. (Google) — Transform Code DiD Jan 2026 54,000 Google engineers (36k treatment, 18k control) Gemini-based internal IDE (2024 data) +17.5% Change List Throughput (CI: [15.9%, 19.0%]); −3.6% investigation session time; ACT/CL not significant TIER 1 (vendor; rigorous DiD)
Holmgren et al. (UCSF/JAMA) — AI Scribe Clinical Cohort Jan 2026 1,202,734 encounters / 1,565 physicians Ambulatory physicians, UCSF Health Ambient AI scribe (clinical documentation) +1.81 RVU/week (P<.001); +0.80 encounters/week (P=.04); ~$3,044 annual revenue/physician; 0 denial increase HIGH / TIER 1 (JAMA Network Open; DiD design; single-site)
Afshar et al. — Ambient AI Scribe Burnout Stepped-Wedge RCT Nov 2025 66 Health care practitioners, ambulatory, 2 states Ambient AI scribe (generative) ~30 min/day documentation reduction; clinically meaningful burnout reduction (Stanford PFI); billing accuracy improved HIGH / TIER 1 (NEJM AI; stepped-wedge RCT)
Lukac et al. — Ambient AI Scribe 3-Group RCT Nov 2025 238 Outpatient physicians, 14 specialties DAX Copilot vs. Nabla vs. control Nabla: −9.5% log writing time (P=0.02); DAX: −1.7% (P=0.66, NS); both improved Mini-Z burnout vs. control HIGH / TIER 1 (NEJM AI; 3-group parallel RCT)
Shen & Tamkin (Anthropic) — AI Coding Skill Formation Feb 2026 52 Junior developers, Python Trio library (unfamiliar) GPT-4o −17pp comprehension quiz (50% vs. 67%, p=0.010, Cohen’s d=0.738); no significant speed advantage; interaction pattern dominates: delegation 24–39%, engagement 65–86% HIGH / TIER 1 (pre-registered; counter-interest Anthropic finding)

Afshar et al. — Ambient AI Scribe Burnout: First Stepped-Wedge RCT (NEJM AI, November 2025)

Source: research/06-industry-verticals/afshar-nejm-ai-ambient-scribe-stepped-wedge-rct-2025.md

24-week stepped-wedge pragmatic RCT, n=66 health care practitioners, ambulatory clinics, 2 US states. Published NEJM AI, November 26, 2025. This is the first randomized evidence (not observational) that ambient AI scribes reduce physician burnout.

Key results:

  • ~30 minutes per day per provider documentation reduction
  • Clinically meaningful burnout reduction (Stanford Professional Fulfillment Index)
  • Billing code accuracy improved

Why it matters for the RCT corpus: The stepped-wedge design eliminates self-selection bias that limited prior observational evidence. Physicians who adopt AI scribes early are different from those who don’t — removing that confound makes the burnout finding causal, not correlational. Complements Holmgren et al. (financial output) and Lukac et al. (product comparison) to form the strongest evidence stack for any single AI deployment category in any industry.

For executives: This is the human capital input side of the ambient AI scribe ROI equation. Holmgren measures revenue per physician. Afshar measures practitioner sustainability. AAMC estimates physician burnout turnover costs $500K–$1M per departure.


Lukac et al. — First Head-to-Head Ambient AI Scribe RCT (NEJM AI, November 2025)

Source: research/06-industry-verticals/lukac-nejm-ai-ambient-scribe-three-group-rct-2025.md

Parallel 3-group pragmatic RCT, n=238 outpatient physicians, 14 specialties, 60 days. DAX Copilot vs. Nabla vs. usual-care control. 1:1:1 covariate-constrained randomization. Published NEJM AI, November 26, 2025. First comparative effectiveness RCT of commercial ambient AI scribe products.

Key results:

  • Nabla: −9.5% log writing time vs. control (P=0.02) — statistically significant
  • DAX Copilot: −1.7% vs. control (P=0.66) — not significant
  • Both products: burnout improvement (Mini-Z scores, task load, work exhaustion) vs. control
  • Usage rates: DAX 33.5%, Nabla 29.5% (intent-to-treat; per-protocol would show larger effects)
  • One mild adverse event; occasional transcription inaccuracies on both platforms

For executives: The DAX non-significance does not definitively establish product inferiority (60-day period, low usage rates, specialty mix unknown) but it does establish that brand size is not a proxy for RCT performance. Procurement should require structured 90-day pilots with objective EHR log data, not self-reported time savings.


Holmgren et al. — AI Scribe and Physician Financial Productivity (JAMA Network Open, January 2026)

Source: research/06-industry-verticals/holmgren-jama-ambient-ai-scribe-physician-productivity-2026.md

The largest published study of AI ambient scribe financial impact. Cohort study using difference-in-differences regression on UCSF Health EHR billing data: 1,202,734 encounters, 1,565 physicians (698 adopters, 867 non-adopters), January 2023–April 2025.

This is not a developer productivity study — the domain is clinical documentation, and the output metric is healthcare reimbursement (RVUs). But it belongs in the RCT corpus as a rigorous natural experiment in AI-assisted knowledge work productivity, with a hard financial outcome measure (revenue) rather than self-reported time savings.

Key results:

  • +1.81 RVUs per week (95% CI 0.86–2.75; P<.001)
  • +0.80 encounters per week (P=.04)
  • ~$3,044 additional annual revenue per physician at Medicare rates
  • Zero increase in claim denial rates

Design strengths: Real billing data (not self-report), difference-in-differences controls for physician fixed effects and system-wide trends, large n, pre/post deployment structure.

Design limitations: Single site (UCSF academic medical center), self-selection among early adopters, cannot fully separate capacity gains from coding improvements.

Relevance for executives: This is the first peer-reviewed financial ROI evidence for a category of AI (documentation) that spans healthcare but also applies wherever knowledge workers spend significant time on structured documentation — legal, compliance, HR, finance. The mechanism (AI reduces low-value documentation burden, freeing capacity for high-value output) generalizes beyond medicine.


Google Transform Code: Large-Scale Observational Study (Tabachnyk et al., January 2026)

54,000 Google engineers (36k treatment, 18k control), 13 months of data, Difference-in-Differences causal inference — the largest longitudinal AI coding productivity study published to date. Vendor-authored; Google studying its own tools on its own engineers.

  • +17.5% Change List Throughput (95% CI: [15.9%, 19.0%]) — tight interval, robust across multiple alternative specifications and cohort exclusions
  • −3.6% Mean Investigation Session Duration (95% CI: [−3.9%, −3.2%]) — developers spent less time leaving the IDE to look things up
  • Active Coding Time per CL: not statistically significant — gains came from more CLs completed, not faster typing
  • Placebo check passed: no significant pre-treatment differences between treatment and control; parallel trends assumption validated
  • Results persisted over multiple post-adoption months (biggest effect in first two months, smaller but positive thereafter)

The key distinction from lab RCTs: authors explicitly note that lab studies “analyze productivity based on specific tasks… while those tasks are very limited in representing an actual developer’s work week.” This study measures real-world throughput over 13 months of normal work.

Credibility caveat: Google studying its own tools on its own engineers with its own training data. The Gemini models fine-tuned on internal Google code — non-replicable by external deployers. Results reflect a years-long iteration on latency, UX, model quality, and measurement — not the out-of-box experience of deploying an off-the-shelf tool.

Source: research/02-corporate-tools/google-ai-ide-productivity-2026.md

Microsoft Security Copilot Agent RCTs (Bono et al., November 2025)

Two vendor-authored RCTs measuring AI agent impact in enterprise security operations. Both used Upwork freelance participants (not in-house enterprise staff). TIER 1 methodology; vendor result bias acknowledged.

Phishing Triage Agent (arXiv:2511.13860, n=167): Three-arm design isolating queue prioritization from verdict labels.

  • 6.5× true positives per analyst minute (ground truth); 3.1× under 80% accuracy assumption
  • +77% F1 improvement; +106–134% precision; +17% recall
  • 53% more analyst time spent on malicious emails — attention reallocation, not rubber-stamping
  • 79–84% of gain came from queue prioritization alone (not verdict labels)

Conditional Access Optimization Agent (arXiv:2511.13865, n=162): Four-task identity governance study in simulated Microsoft Entra environment.

  • +48% accuracy overall; +204% on Zero-Trust gap detection (hardest task)
  • −43% task completion time across all four tasks
  • Largest gains on highest-cognitive-load tasks: consistent with HBS jagged frontier

What this adds to the corpus: These are the first RCTs measuring AI agent productivity in security operations specifically. Pattern matches the jagged frontier: large gains on structured triage tasks (phishing queue, policy checklists) where ground truth is knowable. Does not contradict METR (open-ended developer work) — different task type.

Source: research/06-security-frontier/microsoft-security-copilot-agent-rcts-2025.md · research/12-agent-workers/microsoft-entra-conditional-access-agent-rct-2025.md (Conditional Access deep-dive)


Lång, Gommers et al. (MASAI) — AI-Supported Mammography Screening (The Lancet, January 2026)

The largest completed RCT in clinical AI screening. n=105,934 women ages 40–74 in Sweden’s national population-based mammography screening program. Randomized 1:1 to AI-supported reading (Transpara Detection) vs. standard double-reading without AI. Two-year follow-up completed December 2025. Published The Lancet, January 31, 2026 (PMID 41620232). HIGH credibility; TIER 1.

  • −12% interval cancer rate: 1.55 vs. 1.76 per 1,000 participants (p=0.41, non-inferiority demonstrated)
  • +6.7 pp sensitivity: 80.5% vs. 73.8% (p=0.031) — statistically significant improvement
  • Identical false positives: 98.5% specificity in both arms — the key tradeoff was avoided
  • −27% aggressive (non-luminal A) interval cancers: 43 vs. 59 cases
  • −44% radiologist workload (documented in prior Lång et al., Lancet Oncology 2023 paper, confirmed by Hernström et al., Lancet Digital Health 2025)
  • +29% screen-detected cancers (Lång 2023)
  • Pattern: AI triaged low-risk to single reading, high-risk to double reading — radiologists concentrated effort where signal was highest. Attention reallocation, not replacement. Mirrors the phishing triage mechanism (Microsoft Bono RCTs).
  • What this adds: First domain where AI RCT evidence is unambiguous at clinical scale. Task-type criterion: high volume, defined protocol, measurable ground truth. Extends the jagged frontier into clinical medicine — structured screening tasks respond to AI; open-ended diagnostic reasoning remains unproven at this evidence level.

Source: research/06-industry-verticals/masai-lancet-ai-mammography-rct-2026.md


Ju & Aral (MIT/Johns Hopkins) — The Productivity-Diversity Tradeoff in Creative Teamwork (February 2026)

Largest human-AI teamwork RCT to date. n=2,234 randomly assigned to human-human vs. human-AI teams producing advertisements. GPT-4o (Aug 2024). Validated with 1,200 independent raters and a real-world field experiment on X (~4.9M impressions).

  • +50% output per worker (β=1.954, p<0.001) — robust, large effect, replicates across both study phases.
  • Text quality higher in H-AI (β=0.327, p<0.001). Image quality lower in H-AI (β=−0.133, p<0.001). Gains are dimension-specific, not universal.
  • Diversity collapse: AI-assisted teams produced outputs converging toward a centroid (β=−0.094, p<0.001). Output homogenizes. Delegation to AI drives this (β=−0.083, p<0.001).
  • Real-world null: Despite 50% more output and higher text ratings, H-AI teams showed no statistically significant advantage over H-H teams on CTR, CPC, or view duration in the field experiment. More ≠ better when it comes to market outcomes.
  • Behavioral shift: H-AI teams made 62% fewer direct text edits (β=−1032, p<0.001), sent 25% more task-oriented messages and 18% fewer interpersonal messages. Workers outsource editing to AI and communicate transactionally rather than building shared context.
  • Implication: AI raises volume and text quality floor; compresses creative variance. Organizations competing on differentiated output must actively design against homogenization — not assume volume gains translate to outcome gains.

Source: research/01-ai-native-landscape/ju-aral-collaborating-ai-agents-rct-2026.md


AISI UK Government RCT — The Jagged Frontier Confirmed Across O*NET Tasks (February 2026)

First government-run RCT using the O*NET Generalised Work Activities taxonomy across a broad workforce sample (not just developers or customer service agents). UK AI Security Institute, n=500, treatment = access to early-2025 frontier LLM.

  • Overall: +25% quality, +61% Points per Minute vs. control. No significant time difference on average.
  • Task 1 (Monitoring): +22% quality (p<0.05); no time or PPM effect.
  • Task 2 (Technical Drafting): +23% quality; no time or PPM effect.
  • Task 3 (Planning/Prioritizing): No significant improvement on any metric. Open-ended strategic planning = AI null zone.
  • Task 4 (Interpreting Information for Others): −42% time; +102% PPM (statistically significant); no quality change. Pure throughput leverage.
  • Grading reliability (Krippendorff’s α): 0.693–0.919, with human-grader validation.
  • Pattern: “Jagged frontier” confirmed at government level — gains are large on analysis/synthesis of structured inputs; zero on open-ended judgment. Consistent with METR (complex tasks −19%) and MIT FutureTech (~60% on all tasks, legal 46.8% lowest).
  • Limitation: n=500 across four tasks means ~125 per arm per task; authors flag “high degree of uncertainty.” Preliminary findings; UK workforce population.

Source: research/01-ai-native-landscape/aisi-uk-ai-productivity-rct-2026.md

Sun, Li, Foo, Zhou & Lu (Journal of Applied Psychology) — Creativity Gains Require Metacognitive Infrastructure (December 2025)

First peer-reviewed RCT specifically measuring AI’s impact on creativity (vs. throughput) in a knowledge-work setting. n=250 employees at a technology consulting firm; ChatGPT vs. control; dual-rating design (supervisor + external evaluator). DOI: 10.1037/apl0001296.

  • Overall: LLM access produces higher creativity ratings vs. control — but only for employees with high metacognitive strategy use.
  • High metacognition + LLM: Significantly more novel and useful ideas (both supervisor and external evaluator confirm).
  • Low metacognition + LLM: No significant creativity gain vs. control.
  • Mechanism: LLM as “cognitive job resource” — it supplies information and structural scaffolding. Metacognitive ability determines how effectively that resource is extracted vs. accepted uncritically.
  • Practical implication: AI access is not the bottleneck for creativity-intensive roles. Metacognitive skill development is. Training programs focused on prompt syntax miss the higher-leverage investment.
  • Consistency with corpus: Matches pattern across all RCTs — AI benefits are not uniformly distributed; individual characteristics interact with tool access; experience and cognitive strategy moderate outcomes.

Source: research/07-adoption-challenges/sun-li-genai-creativity-rct-2025.md


Reading Them Together

  • Task type dominates. Greenfield isolated tasks: large gains. Brownfield mature codebases: zero or negative.
  • Experience dominates inversely by population. In customer service (Brynjolfsson), novices gain most. In open-ended developer work (METR), experts lose most. Same direction: AI substitutes for missing skill where the task is bounded; AI interferes with tacit expertise where the task is open-ended.
  • Perception ≠ reality. METR’s 39-percentage-point perception gap is the corpus’s most-cited single finding. Survey-based ROI claims should be discounted accordingly.
  • No RCT has reproduced organizational productivity gains. Individual gains do not aggregate to team or company throughput (Faros 10,000+ developers; NBER Copilot RCT n=7,137 — zero coordination effect).

Statistics Canada SDTIU — The Capability-Dependent Productivity Premium (April 2026)

The only non-US government microdata study in the corpus using a controls design on the AI productivity premium. Mandatory participation (SDTIU) eliminates self-selection bias; administrative linkage to CRA business microdata provides objective productivity measures rather than self-reports. Authors: Jiang Li & Huju Liu, Statistics Canada, Economic and Social Reports, Vol. 6, No. 4, April 22, 2026.

  • Raw premium: AI adopters are 16.8% more productive than non-adopters — this is the number in vendor decks.
  • Controlled for firm quality (size, sector, human capital): Falls to 10.2%.
  • Controlled for complementary capabilities (data analytics + advanced robotics): Falls to 5.1% — statistically insignificant. Confidence interval crosses zero.
  • Adoption gradient: Firms already using data analytics were 15.0 pp more likely to adopt AI; robotics users 8.1 pp more likely.
  • Core finding: The productivity premium reflects the firm’s pre-existing capability stack, not the AI tool. AI adoption follows capability maturity; it does not create it.
  • Practical implication: Infrastructure-before-tooling sequencing is not a planning preference — it is the order in which successful adopters actually proceed, visible in the data.

Source: research/01-ai-native-landscape/statcan-ai-productivity-capability-maturity-2026.md


Otis, Clarke, Delecourt, Holtz, Koning (MIT Sloan MR) — AI Amplifies Pre-Existing Judgment Gaps (April 2026)

The cleanest causal evidence that AI does not produce uniform outcomes — it amplifies the judgment differential that already exists in a population. RCT, n=640 small business owners in Kenya, GPT-4 via WhatsApp, May–November 2023. Published MIT Sloan Management Review, April 20, 2026. HIGH credibility, TIER 1.

Context note: This is a developing-economy small-business study. The mechanism is generalizable; the specific percentages (+15%/−10%) should not be applied as benchmarks to enterprise deployments. Use as mechanism evidence.

  • Top 50% performers at baseline: +15% revenue and profit with AI access
  • Bottom 50% performers at baseline: −10% revenue and profit with AI access
  • Average effect: ~0%, not statistically significant — masks the polarized distribution
  • Both groups received similar AI suggestions and asked similar questions. The difference was entirely in application.
  • High performers identified context-specific, actionable advice and discarded generic suggestions
  • Low performers acted on generic recommendations (lower prices, advertise more) that eroded margins
  • Core finding: “In contexts where problems are broad and fuzzy, generative AI amplifies the role of human judgment.”

Enterprise translation: An undifferentiated enterprise AI rollout measured only by adoption rates will look successful in aggregate while degrading output quality among employees who lack the domain judgment to filter AI recommendations. The Statistics Canada study shows this at the firm level (the productivity premium reflects pre-existing capability); this study shows it at the individual level.

Consistent with:

  • Sun et al. (JAP, 2025): creativity gains only for high-metacognition employees
  • Dell’Acqua et al. (BCG/HBS, 2023): consultants over-trusted AI off the jagged frontier
  • Ni, Wang et al. (Alibaba, 2025): top customer service agents declined with AI; bottom agents improved (different direction but same judgment-moderation mechanism)
  • Daniotti et al. (Science, 2026): experienced developers capture coding AI gains; beginners see no benefit

Source: research/07-adoption-challenges/mit-sloan-ai-judgment-gap-rct-2026.md


NBER WP 34984 — CFO-Level Productivity Paradox (March 2026)

The most direct evidence that perception-vs-measurement gaps extend to the firm level. 748 CFOs (Duke/Fed Atlanta/Fed Richmond) report both their perceived AI productivity impact AND allow researchers to calculate what revenue-per-employee data actually implies.

  • Mean CFO-perceived LP gain (2025): 1.8%
  • Mean implied (measured) LP gain (2025): 0.6%
  • Gap: 1.2 pp — CFOs overestimate their AI productivity gains by roughly 3x
  • Finance sector (where gains are largest): ~0.8% implied LP (2025) → >2% (2026)
  • Cost reduction is NOT correlated with actual gains. Only innovation/demand motivations predict measured revenue productivity improvements.
  • The authors call this the “AI productivity paradox” — explicitly citing Solow’s 1987 computer paradox as a structural template.

Source: research/01-ai-native-landscape/nber-w34984-ai-productivity-corporate-executives-2026.md

HBS/Stanford — The GenAI Wall Effect (September 2025)

The first RCT to test whether GenAI can transfer expertise horizontally across occupations (not just help low performers catch up within the same role). Vendraminelli, DosSantos DiSorbo, Hildebrandt, McFowland III, Karunakaran, and Bojinov (HBS Working Paper 26-011, September 2025; n=78, IG Group UK, 2024 experiment).

Study design: Three groups — web analysts (insiders, n=12), marketing specialists (adjacent outsiders, n=26), technology specialists (distant outsiders, n=40) — attempted to write corporate web articles using a bespoke LLM tool. Two task phases: conceptualization (brief/outline) and article execution (full prose).

Key findings:

  • Conceptualization (planning tasks): AI fully equalizes all groups. Without AI: web analysts 3.82/5.0, marketing specialists 3.04, tech specialists 3.02 (both significantly lower). With AI: web analysts 4.12, marketing specialists 4.18, tech specialists 4.05 — differences statistically indistinguishable (p=0.66 and p=0.63). Time cut from 62.8 → 23.0 minutes (63%, p<.0001).
  • Execution (prose writing): AI equalizes adjacent outsiders only. With AI: web analysts 3.96, marketing specialists 3.92 (p=0.459, not different); tech specialists 3.42 (p=.007 below marketing specialists). Time cut from 87.2 → 22.4 minutes (74%) for all groups — but quality diverged.
  • The mechanism: Marketing specialists could evaluate AI output quality (domain judgment intact). Tech specialists lacked the foundational marketing knowledge to distinguish good copy from bad — they edited AI output in ways that degraded quality while believing they were improving it. “I ended up reversing most of its suggestions because I prefer articles that are clear and direct” — a data scientist removing SEO elements the rubric required.
  • GenAI ceiling vs. GenAI wall: Prior literature documented the “GenAI ceiling” (low performers catching up with high performers in the same role). This paper identifies the “GenAI wall” — the horizontal limit where knowledge distance makes AI output uninterpretable to the user.

Limitation: Single firm, single task domain (web article writing), bespoke AI tool (not generic LLM), n=12 insiders limits statistical power.

Source: research/01-ai-native-landscape/hbs-stanford-genai-wall-effect-expertise-transfer-2025.md

Daniotti, Wachs, Feng & Neffke (Science) — AI Coding Diffusion and the Experience Divide (February 2026)

The largest observational study of AI adoption in software development. Not an RCT — observational — but the dataset scale (30M commits, 160K developers, 2019–2024) and peer-reviewed Science publication make it the most robust measurement of real-world AI coding adoption available.

  • Adoption: AI tools generate ~29% of Python functions written by US developers by end of 2024, up from ~5% in 2022. France (24%), Germany (23%), India (20% and rising). US lead is “modest and shrinking.”
  • Aggregate productivity: +3.6% quarterly output by end of 2024. Estimated $23–38B annual value in US software sector.
  • Experience divide: Less experienced developers use AI in ~37% of code; experienced in ~27%. Yet experienced developers capture nearly all productivity gains. Beginners show no statistically significant productivity improvement.
  • Mechanism: Experienced developers use AI to explore new domains and libraries — amplifying range. Beginners use AI to generate code they cannot evaluate or debug — producing volume without shipping value.
  • No gender gap: AI usage rates show no gender differences.
  • Skill gap implication: AI adoption is widening, not closing, the gap between experienced and novice developers.

Key quote: “Beginners hardly benefit at all.” — Simone Daniotti (EurekAlert, February 2026)

Limitation: Observational, not experimental. Python only. GitHub public repo data; enterprise private repos excluded. 2019–2024 dataset predates 2025–2026 model generation.

Source: research/02-corporate-tools/daniotti-wachs-who-is-using-ai-to-code-2026.md

Cui, Demirer, Jaffe et al. (Management Science) — The 26% Real-World Baseline (Experiments 2022–2023, Published 2025)

The largest multi-site corporate RCT of AI coding assistants. Three experiments at Microsoft (n=1,746), Accenture (n=320), and an anonymous Fortune 100 electronics manufacturer (n=3,054) — 4,867 developers total. Published in Management Science (DOI: 10.1287/mnsc.2025.00535). Tool tested: GitHub Copilot, 2022–2023 versions. TIER 3 data — predates current model capabilities by 2–3 generations, but strongest available real-world RCT evidence.

Key findings:

  • +26.08% (SE: 10.3%) increase in weekly pull requests for Copilot adopters (LATE estimate, IV specification). Secondary: +13.55% commits, +38.38% builds.
  • Experience gap is the headline: Junior/short-tenure developers gain 21–39%. Senior/long-tenure developers gain ~8%, with some abandoning the tool. Short-tenure devs 9.4pp more likely to adopt.
  • No quality decline: PR approval rates held or improved at Microsoft (only site with quality data). No evidence of increased defect rates over experiment window.
  • Lab vs. real-world gap: Lab studies find 58% time reduction (Peng et al.); real-world is 26%. Gap explained by non-coding tasks absorbing time savings and organizational friction limiting workflow change.
  • External validity argument: Three different company types across 8 months of observation, random assignment, IV identification — strongest design in the enterprise coding AI literature.

Read against Daniotti et al.: Both studies find experienced developers drive gains while less experienced developers lag — but from opposite directions. Cui finds junior developers gain more (+27–39%) while Daniotti finds experienced developers capture all aggregate gains. These are not contradictory: Cui measures individual-level productivity lift within enterprise settings where junior developers have defined tasks; Daniotti measures aggregate output across open-source contributions where beginners generate code they can’t ship.

Source: research/02-corporate-tools/demirer-cui-three-field-experiments-developer-productivity-2025.md

INSEAD/HBS RCT — The Mapping Problem (March 2026)

The first rigorous field experiment to identify WHY individual AI task gains don’t aggregate to firm-level performance. Kim, Kim, and Koning (INSEAD/HBS, n=515 startups, SSRN #6513481, March 2026) isolated the “mapping problem” — the failure to identify where in a production chain AI creates compound (not additive) value.

  • Treatment: Firms receiving case studies on how other companies reorganized production around AI (informational intervention only — same tools/training as control)
  • Revenue: 1.9× higher; Customer acquisition: +18%; Tasks completed: +12%
  • Capital required: −$224,000 (−39.5%) with no headcount increase
  • Treatment effect by technical skill: Zero variation — the bottleneck is cognitive/informational, not technical
  • Distribution: Revenue gains concentrated at 90th–95th percentile — only firms that reorganized complete production chains captured structural economic shifts
  • Key mechanism: AI creates compound value across interconnected workflow steps (O-Ring/complementarity structure). Adding AI to one step in a manual chain produces isolated gains and often just moves the bottleneck.
  • Implication for enterprise programs: Tool subsidies and prompting workshops address the wrong constraint. The constraint is production-chain mapping.

Source: research/01-ai-native-landscape/insead-hbs-mapping-ai-production-rct-2026.md

Stanford / ADP — Canaries in the Coal Mine: Large-Scale Employment Evidence (November 2025)

Source: research/01-ai-native-landscape/stanford-canaries-coal-mine-ai-employment-2025.md

Not an RCT, but the most rigorous observational study of AI’s labor market effects. ADP payroll data, 3.5–5M workers/month, causal design with firm-time fixed effects. Brynjolfsson, Chandar, Chen (Stanford DEL / NBER). Published November 2025.

The central finding shifts the question from “does AI improve individual productivity?” to “what does AI adoption do to the workforce that builds the next generation of productive workers?”

  • 16% relative employment decline for ages 22–25 in highest AI-exposure quintile vs. lowest, after firm-time fixed effects. Software developers 22–25: ~20% employment decline from late 2022 peak.
  • Augmentation doesn’t kill entry-level; automation does. The Anthropic Economic Index classification (Claude usage patterns) distinguishes automative from augmentative AI use at the occupational level. Only automative AI correlates with entry-level employment declines. This is the critical design implication: choice of AI deployment mode determines whether junior hiring contracts.
  • Wages don’t warn you; headcount does. No compensation divergence by age or exposure. Firms reduce inflows quietly. Workforce planning dashboards that only monitor compensation miss the structural shift until it compounds.
  • Mechanism: AI replaces codified knowledge (what entry-level roles supply) while complementing tacit knowledge (what experienced workers hold). Experience is now the buffer against displacement.

Stanford HAI AI Index 2026 — Task-Level and Macro-Level Productivity Synthesis (April 2026, HIGH)

Full research file: research/01-ai-native-landscape/stanford-hai-economy-2026.md

The Stanford HAI synthesis aggregates the RCT corpus alongside macro BLS data — the first authoritative annual synthesis to combine both levels:

  • Task-level gains: 14%–50% in narrow structured tasks (customer support, software development, marketing). Zero or negative in tasks requiring judgment. This range brackets the individual studies above: it is consistent with the Brynjolfsson customer-support RCT (+15%), the BCG/HBS consultant study (+12%), and the software coding studies — and explicitly excludes judgment-intensive work, consistent with AISI UK (zero gain on planning tasks) and METR (−19% for experienced developers on complex real-world tasks).
  • US macro labor productivity: +2.7% in 2025, nearly double the prior-decade average (BLS data). This is the first macro signal that AI task-level gains are beginning to aggregate across the economy. However, the distribution is highly uneven — the aggregate rise does not imply that most individual firms are seeing commensurate gains. It is consistent with a small fraction of high performers capturing large gains that shift the average.
  • 88% organizational adoption, single-digit EBIT impact share. The deployment-to-value gap in the RCT evidence (no RCT has reproduced organizational productivity gains) is visible in the macro survey data: 88% of organizations use AI in at least one function; fewer than 6% capture significant financial value.

The task-level/macro-level pattern is the key interpretive frame for citing productivity evidence: individual task gains are real and large, but do not automatically aggregate to firm-level returns. Workflow redesign is the documented mechanism for crossing that gap (McKinsey March 2025: workflow redesign = #1 EBIT predictor out of 25 attributes).

Source: research/01-ai-native-landscape/stanford-hai-economy-2026.md — HIGH credibility, April 2026

Cui, Demirer, Jaffe et al. — Management Science Enterprise RCT (February 2026)

The largest peer-reviewed RCT on AI coding tools to date. Three pre-registered field experiments at Microsoft, Accenture, and an anonymous Fortune 100 company. n=4,867 software developers in ordinary production conditions. Published Management Science (INFORMS), February 27, 2026.

  • +26.08% completed tasks (SE: 10.3%) — statistically significant combined effect across three enterprise sites
  • Less experienced developers drove the result: higher adoption rates and greater productivity gains among junior/mid-level developers
  • Senior developers: smaller or negligible gains — reconciling with METR (experienced developers 19% slower), not contradicting it
  • Cross-site variation is real — “noisy” individual experiments; significance emerges from meta-analysis
  • Volume metric only: “completed tasks” does not capture code quality; CMU/Yetistiren found 40.7% code complexity increase with Copilot — a downstream maintenance cost invisible here
  • Funded by MIT GenAI Initiative; Microsoft author affiliation is a limitation for unbiased framing

Synthesis with METR: Both studies are correct. METR tests experienced developers on open-ended, complex mature-codebase tasks. Cui et al. tests a mixed-experience population on ordinary production work. The experience moderator reconciles both findings: AI coding assistance levels up less experienced developers on defined tasks; it adds overhead for senior developers on complex judgment-intensive work.

Source: research/07-adoption-challenges/management-science-genai-software-developers-rct-2026.md


Suh & Oh (Bank of Korea) — The Leisure Channel: Where AI Efficiency Goes (February 2026)

The first nationally representative worker survey to simultaneously measure time savings, output changes, and worker welfare from GenAI. n=5,512 Korean workers. Bank of Korea funded, independent academic. Fieldwork May–June 2025. Published February 12, 2026 (arXiv).

  • 51.8% of Korean workers use GenAI for work — roughly twice the U.S. comparable rate
  • 3.8% time reduction for active users; 1.4% economy-wide (including non-users)
  • Weighted correlation (time savings ↔ output): 0.008 — effectively zero. Confirmed in regression across multiple specifications with occupation and industry fixed effects
  • Workers capture efficiency gains as on-the-job leisure (+1.3 pp share of workday) rather than expanded output
  • 36.1% of users report output declines — the friction and learning-curve portion
  • Less experienced workers save more time: 20-year experience gap = 1.9 pp advantage in time savings, favoring junior workers — consistent with Cui et al. and METR direction
  • Job satisfaction rises with time savings (β=0.039, p<0.01) even when output is flat — welfare gains are real but invisible in GDP
  • GenAI is a targeted accelerant: top-saving task = ~70% of total time reduction

The structural implication: Efficiency gains from AI become organizational output only when incentive structures explicitly connect time savings to deliverable scope expansion. Without that redesign, workers — rationally — capture gains as reduced workload. This is the micro-level mechanism behind the macro finding in NBER w34836 (six-country executive survey, n~6,000): 90% of executives report no AI impact on employment or productivity over three years.

Korea scope note: Korea’s GenAI adoption is ~2× U.S. rates. Directional mechanisms generalize; absolute percentages should not be transferred directly to a U.S. enterprise workforce.

Source: research/07-adoption-challenges/suh-oh-genai-time-reallocation-korea-2026.md


Hartley, Jolevski, Melo & Moore — Labor Market-Level Evidence (January 2026)

U.S. occupation-level analysis linking worker surveys to administrative labor-market data. SSRN 5136877 (Jan 2026). Authors: Hoover/Stanford, World Bank, RAND/Clemson, Stanford — no AI vendor interest. TIER 1.

  • 35.9% of U.S. workers used generative AI tools by December 2025, up from ~30% twelve months earlier. Adoption is accelerating, not stalling.
  • Adoption is concentrated among younger, college-educated, and higher-earning workers. Industries: customer service, marketing, IT lead; trades, manufacturing, healthcare lag.
  • Small but statistically significant positive wage effects in AI-exposed occupations through Dec 2025. AI augments rather than displaces at current adoption levels.
  • No statistically significant decline in job openings or aggregate employment in AI-exposed occupations — macro displacement has not materialized in administrative data through end of 2025.
  • Cross-reading with Suh & Oh (Bank of Korea): both studies find the macro labor market is absorbing AI tools without disruption at the aggregate level — but Suh & Oh shows why (time savings captured as leisure, not output). Hartley et al. shows wages are rising, not the output they should be generating.
  • The reconciliation with METR (−19%): individual-level and aggregate-level findings can both be true. High-exposure workers as a class are earning more; individual workers using AI incorrectly may be slower.

Source: research/07-adoption-challenges/hartley-jolevski-melo-moore-genai-labor-market-2026.md


Foxit / Sapio Research — The Verification Burden (March 2026)

Large-scale behavioral survey that quantifies the mechanism by which individual AI time savings are erased at the workflow level. n=1,400 (1,000 end users + 400 executives), US + UK, independent Sapio Research fieldwork. Published March 11, 2026 (TIER 1).

  • Executives perceive 4.6 hrs/week saved from AI; spend 4 hrs 20 min validating outputs → net gain: 16 minutes/week
  • End users perceive 3.6 hrs/week saved; spend 3 hrs 50 min reviewing outputs → net loss: 14 minutes/week
  • U.S. respondents net −10 min/week; U.K. respondents net +2 min/week
  • 89% of executives “believe AI boosts productivity” — the perception gap is 4+ hours wide
  • 60% of executives express high confidence in AI accuracy; only 33% of end users share that confidence — the trust divergence drives verification behavior
  • 68% of executives report AI has triggered restructuring or headcount changes; only 12% of end users are “very concerned” about job security
  • 93% of organizations now track Return on Employee (ROE) alongside ROI metrics

Mechanism: The “verification burden” — time spent reviewing, fact-checking, and correcting AI outputs before confident use. The study directly measures the delta between perceived savings and actual net time after checking. This is the behavioral analog to METR’s perception gap (developers believed they were 20% faster, were actually 19% slower): the tools feel faster; the downstream checking cost is invisible in self-reports.

Corroborates: Workday/Hanover Research (40% of AI time savings consumed by rework, only 14% net-positive outcomes), ActivTrak (every work category increased post-AI, focus efficiency at 3-year low), Goldman Sachs (“basically zero” GDP contribution).

Implication for program design: Verification burden is a workflow design problem, not an accuracy problem. Solving it requires: (1) accuracy investment before broad deployment, (2) explicit workflow redesign of the checking step, (3) trust calibration policies that define verification intensity by task type and risk level.

Source: research/07-adoption-challenges/foxit-sapio-document-intelligence-ai-productivity-2026.md


Supporting Research

File Angle
research/01-ai-native-landscape/nber-generative-ai-at-work-brynjolfsson-2023.md Customer service RCT, mechanism analysis
research/13-multimodal-sources/research-papers/brynjolfsson-generative-ai-at-work-rct-2025.md QJE 2025 large-scale RCT update: 15% productivity floor, novice vs. expert divergence, RAG-grounded support
research/01-ai-native-landscape/hbs-cybernetic-teammate-rct-2025.md P&G n=776, AI replaces team coordination
research/01-ai-native-landscape/ai-pair-programming-research.md METR + 6 other developer studies, methodology critique
research/01-ai-native-landscape/academic-ai-productivity-papers.md Eight-study landscape, RCT-vs-observational
research/04-consulting-firms/mckinsey-ai-developer-productivity.md McKinsey methodology critique
research/04-consulting-firms/mckinsey-methodology-critique.md Deep standalone critique: ~40-developer lab → $2.6–4.4T extrapolation; METR 19% slower contradicts McKinsey; six independent studies converge on ~10% org productivity gain; Kent Beck/Orosz metric-gaming critique
research/07-adoption-challenges/measuring-ai-developer-tool-roi.md Enterprise measurement frameworks (DORA, SPACE); METR follow-up: task-level experiments face “insurmountable obstacles”; ROI metric shift from productivity to P&L
research/08-radical-vs-tablestakes/spectrum-analysis.md Maps AI engineering maturity from table stakes (code autocomplete) to emerging standard (test generation, AI code review) to radical frontier; includes METR RCT context on where productivity gains and losses cluster by task type and AI maturity tier
research/01-ai-native-landscape/real-roi-by-function.md Where the RCT signal generalizes
research/01-ai-native-landscape/openai-state-of-enterprise-ai-2025.md OpenAI n=9,000 self-report: 40–60 min/day saved on average; 6x gap between frontier and median workers — apply vendor caveat + perception-vs-reality critique
research/01-ai-native-landscape/stanford-enterprise-ai-playbook-2026.md 51-deployment observational study: 71% median productivity with high-automation workflow vs. 30% with human-primary — not RCT but most rigorous deployment-level analysis
research/01-ai-native-landscape/enterprise-rag-guide.md RAG-grounded customer support: Brynjolfsson QJE 2025 +14% avg (+34% novices); RAG failure mode taxonomy; Gartner Q4 2025 n=800 on context-stuffing vs. retrieval
research/07-adoption-challenges/activtrak-state-of-workplace-2026.md Longitudinal behavioral study (n=163,638, 443M hours): AI intensifies work density rather than reducing workload; 80% adoption → 3-year-low focus efficiency; only 3% of users reach optimal AI usage window — largest behavioral observational dataset in corpus
research/07-adoption-challenges/betterup-stanford-workslop-ai-quality-cost-2025.md The hidden cost of AI-generated volume without quality: 40% of workers received AI-generated “workslop” in a given month; $186/employee/month in cleanup cost; 40% of AI time savings consumed by rework (corroborated by Workday/Hanover n=3,200) — TIER 2 (Sep 2025 fieldwork, Stanford/BetterUp)

BIS/EIB European Firms Study (January 2026)

The largest causal firm-level productivity study in the corpus — not an RCT but the closest thing at enterprise scale.

  • n=12,000+ EU non-financial firms + 800 US firms, EIBIS survey 2019–2024 matched to ORBIS financial data.
  • Causal identification: Instrumental variables — each EU firm matched to comparable US firms; US AI adoption rate assigned as exogenous instrument. This sidesteps the selection bias problem (productive firms adopt AI first) that plagues observational studies.
  • Headline finding: +4% causal labor productivity (turnover per employee, from administrative records; not self-report).
  • No employment effect after instrumentation — the productivity gain comes from workers producing more, not from hiring fewer.
  • Workers benefit: ~7% higher total wages in AI-adopting firms.
  • Size stratification is critical:
    • Micro/small (<50 employees): no statistically significant gain
    • Medium (50–249): +4% (marginal)
    • Large (250+): +7.9% (strong)
  • Complementary investment multiplier:
    • Training: +5.9% additional productivity per additional percentage point of investment share
    • Software/data: +2.4% additional productivity per pp
  • Interpretation: The 4% firm-level finding is consistent with “mid-range” macro projections (Bergeaud: 0.3pp TFP/year; Aghion-Bunel: 0.7pp). It reconciles task-level optimism (40% writing gains, Noy and Zhang 2023) with macro skepticism (Acemoglu: 0.07% TFP/decade) — adoption frictions and skill gaps absorb most of the task-level signal before it reaches firm productivity.

Source: research/01-ai-native-landscape/bis-eib-european-firms-ai-productivity-2026.md


Anthropic Economic Index (Nov 2025 / Mar 2026)

  • Not an RCT — observational analysis of 100,000 (Nov 2025) then 1,000,000 (Feb 2026) Claude conversations via CLIO.
  • Reports ~80% average task-time reduction (median 84%), projected to 1.8% annual US productivity growth over a decade under universal adoption.
  • Cannot be read as comparable to METR’s –19%: measures only successful in-conversation time, no post-conversation rework, no control group, selection bias toward favorable tasks.
  • High-tenure users show +3–4pp success advantage after task-type controls — learning curve compounds.
  • 14% drop in job-finding rates for ages 22–25 entering exposed occupations post-ChatGPT (Labor Market paper, Mar 2026).

Sources: research/01-ai-native-landscape/anthropic-economic-index-2026.md · research/01-ai-native-landscape/anthropic-economic-index-learning-curves-2026.md (dedicated learning-curve and capability-scaling analysis)

MIT FutureTech — Capability Trajectory Across Real Labor-Market Tasks (April 2026)

Not an RCT measuring productivity gains — rather, the largest published evaluation of AI capability levels on real O*NET job tasks. Distinct contribution: documents how fast AI capabilities are growing across the full range of white-collar work, not just coding or customer service.

  • n=17,000+ worker evaluations across >3,000 text-based O*NET tasks; >40 LLMs tested; domain-expert human evaluators; April 2026 preliminary findings (ongoing survey).
  • Mean success rate: ~60% — frontier AI models produce minimally-sufficient outputs on roughly six in ten real professional tasks without edits.
  • Rising tide, not crashing wave: Success-duration slope is flat across most job families (β = -0.31 pooled) — AI improves similarly on short and long tasks simultaneously. Contradicts METR/Kwa software-engineering benchmark slopes (β ≈ -1.0) which suggest sudden capability jumps on narrow task types.
  • Pace: Failure rates halving every 2.4–3.2 years. Success rates increasing 8–11 pp annually across 5-min–24-hr task range. Task-duration doubling time at 50% success: 3.8 months.
  • 2024 → 2025 jump: Frontier models went from 50% success on 3–4 hour tasks (Q2 2024) to 50% success on ~1 week tasks (Q3 2025).
  • Projection: 80–95% success on most text-based tasks by 2029 (upper-bound estimate under continued trend growth).
  • For productivity debates: This paper measures capability, not deployment productivity. The gap between 60% task-level success and 0.6% firm-level LP gains (Fed Atlanta CFO data) is the measurement gap in action — integration cost, last-mile friction, and workflow redesign absorb most capability before it reaches P&L.

Source: research/01-ai-native-landscape/mit-futuretek-crashing-waves-rising-tides-ai-automation-2026.md

Longitudinal Evidence (18+ months)

  • Brynjolfsson/Li/Raymond (QJE 2025, n=5,179 support agents, Nov 2020–2022): 14% gain appears in month 2 post-adoption and remains stable and persistent through end of sample — not a novelty effect.
  • Outage event studies: workers exposed to AI for 3+ months handle chats 15–25% faster than pre-AI baseline even when AI is unavailable. Durable human-capital transfer.
  • NAV IT Copilot panel (Stray et al., arXiv 2509.20353, Sep 2025, n=39 devs, 26,317 commits, 105 weeks): zero statistically significant change in commit output over 20 months. Self-reported vs actual productivity correlation ρ=0.17 (p=0.40).
  • The two studies disagree because they measure different deployments: narrow repeatable task with clear metric (support) delivers durable gains; open-ended task with ambiguous metric (coding) does not move measurable output.
  • Market-level retention: Copilot paid-AI share fell 18.8% → 11.5% in six months (Jul 2025 → Jan 2026). 64% of employees with Copilot access do not use it. Tool-specific churn is material even as aggregate adoption grows.

Source: research/07-adoption-challenges/longitudinal-ai-adoption-18-months.md

NBER “Shifting Work Patterns with Generative AI” (Dillon, Jaffe, Immorlica, Stanton — May 2025)

  • Largest randomized field experiment of generative AI in real workplaces to date: 7,137 knowledge workers, 6 months, cross-industry, Microsoft 365 Copilot.
  • Regular users (≥50% of weeks): -3.56 hrs/week on email (-31%). ITT: -1.29 hrs/week. Clear individual productivity signal.
  • Zero statistically significant change in time spent in meetings, recurring meetings, number of documents completed, or emails replied to.
  • Authors’ mechanism: workers changed behaviors they could change unilaterally; did not change behaviors requiring coordination with colleagues.
  • Coworker Copilot share had no significant effect on adoption after firm fixed effects — firm identity was the dominant predictor of usage, suggesting managerial practices and training drive outcomes more than peer density.
  • Collaborative documents with multiple contributors showed the largest time-to-complete reduction (~25%, small subsample) — the one coordination-adjacent result, consistent with AI lowering the cost of tracking colleagues’ edits.

Source: research/07-adoption-challenges/ai-team-coordination-evidence.md

Bain Coding Productivity Data (Sep 2025)

  • AI coding tools: 10–15% productivity gains when applied narrowly to code generation; 25–30% with lifecycle-wide transformation (Technology Report 2025, no sample disclosed).
  • Financial services RCT: 26% increase in completed tasks among 4,900 coders at three large companies (Bain FinServ survey, Jul 2024).
  • Both findings consistent with METR’s result: narrow code-assist gains are real but modest; the lifecycle bottleneck is non-coding work.

Source: research/04-consulting-firms/bain-ai-research-2026.md

Forrester State of AI Survey (n=1,400+, 2025)

  • Not an RCT but the largest analyst-firm survey specifically measuring EBITDA impact: only 13-15% of AI decision-makers report positive EBITDA lift from AI. Fewer than one-third can tie AI to any P&L change.
  • Corroborates the RCT finding that individual productivity gains do not aggregate to organizational financial outcomes without workflow redesign.
  • 48% of firms cut headcount on projected AI gains; most cannot demonstrate those gains materialized — the gap between deployment (68% in production) and value capture (<15% EBITDA lift) is the widest documented by any analyst firm.

Source: research/04-consulting-firms/forrester-ai-research-2026.md

Forrester Developer Experience & AI Tools (Developer Survey 2025)

  • Coding (48%) and testing (47%) are the top SDLC phases for AI adoption; trailing phases like development insights (33%) confirm AI remains a narrow coding accelerator.
  • Analyst Devin Dickerson documented the “honeymoon” effect: initial productivity felt dramatic, but architectural drift, cascading multi-file errors, and test coverage collapse followed — aligning with METR’s finding that perceived gains diverge from measured outcomes.
  • Forrester predicts CS enrollment drops 20% and time-to-hire for developers doubles, creating a paradox where routine coding gets cheaper while the architects who govern AI output become scarcer.

Source: research/05-analyst-firms/forrester-developer-experience-ai-tools.md

PwC 29th CEO Survey + AI Jobs Barometer (Jan 2026 / Jun 2025)

  • 56% of CEOs (n=4,454) report zero financial return from AI — neither revenue nor cost benefit. Only 12% achieved both.
  • The vanguard (12%) deploys AI more extensively and has built foundations first: technology integration, road maps, responsible AI, culture. They are 3x more likely to report meaningful returns.
  • Industry-level data (~1B job ads) shows 3x higher revenue-per-employee growth in AI-exposed industries (27% vs 8.5%, 2018–2024). The value exists — most companies fail to capture it.
  • Only 14% of workers use GenAI daily (PwC Workforce Hopes & Fears 2025). The adoption floor remains low.
  • PwC frames the gap as a foundations problem, not a technology problem — consistent with METR (wrong tasks), Faros (bottleneck migration), and MIT CISR (maturity stages).

Source: research/04-consulting-firms/pwc-ai-research-2026.md

METR Self-Reported Productivity Survey (May 2026, n=349)

  • Self-report survey of 349 technical workers (87 software engineers, 71 researchers, 129 academics/PhD students, 48 founders/managers), Feb–Apr 2026.
  • Median self-reported value gain: 1.4–2x (depending on question framing). Retrospective March 2025 estimate: 1.3x; current (March 2026): 2x; forecast March 2027: 2.5x.
  • Speed median: 3x — but METR explicitly flags this overstates value gains, as workers substitute toward lower-value tasks when AI makes them faster.
  • Critical caveat from METR itself: a prior study found respondents overestimate AI’s time impact by 40 percentage points on average. METR staff — most aware of this gap — reported the lowest value gains.
  • Seven respondents claiming ≥10x gains were reviewed qualitatively; external productivity indicators did not match their self-reports.
  • Why METR abandoned RCT design (Feb 2026): 30–50% of developers refused to submit tasks they thought AI could accelerate; developers refused to work without AI even at $50/hour compensation; agentic multitasking made time-tracking unreliable.
  • Reading with the July 2025 RCT: Not contradictory — they measure different eras, models, and task types. The survey captures rising perceived gains; the RCT captures controlled measurement. The honest range is between −19% (RCT floor, mid-2025 models, mature OSS codebases) and 2x (self-reported ceiling, 2026 models, mixed work).
  • Key signal: The behavioral finding — that experienced developers will not work without AI — is more durable than any specific number.

Source: research/01-ai-native-landscape/metr-self-reported-productivity-survey-2026.md

BCG Tech Function Productivity Claims (Jan 2026, n=1,250)

  • Self-reported survey: SDLC 25% current productivity gain, 44% expected at scale. Data management 25%+, expected 45%+. Compliance 20%+, expected 45%.
  • These are respondent estimates, not independently measured outcomes. The METR perception gap (perceived +20% vs. actual −19%) applies directly: survey-based productivity claims should be discounted.
  • The 5%/35%/60% value distribution (measurable value / scaling / no material value) is consistent with Gartner’s 72% breaking even or losing and PwC’s 56% zero return.
  • Useful as a market sentiment indicator, not as productivity evidence. The tech function is where companies report the most AI value — but only 5% report measurable value at all.

Source: research/04-consulting-firms/bcg-ai-tech-function-payoff-2026.md

AWS re:Invent 2024–2025 Customer Sessions

  • Infosys survey (n=1,500 CXOs, Dec 2024): fewer than 50% of executives are satisfied with their AI results; fewer than 20% of PoCs reach production; only 12% have an AI strategy in place; only 11% have AI governance. The sample is CXO-level, the methodology is a vendor-commissioned survey, and the source is a conference keynote — treat as MEDIUM credibility.
  • These numbers corroborate the BCG 5% / McKinsey 6% high-performer findings from independent sources: the pilot-to-production failure rate and the strategy gap are consistent across independent surveys, analyst firm research, and now this vendor-disclosed CXO sample.
  • The corroboration matters not because it upgrades the Infosys numbers, but because three methodologically distinct sources converging on the same range (<20% production, ~10–12% high performers) reduces the probability the independent findings are methodology artifacts.
  • Credibility note: vendor conference source. Do not cite standalone. Cite as corroborating context when paired with BCG or McKinsey primary findings.

Source: research/13-multimodal-sources/aws-reinvent/aws-reinvent-2024-2025-enterprise-ai-sessions.md

Microsoft Ignite / Build 2024–2025 (Vendor-Published)

  • Forrester TEI (n=367, Mar 2025, Microsoft-funded): 9 hrs/month saved per Copilot user, 116% ROI over 3 years, 2.6% revenue lift. Vendor-funded; directionally consistent with self-reported savings from Campari (16–30 min/day) and Vodafone (3 hrs/week). Higher than independent studies find.
  • Vodafone (n=300 initial cohort): 3 hrs/week per employee; legal staff 4 hrs/week. Self-reported, no control group.
  • Lumen: 96% reduction in customer outreach research time. Single-workflow measurement.
  • These vendor-published figures contrast with NBER Copilot RCT (n=7,137): email time saved but zero change in meeting load, task composition, or delivery throughput. Time saved ≠ value captured remains the core tension.

Source: research/13-multimodal-sources/microsoft-ignite-build/microsoft-ignite-build-2024-2025-enterprise-ai-sessions.md

Gartner AI-Augmented Software Engineering Projections (2025–2026)

  • 90% of enterprise software engineers will use AI code assistants by 2028 (up from <14% in early 2024) — the fastest enterprise technology adoption curve in decades.
  • Gartner’s defect warning: prompt-to-app development will increase software defects by 2,500% by 2028 without governance, architectural checkpoints, and human review. This directly challenges the naive productivity narrative: speed without quality governance is a net negative.
  • Average GenAI project cost: $1.9M per initiative, with less than 30% of CEOs satisfied with ROI. This is consistent with the BCG 5% / McKinsey 6% high-performer finding — most organizations spend without capturing returns.
  • First Gartner MQ for AI Code Assistants (Sep 2025): 14 vendors evaluated; Leaders: GitHub, AWS, GitLab, Google Cloud, Cognition (Windsurf).

Source: research/05-analyst-firms/gartner-ai-software-engineering.md

Vendor Claim Critical Analyses

  • GitHub Copilot: The 55% headline (Peng 2023, n=95, greenfield JS) does not reflect enterprise development. Accenture enterprise study (May 2024, vendor-led) found 8.69% PR increase. MIT/Microsoft field experiments (n=1,663 + n=311) found 7–22% PR increases but researchers flag both as “poorly powered.” Forrester’s “376% ROI” is based on interviews with six people. The perception-vs-reality gap is the most important finding: developers believe AI makes them faster even when data shows the opposite.

Source: research/02-corporate-tools/github-copilot-economic-impact-critical-analysis.md

  • Cursor: Sarkar (U Chicago, Nov 2025) finds 39% more merged PRs using Cursor platform data — but does not measure code quality. Carnegie Mellon (n=807 repos, Nov 2025) directly contradicts: velocity gains are transient (first 2 months), while static analysis warnings rise 29.7% and complexity rises 40.7% persistently. All vendor case studies (Coinbase, Stripe, NVIDIA, Upwork) are executive testimonials with no control groups or disclosed methodology.

Source: research/02-corporate-tools/cursor-enterprise-case-studies-critical-analysis.md

  • Google Gemini Code Assist: Pichai’s “30% of new code is AI-generated” (Q1 2025 earnings call) and “10% engineering velocity increase” (mid-2025) have no published methodology, no sample size, no independent verification. Customer case studies are thin — ComplyAdvantage 37% dev-time reduction is a single-company self-report; Delivery Hero’s 4,000-engineer rollout reports only “developer satisfaction” perception metrics. The most credible data carrying Google’s name is DORA 2025 (n~5,000) — and it directly contradicts the earnings-call narrative: a 25% increase in AI adoption correlates with –1.5% throughput and –7.2% delivery stability. DORA’s AI Capabilities Model finding (“AI doesn’t fix a team; it amplifies what’s already there”) is the actionable insight.

Source: research/02-corporate-tools/google-gemini-code-assist-critical-analysis.md

Practitioner voices (pillar 13)

“The junior software engineering market is terrible. If you know anyone who’s graduated, they can’t find jobs. It’s really bad. And that’s because AI has hit its stride there.” — Dylan Patel, Founder, CEO and Chief Analyst, SemiAnalysis (2026) Source: research/13-multimodal-sources/beyond-the-pilot/2026-04-13-ai-inference-is-reshaping-enterprise-ai-economics-heres-how-.md

“We took the intelligence of a massive model, which is typically on the order of hundreds of billions of parameters in size, and distilled it down into a tiny 600 million parameter model, and then even later down to a 220 million parameter model.” — Aaron Berger, VP of Product Engineering, LinkedIn (2026) Source: research/13-multimodal-sources/beyond-the-pilot/2026-01-21-inside-linkedins-ai-engineering-playbook.md

“Anything but business process re-engineering or re-imagination is a band-aid. The limiting factor is not really technology, but the limiting factor is the human mind. The human mind just recreates the old process again.” — Arya Bolurfrushan, Founder and CEO, Applied AI (2026) Source: research/13-multimodal-sources/ai-for-the-c-suite/2026-04-14-ayra-bolurfrushan-most-companies-are-thinking-about-ai-compl.md

Me, Myself, and AI — Raffaella Sadun, Harvard Business School (Mar 2025)

Source: research/13-multimodal-sources/me-myself-and-ai/2025-03-18-reskilling-the-workforce-with-ai-harvard-business-schools-ra.md · MIT SMR + BCG joint production · HIGH / TIER 2 (2025 fieldwork)

“The average half-life of skills is now less than five years, and in some fields was less than two and a half years.” — Raffaella Sadun, Professor of Business Administration, Harvard Business School (2025)

  • The accelerating skill obsolescence rate is a direct constraint on AI productivity gains: workers trained on pre-AI workflows need continuous reskilling just to maintain current productivity, let alone capture AI productivity upside.
  • Sadun frames AI adoption as requiring significant organizational change management, not just technology implementation — consistent with the RCT evidence that tool access alone (see: Dillon et al., n=7,137) does not shift task distribution or produce measurable productivity at the firm level.
  • The half-life finding reframes AI training budgets: companies that treat AI literacy as a one-time investment will see returns decay within 2–5 years. Continuous learning infrastructure is the durable investment, not a specific tool certification.

Federal Reserve Atlanta / Duke / NBER CFO Survey (n=748, Nov 2025–Jan 2026)

  • CFO-reported (perceived) AI labor productivity gain: 1.8% in 2025, 3.0% expected 2026. Revenue-based (measured) labor productivity gain: 0.6% in 2025, 1.9% expected 2026. The 3:1 gap between perceived and measured is the most precise quantification of the productivity-paradox dynamic in the 2026 corpus — and directly corroborates the METR finding that developers believe they are 20% faster while being 19% slower.
  • By sector (2025 implied measured gain): high-skill services/finance ~0.8%; low-skill services/manufacturing/construction ~0.4%. Expected to roughly double in 2026 for all sectors; finance/high-skill exceeds 2%.
  • Productivity gains driven primarily by innovation- and demand-oriented channels (new products, reaching customers), NOT cost reduction alone — which explains why deployments designed only for headcount elimination underperform.
  • Aggregate employment expected to decline <0.4% due to AI in 2026 (economy-wide). Large firms expect modest net job loss; small firms expect modest net gains. Routine clerical employment expected to decline >2pp over 3 years; skilled technical roles increasing.
  • Note: This is a Federal Reserve working paper (not peer-reviewed), not an RCT. It measures CFO-reported AI-attributed changes, not a randomized experimental comparison. Treat as high-credibility primary survey, not causal proof.

Source: research/01-ai-native-landscape/fed-atlanta-ai-productivity-workforce-2026.md

Stanford AI Index 2026 — Productivity Benchmarks (HAI, published April 2026)

Aggregate secondary analysis across published productivity studies — not a new RCT, but the most comprehensive synthesis of measured AI productivity gains available as of April 2026.

  • 14–15% productivity gain in customer support (replicates Brynjolfsson/Li 2023 Tier 4 baseline; now supported by Tier 1 data from multiple call-center deployments)
  • 26% gain in software development (consistent with GitHub Copilot enterprise telemetry but absent individual-task controls; contrast with METR RCT -19% on open-ended developer tasks)
  • 73% increase in marketing output volume (measured output, not quality-adjusted)
  • Employment for software developers ages 22–25 has fallen nearly 20% from 2024 — consistent with Fed Atlanta CFO survey expectation of routine clerical decline >2pp over 3 years

Temporal tier: TIER 1 (April 2026). Credibility: MEDIUM-HIGH — HAI is an independent academic institution; synthesis methodology varies by underlying source quality. The 14–15% / 26% productivity figures aggregate heterogeneous study designs, not a single RCT.

Source: research/01-ai-native-landscape/stanford-ai-index-2026.md

AI Code Quality: The Hidden Cost Side of the Productivity Equation

Velocity gains from AI coding tools are well-documented; what is less visible is the structural degradation that follows. Three corpus files anchor the evidence on the cost side:

  • AI Code Rot (GitClear, 211M lines, 2020–2024): 8x increase in duplicated code blocks, 60%+ collapse in refactoring activity, 41% rise in code churn. Maintenance costs run 12% higher in year one; 4x traditional levels by year two. Forrester predicts 75% of technology decision-makers will face moderate-to-severe technical debt by 2026. The “18-month wall” — where delivery cycles stall as debugging becomes the primary bottleneck — is consistent across multiple independent datasets.

    Source: research/07-adoption-challenges/ai-code-rot-maintainability-costs.md

  • Production Failure Modes (Aikido Security, n=450 CISOs/developers, 2026): 20% of organizations suffered a serious security incident caused by AI-generated code; 69% discovered AI-introduced vulnerabilities in production. Apiiro’s Fortune 50 study: 4x developer velocity produced 10x more security findings (privilege escalation paths +322%, architectural design flaws +153%). 20% of AI-recommended packages do not exist — creating a “slopsquatting” supply chain attack vector. Gartner: prompt-to-app approaches will increase software defects 2,500% by 2028 without governance.

    Source: research/07-adoption-challenges/ai-code-production-failure-modes.md

  • SWE-CI Benchmark (Sun Yat-sen University / Alibaba, March 2026; 18 models, 100 repos, 233 days): 75% of AI models break previously working code during long-term maintenance. Only Claude Opus 4.6 exceeds a 50% zero-regression rate (0.76); all other model families score below 0.25. A model scoring 70%+ on SWE-bench (one-shot bug-fix) can still be a maintenance liability within weeks. Cognitive debt — AI generating code 5–7x faster than developers can comprehend it — is the least-measured failure mode.

    Source: research/07-adoption-challenges/swe-ci-ai-code-maintainability.md

The honest synthesis: velocity gains in months 1–3 are real and measurable. The structural costs in months 12–18 are also real and measurable. Both are true. A CTO who optimizes only for throughput speed is building a maintenance crisis. A governance framework that requires maintainability benchmarks (SWE-CI class) alongside velocity metrics is the intervention that holds both in balance.

Behavioral Sophistication Gap: UT Austin / KPMG (March 2026)

The largest behavioral study of real-world AI use finds that RCT productivity divergence may reflect behavioral sophistication as much as task type. Published HBR March 19, 2026 (n=2,597, 1.4 million actual interactions, 8 months).

  • Only ~5% of users consistently demonstrate sophisticated AI engagement across months of usage — despite ~90% using AI regularly.
  • Sophisticated users treat AI as a reasoning partner (iterating, assigning roles, demanding verification); routine users treat it as an answer machine (single queries, first-output acceptance).
  • The four measurable behavioral signals: (1) return frequency, (2) output refinement persistence, (3) ambition of initial request, (4) deliberate tool selection.
  • First-prompt length ranged from 26 to 48,670 characters; iteration ranged from 1 to 45 exchanges per conversation. The within-organization behavioral range exceeded the average difference between “AI user” and “non-user” categories.
  • This offers a complementary explanation for the METR −19% finding: if 95% of users in the study were routine users (single-query, first-output acceptance), measured task completion would underperform the 5% sophisticated users who iterate.
  • Most organizations measure adoption via login frequency and seat utilization — neither metric captures behavioral sophistication. This is why dashboards look healthy while financial returns remain elusive.
  • Credibility note: behavioral data from real interactions (not self-report); academic authorship (UT Austin McCombs); KPMG co-authorship introduces consulting-firm caveat; finding runs counter to KPMG’s commercial interest.

Source: research/07-adoption-challenges/kpmg-utaustin-sophisticated-ai-use-behaviors-2026.md

Task-Level Evidence Map: Which Tasks AI Helps, Hurts, or Leaves Unchanged

The BCG/Harvard “jagged frontier” framework (n=758 consultants, Sep 2023), METR RCT (n=16 developers, Jul 2025), and Faros AI throughput study (n=10,000+ developers, Jul 2025) together form a task-level evidence map:

  • First-draft text, idea generation, structured synthesis: AI reliably improves output quality (BCG: +40% quality on frontier tasks) and speed for individual contributors.
  • Experienced developers on complex, open-ended codebases: AI reliably slows work. METR: −19% measured vs. +20% perceived — a 39-point perception gap.
  • Organizational throughput: Individual task gains do not compound into team-level output. Faros: +21% tasks/developer, zero improvement in DORA metrics or quality, +91% review time.
  • Skill compression: The most consistent finding across studies — AI narrows the gap between low and high performers (BCG: +43% gain for bottom-half vs. negligible for top performers; NBER Brynjolfsson: +34% for novice CS agents).

The “AI helps on easy tasks and hurts on hard ones” framing is an oversimplification. The actual determinant is whether the task is inside or outside the AI’s capability frontier for the specific model, tool, workflow, and worker skill level — not apparent task difficulty.

Source: research/07-adoption-challenges/ai-task-performance-helps-vs-hurts.md

NVIDIA State of AI 2026 (n=3,200+, Aug–Dec 2025)

Not an RCT — a vendor-sponsored survey. Useful as a deployment signal, not an outcomes benchmark.

  • 88% report AI increased annual revenue; 87% report reduced costs. These figures conflict with the 0.6% measured gain in the Fed Atlanta CFO survey and McKinsey’s 6% high-performer finding.
  • The conflict is methodological: NVIDIA’s sample skews 40% AI practitioners (already committed to deployment) and is opt-in from the NVIDIA ecosystem. Self-selection inflates positive responses.
  • The challenge data (48% cite data management, 38% talent shortage, 30% unclear ROI) is the credible signal — it replicates consistently across independent studies and names the constraints that explain why 88% can claim gains while only 6% show EBIT impact.
  • Useful as investment-direction calibration (86% increasing budgets, 44% deploying agents) not ROI benchmarking.

Source: research/01-ai-native-landscape/nvidia-state-of-ai-2026.md

DORA: ROI of AI-Assisted Software Development (2026.01) — Task-Type Framework

Not an RCT — a framework report drawing on ~5,000 technology professionals surveyed in the 2025 DORA State of AI study. Credibility MEDIUM-HIGH; Google Cloud vendor caveat applies. Published April 2026.

  • Greenfield code (new, simple tasks): 35–40% productivity gain (Stanford Software Engineering Productivity programme, cited in report). Consistent with the BCG “frontier task” findings: AI helps substantially when the task is inside its capability range.
  • Legacy/complex code: ~10% or less productivity gain. This is where most enterprise engineering work lives. The gap between vendor benchmarks (which use greenfield conditions) and enterprise reality (which runs on 10–20-year-old codebases) explains most AI coding tool disappointments.
  • The J-Curve: Three causes of initial productivity decline — learning curve, verification tax (reviewing AI-generated code), downstream process adaptation (testing, deployment gates). Consistently appears across continuous delivery and platform engineering transitions too.
  • Instability tax: AI adoption associated with rising change failure rates. Sample model: failure rate rises 5% → 6%, producing −$344K downtime cost on a 500-person team. More code moving faster overwhelms existing pipelines.
  • Task throughput in high-AI environments: +33.7% tasks/developer, +66.2% epics/developer — but only in teams that built the surrounding infrastructure to absorb the throughput.
  • Central finding (corroborates METR, Faros, BCG, MIT CISR): “AI magnifies the strengths of high-performing organisations and the dysfunctions of struggling ones.”

Source: research/02-corporate-tools/dora-roi-ai-assisted-software-development-2026.md


Morgan Stanley AlphaWise — Organizational-Level Productivity Outcomes (May 2026)

Not an RCT — a cross-sectional survey of organizations already using AI for 12+ months (n=935, 5 countries, 5 sectors). Credibility MEDIUM; vendor-interest caveat applies (Morgan Stanley covers AI-exposed sectors for institutional investors). Useful as the largest-sample concurrent productivity-and-headcount measurement available at organizational scale.

  • Organizations using AI for 12+ months report 11.5% net productivity gain alongside a 4% net headcount decline — both measured in the same organizations over the same period.
  • The productivity-headcount ratio matters for productivity RCT interpretation: task-level RCTs measure productivity gain in isolation. This survey measures where the organizational dividend actually goes — and a meaningful share went to cost reduction, not reinvestment.
  • The US is the only surveyed country showing net employment gain (+2%). This divergence limits generalizability of any single-country RCT to global enterprise deployments.
  • Selection bias caveat: restricting to organizations already using AI for 12+ months overstates mature-adopter gains vs. the broader enterprise population.

Source: research/07-adoption-challenges/morgan-stanley-ai-efficiency-paradox-2026.md · May 2026 · MEDIUM · TIER 1


HBR — The Competence Penalty: A Social Constraint on Adoption (August 2025)

Not an RCT in the conventional sense but a pre-registered experiment measuring the social perception effect of AI use — a mechanism that shapes who participates in the productivity RCT population.

  • Pre-registered experiment: n=1,026 engineers reviewed identical Python code described as either human-written or AI-assisted. Code quality ratings unchanged. Competence ratings fell 9% on average when AI was disclosed.
  • Observational study: n=28,698 software engineers, 12 months post-AI-assistant rollout. Only 41% had tried the tool. Female adoption: 31%. Engineers 40+: 39%.
  • Gender amplification: Female engineers rated 13% lower competence (vs. 6% for males). Male non-adopters most severe: penalized female AI users 26% worse.

Why this matters for RCT interpretation: Task-level productivity RCTs measure workers who use AI. They say nothing about the social selection process that determines who uses AI on work others will see. If AI adoption suppresses perceived competence, the workers with the most to lose (senior contributors, women, those in evaluation cycles) have a rational incentive to avoid AI on visible work — limiting their contribution to the productivity baseline. This is a partial explanation for why individual-level productivity gains (Brynjolfsson: +15%; Sadun: team-equivalent output) do not aggregate to organization-level financial returns.

Source: research/07-adoption-challenges/hbr-psychological-costs-ai-adoption-2026.md — TIER 2, HIGH (pre-registered, August 2025)


Goldman Sachs — Economy-Wide AI Productivity Analysis (Feb–May 2026)

Not an RCT — macroeconomic analysis and S&P 500/Russell 3000 earnings call review by Goldman Sachs economists (Jan Hatzius, Ronnie Walker, Joseph Briggs, Sarah Dong). Highest-authority investment bank macro lens on whether AI productivity claims are appearing in aggregate data. Credibility HIGH; Goldman commercial interest (investment banking relationships with AI sector) runs counter to bearish findings, adding credibility.

  • No economy-wide relationship found: Goldman economists find “no meaningful relationship between productivity and AI adoption at the economy-wide level” as of early 2026 earnings data.
  • U.S. GDP contribution: AI investment spending added “basically zero” to U.S. GDP in 2025. Mechanism: hardware imports from TSMC/Samsung transfer the domestic demand multiplier to Taiwanese and Korean GDP (Jan Hatzius, Atlantic Council, Feb 23, 2026).
  • The 30% figure is real but narrow: Companies that measured AI impact reported a median 30% productivity gain — but only in two specific use cases: customer support and software development. These are also the use cases with the clearest structured-input / structured-output architecture.
  • Measurement gap is the explanation: Only 10% of S&P 500 companies discussing AI on earnings calls quantified its impact on specific use cases. Only 1% quantified impact on earnings. The gap between 70% of companies discussing AI and 1% quantifying earnings impact is the executive perception gap.
  • Adoption reality: Census Bureau BTOS data shows 19% of U.S. establishments have AI adopted for business functions (March 2026). Large firms (250+): 35.3%. Survey data showing “70% use AI” measures individual employee tool access, not organizational deployment.
  • FOMO as investment driver: Goldman research (May 2026) found competitive insecurity — not ROI signals — is the primary driver of AI investment decisions. This matches the NBER w34836 finding (90% no past impact; same executives predict future gains).

Cross-reference: NBER w34836 (90% no past impact, n=6,000 executives) and McKinsey State of AI 2025 (6% EBIT-impact cohort) both measure the same productivity-visibility gap from different angles. Goldman adds the macro/financial-market lens.

Source: research/01-ai-native-landscape/goldman-sachs-ai-economic-research-2026.md · Feb–May 2026 · HIGH · TIER 1


Stanford SIEPR Consumer AI Productivity (2026)

Not an RCT — population-scale natural experiment using behavioral internet browsing data. Blank (Stanford SIEPR/GSB), Schubert (UCLA Anderson), and Zhang (USC Marshall) tracked 200,000+ U.S. households before and after ChatGPT’s November 2022 launch, using pre-ChatGPT browsing patterns as an instrumental variable for adoption. This is the largest behavioral dataset applied to AI productivity research to date; no self-report, no vendor interest.

  • 76–176% efficiency gains on consumer digital chores: job hunting, travel planning, product research, troubleshooting. For specific tasks (plumbing diagnostics, tax questions, product comparisons), gains exceed 500–1,000%.
  • Leisure substitution, not upskilling: Freed time went to leisure browsing (+31 pp share among adopters), not education or skill development. Productive browsing share fell 21 pp.
  • Total productive time unchanged: ChatGPT adoption did not increase total productive task time — users completed tasks faster, then stopped productive tasks. Output volume did not rise proportionally to efficiency.
  • Digital divide is widening: Younger, higher-income households adopted substantially faster. Older, lower-income households continued traditional browsing patterns without closing the gap over the 2021–2024 observation window.
  • Consumer context ≠ enterprise context: Gains are large for bounded, information-seeking tasks with low coordination overhead and no integration requirement. Enterprise gains require workflow redesign, training, and change management — all absent in consumer home use.

Cross-reference: Pairs with Goldman Sachs (no economy-wide gain; 30% concentrated in customer support + software dev) and BIS/EIB (small firms <50 employees see ~0% gain). Together these three studies establish that productivity gains from AI are real but highly context-dependent — and concentrated where the workflow conditions support them.

Source: research/01-ai-native-landscape/stanford-siepr-consumer-ai-productivity-2026.md · April 2026 · HIGH · TIER 1


Workday / Hanover Research — AI Rework Tax (Nov 2025)

Not an RCT — vendor-commissioned large-sample survey. Hanover Research conducted independent fieldwork (n=3,200 active AI users at $100M+ revenue organizations, North America/APAC/EMEA, November 2025). Workday has direct commercial interest in workforce platform investment findings. The “self-damning” 14% net-positive figure adds credibility — vendor-commissioned surveys rarely foreground findings that undermine adoption narratives. Directionally corroborated by METR 2025 RCT (19% slower) and Goldman Sachs macro analysis (no economy-wide productivity gain). Credibility: MEDIUM-HIGH; TIER 2.

  • 14% of employees consistently achieve net-positive outcomes from AI use — after accounting for rework time (correcting errors, rewriting drafts, verifying outputs).
  • 40% of AI time savings are consumed by rework. For every 10 hours saved, approximately 4 hours are lost to correction and verification. Workday terms this the “AI tax on productivity.”
  • 85% report saving 1–7 hours per week — the metric organizations typically cite as evidence the investment is working. The 40% rework loss is what most AI dashboards never capture.
  • 79% of the net-positive 14% cohort received increased skills training; only 37% of heavy AI users overall received it — the gap is training and role redesign, not technology.
  • 89% of organizations updated fewer than half their roles to reflect AI-augmented work. Tool deployed; job not redesigned.

Cross-reference: METR 2025 RCT (experienced developers 19% slower despite believing they were 20% faster — same dynamics, controlled evidence). Goldman Sachs (no economy-wide productivity gain). BCG “brain fry” study (40% of AI-heavy workers experience cognitive strain). Together these establish that self-reported productivity gains and measured net outcomes diverge systematically.

Source: research/07-adoption-challenges/workday-beyond-productivity-ai-rework-2026.md · Nov 2025 fieldwork / May 2026 write-up · MEDIUM-HIGH · TIER 2

Morgan Stanley AI Adoption Survey 2026 — Sector-Level Productivity and Employment Evidence

Survey of n=935 corporate executives in the US, Germany, Japan, and Australia (5 AI-exposed sectors: automotive, healthcare equipment, consumer staples, retail/consumer, real estate/transport). Firms that had been using AI for 12+ months. Published February 5, 2026. Named analysts: Michelle Weaver and Stephen Byrd. Source credibility: MEDIUM (investment bank, survivorship-biased sample limited to active deployers, self-reported; directionally consistent with BIS/EIB and NBER w34984). TIER 1.

  • 11.5% average net productivity gain among companies using AI 12+ months — but this is a survivorship-biased sample (non-deployers excluded). Distribution: ~50% in 1–10% range; ~33% in 11–20% range; 14% above 20%. Healthcare sector highest gains; real estate lowest.
  • 4% net headcount decline globally — the employment effect is quiet, not dramatic: 11% jobs eliminated, 12% vacancies left unfilled, 18% new hires created. The mechanism is attrition management and non-backfill, not mass layoffs.
  • UK worst among peer nations at −8% net jobs. US the only country with slight net job growth. Germany suppressed on both sides by labor regulation.
  • Entry-level roles bear the highest elimination risk. Mid-career workers (2–10 years experience) are being retrained — 27% retrained in the past year. The Korn Ferry pipeline-risk finding (37% eliminating entry-level roles, creating future leadership gap) is corroborated.
  • The “efficiency paradox”: Productivity gains are real in the high-deployer cohort, but job creation does not follow at rates comparable to prior technology cycles. Morgan Stanley’s own analyst Byrd described the net job loss magnitude as “surprising” and called the data an “early warning.”

Cross-reference: BIS/EIB (n=12,000+ EU firms, IV-identified, +4% productivity, no employment effect at aggregate) and NBER w34984 (748 CFOs, −0.37% aggregate employment, 3x perception gap) are the independent academic checks. Together they bound the Morgan Stanley figures: the 11.5% productivity gain is at the high end because of survivorship selection, and the employment effect is real but lower than Morgan Stanley’s sample implies at the economy-wide level.

Source: research/05-analyst-firms/morgan-stanley-ai-adoption-survey-2026.md · Feb 2026 · MEDIUM · TIER 1

METR Time Horizons Benchmark — Autonomous Task-Completion Capability Trajectory (TH1.1, Jan 2026)

Independent benchmark from METR (Model Evaluation & Threat Research), no commercial AI interest. arXiv:2503.14499. 228 real tasks across software engineering, ML, and cybersecurity. Human baseline: skilled professionals (5+ years experience). Source credibility: HIGH. TIER 1 (TH1.1 January 2026, dashboard updated May 2026).

  • Frontier AI (Claude Opus 4.5) achieves 50% success on tasks taking a skilled human 5.3 hours. Claude Sonnet 3.7 was at 1 hour in early 2025. Two frontier generations represent a 5x expansion in autonomous task-completion range within approximately one year.
  • Doubling time is 89–131 days (since 2024 and 2023 respectively) — faster than the original 7-month (210-day) estimate from TH1.0 (March 2025). The trajectory is accelerating, not plateauing.
  • The benchmark is bumping against its own ceiling. METR doubled the number of 8+ hour tasks from 14 to 31 in TH1.1 because frontier models are outperforming the original suite. Models above 16-hour task horizons are now listed as “unreliably measured” — the next measurement challenge.
  • Extrapolation (from TH1.0): Week-long autonomous tasks arrive within 2–4 years if trends hold. Month-long autonomous projects by end of decade. At 89-day doubling, the week-long threshold arrives faster.
  • Implication for RCT research: METR’s own RCT design is breaking down because 30–50% of experienced developers now refuse to work without AI tools. The “no AI” control condition is becoming unenforceable. This is corroborating evidence that adoption has moved from opt-in to infrastructure for technical workers.

Cross-reference: METR 2025 RCT (19% slower, July 2025) measured a specific population at a moment when time horizons were under 1 hour. The task-level autonomy data explains why that finding is period-specific: the models used in that RCT operated in a fundamentally different capability regime than current frontier models.

Source: research/12-agent-workers/metr-time-horizons-autonomous-ai-capability-2026.md · TH1.1 Jan 2026, dashboard May 2026 · HIGH · TIER 1

Stanford CS224N 2024 — Benchmark Fragility and Evaluation Reliability Evidence

Academic lecture content from Stanford’s NLP flagship course (CS224N Spring 2024). Speaker: Yan, 3rd-year Stanford PhD student advised by Tatsu and Percy Liang. Credibility: HIGH for quantified findings with named institutional affiliation. TIER 3 (2024 academic lecture; evaluation methodology insights remain structurally valid as the benchmarks described are still in active use).

  • MMLU had three different implementations coexisting for nearly a year, producing a 15-point score spread on the same model. Llama 65B ranged from 48.8% to 63.6% depending on which implementation was used. This is the most widely cited LLM benchmark — a 15-point swing represents the difference between “competitive with GPT-3.5” and “approaching GPT-4” on public leaderboards.
  • GPT-4 shows a complete performance cliff on CodeForces problems released after its training cutoff: 10/10 on pre-2021 problems, 0/10 on post-2021 problems. This is direct evidence of benchmark contamination rather than genuine reasoning transfer.
  • 70% of top AI conference papers evaluate only on English-language tasks; 40% measure only accuracy. This creates a publication monoculture that systematically overstates generalization.
  • Implication for enterprise buyers: Published benchmark scores are not reliable proxies for task performance. The same model can score 15 points higher or lower on “the same benchmark” depending on implementation. Enterprises should require task-specific evaluation on representative samples of their actual workload — not vendor-provided leaderboard rankings.

Source: research/13-multimodal-sources/stanford-cs224n/2026-05-18-stanford-cs224n-llm-evaluation-benchmarking.md · Stanford CS224N Spring 2024 · HIGH · TIER 3

NBER WP35046 — Forecasting the Economic Effects of AI: Diffusion Lag as the Productivity Puzzle (April 2026)

Source: research/01-ai-native-landscape/nber-karger-forecasting-economic-effects-ai-2026.md · NBER Working Paper 35046 · Karger, Kuusela, Abaluck, Tetlock et al. · n=160 specialists + 401 general public · October 2025–February 2026 · HIGH / TIER 1

Multi-group forecasting tournament across economists, AI industry professionals, policy researchers, and superforecasters. Directly relevant to interpreting the RCT productivity literature: the capability-vs-impact gap that individual RCTs capture is a microeconomic manifestation of the macroeconomic diffusion lag this study quantifies.

  • 61.4% of specialist economists expect moderate or rapid AI progress by 2030, yet the unconditional GDP forecast is only 2.5% — barely above government baselines. The individual-task gains measured in RCTs (25% quality: AISI; +50% output: Ju/Aral; +15%: Brynjolfsson/Li/Raymond) are not yet aggregating to economy-wide productivity signal. This study provides the structural explanation: diffusion lags of the type seen in electrification and personal computing.
  • Superforecasters (most accurate in the sample) assigned 45% probability to the “slow AI” scenario — higher than any other group. Independent forecasting skill, not AI domain knowledge, produces the most conservative near-term outlook.
  • The capability-forecast/GDP-forecast gap is not contradiction — it is the core finding for enterprise planning. Organizations using AI capability announcements as a proxy for economic impact are working with the wrong variable. The RCT evidence on individual-task gains is real; the macro translation takes decades, not quarters.
  • Cross-reference: NBER WP 34984 (748 CFOs, 1.8% perceived vs. 0.6% measured productivity gain) provides the enterprise-level version of this same gap. The forecasting tournament provides the macro framing; the CFO survey provides the firm-level measurement.

Microsoft NFOW 2025 — Observational Evidence on Skill Degradation and Team-Level Productivity

Source: research/07-adoption-challenges/microsoft-new-future-of-work-2025.md · Microsoft Research MSR-TR-2025-58 · 2025 · MEDIUM-HIGH / TIER 2

Not RCTs themselves, but the NFOW 2025 report aggregates observational evidence on how productivity measurements fail at the team/org level — directly relevant to interpreting the RCT corpus:

  • 60% of employees skip accuracy checks on AI output despite 67% reporting they trust AI agents (Benzing et al., 2025, n=1,800, Microsoft Agentic Team AI & Trust Report) — the gap between stated trust and practiced verification that RCTs don’t capture
  • Skill degradation (clinical peer-reviewed): Colonoscopy clinicians showed measurable decline in independent polyp-detection ability after just three months of AI-assisted practice (Budzyn et al., 2025, The Lancet Gastroenterology & Hepatology) — the first peer-reviewed clinical deskilling finding; directly relevant to interpreting long-run RCT productivity gains vs. skill formation
  • Group facilitation gains ≠ outcome gains: An AI facilitator boosted information sharing in group tasks by 22%, but neither human nor AI facilitators changed group decisions (Alsobay et al., 2025, Microsoft Study) — consistent with Ju & Aral’s finding that 50% more output from H-AI teams produced zero market advantage
  • Software engineers spend only 15–25% of their time writing code (Kumar et al., 2025; Meyer et al., 2017, 2019) — reframes most coding RCTs: the measured task (code writing) is not the primary productivity bottleneck, so gains there need not convert to overall SDLC improvements

Harness / Sapio Research 2026 — The Measurement Infrastructure Breakdown

Source: research/01-ai-native-landscape/harness-engineering-excellence-2026.md · Harness/Sapio Research, n=700 engineering practitioners and managers, April 2026 · MEDIUM / TIER 1

This survey does not measure productivity directly — it measures the organizational measurement infrastructure that organizations use to assess AI’s productivity impact. The core finding: that infrastructure is broken in ways that organizations acknowledge but haven’t acted on.

  • 89% of engineering leaders say current metrics accurately reflect AI’s impact; 94% simultaneously say key factors (tech debt, validation time, burnout) are missing from those metrics. Both cannot be true. The gap describes confident but structurally incomplete information used for multi-year investment decisions.
  • 81% of leaders say code review time increased post-AI adoption; 28% say the increase exceeds 30%. This aligns with Faros AI’s finding (98% more PRs, zero delivery improvement): AI moved the bottleneck from code generation to review and integration — and most dashboards track generation, not integration.
  • ~31% of developer time is “invisible work” — reviewing AI-generated code, fixing AI-introduced bugs, context switching — that doesn’t register in standard productivity metrics. GitClear’s codebase data (211M lines, 153 repos) provides independent corroboration: code churn up 39%, useful code per developer down 59%.
  • The manager-developer comfort gap on measurement: 15% of managers vs. 4% of developers report no concerns about AI productivity data usage. Managers are 4x more comfortable with the same measurement infrastructure that developers distrust. 54% of developers fear individual performance evaluations based on AI metrics.
  • Implication for the RCT corpus: The METR July 2025 RCT (−19% for experienced developers) and AISI February 2026 RCT (+25% quality, +61% throughput) measured specific task types under controlled conditions. The Harness finding is that real-world organizations are not measuring comparable task types — they’re measuring PR volume and code generation velocity while invisible review work accumulates. This explains part of the self-report/RCT gap that METR documents: workers report 2x value gains while RCTs show more modest outcomes because organizations have shifted to measuring the part of the workflow where AI looks good.

ADP Global Workforce Survey 2026 — The Productivity Perception Gap at Scale (n=39,000)

Source: research/07-adoption-challenges/adp-people-at-work-ai-productivity-paradox-2026.md · ADP Research, n=39,000 working adults, 36 markets, fieldwork July–August 2025, published March 25, 2026 · HIGH / TIER 1

The largest workforce dataset in the 2026 corpus adds a critical piece to the RCT interpretation: the AI productivity paradox at scale, from the worker perception side.

  • Daily AI users are 4x more likely to feel less productive than they could be, despite being 2x more engaged (30% vs. 14% fully engaged) and half as stressed (11% vs. 23% experiencing high stress overload). The dissociation between engagement/stress outcomes and productivity perception is the core finding.
  • ADP’s interpretation: “Productivity might in fact have increased, but people might perceive that they are doing less of the work themselves.” This is the perception-side mechanism for the self-report/RCT gap: volume-based self-assessment penalizes AI-assisted workers because the visible effort (keystrokes, hours-to-draft) has moved to the AI, while the human contribution has shifted to oversight, judgment, and direction — which workers don’t count as “productive” work.
  • Implications for measurement: Organizations using task-completion volume, word count, or first-draft speed to measure AI productivity will systematically underrate their AI programs. The Harness finding (31% of developer time is invisible work post-AI) and the ADP finding (daily users feel less productive) are two sides of the same measurement failure — the metrics are calibrated to pre-AI work patterns.
  • Cross-calibration with METR RCT: METR measured −19% task-count productivity for experienced developers. ADP measures a 4x increased perception of underperformance. Neither finding indicates AI failure; both indicate that standard productivity metrics — whether objective (task count) or subjective (self-report) — fail to capture what AI-augmented work actually produces.

NBER WP 34851 — Cruces et al.: The Education Gap RCT (February/May 2026)

Source: research/07-adoption-challenges/nber-w34851-cruces-genai-education-productivity-gap-2026.md · NBER WP 34851; preregistered RCT, AEA-RCT #0016607; n=1,174, Argentina, GPT-4.1, Feb 2026/revised May 2026 · HIGH / TIER 1

The first RCT to explicitly measure AI’s effect on education-based productivity gaps (not within-occupation performance variation). Design: workers recruited outside firms to avoid organizational compression of educational heterogeneity; incentivized business problem-solving task (email reply requiring multi-source analysis); GPT-4.1 assistant for treatment group; non-AI follow-up module for all participants.

  • Baseline gap (no AI): High-education participants outperform low-education by 0.548 SD on the business task.
  • Gap with AI access: 0.139 SD — approximately 75% closure. Low-education workers gained +1.242 SD vs. high-education +0.834 SD.
  • Gains are not pure delegation: Low-education treated participants scored +0.171 SD above their controls in the non-AI follow-up module (significant). The carry-over is real, but a residual 0.200 SD gap versus high-education treated participants persists — underlying human capital still matters for unassisted work.
  • Engagement determines whether gains last: High-AI-assistance + high-engagement users: 0.732 SD follow-up score. High-AI-assistance + low-engagement (pure delegators): 0.065 SD. This is the most actionable finding for enterprise training design — the tool is not the intervention; the engagement model is.
  • Mechanism: Lower-education workers obtained more AI assistance (10 pp more likely to copy AI output directly); higher-education workers used the tool more effectively (structured workflow, specific questions, reviewing rather than delegating). Equal access ≠ equal effective use.
  • Cross-calibration: Consistent with Dell’Acqua et al. BCG (larger gains for lower-skilled consultants) and Brynjolfsson et al. customer service (larger gains for novices). This study extends those findings outside the firm and outside a single occupation, targeting the broader educational distribution.

Addition to RCT table:

Study Date n Population Model Era Headline Tier
Cruces et al. (NBER WP 34851) Feb/May 2026 1,174 General adults 25–45, Argentina (mixed education) GPT-4.1 75% education gap closure; engaged users +0.732 SD follow-up vs. delegators +0.065 SD TIER 1

Epoch AI / Ipsos (n=2,021, March–April 2026) — Population-Level Behavioral Usage Data

Source: research/07-adoption-challenges/epoch-ai-ipsos-ai-workplace-usage-2026.md · Epoch AI / Ipsos KnowledgePanel, probability-based, n=2,021, March–April 2026 · HIGH / TIER 1

The first probability-based (non-convenience) national survey on AI workplace usage provides population-level context for interpreting RCT findings:

  • 80% of AI users’ top use case is information lookup; 59% is writing/editing text. Only 37% use AI for data analysis or programming — the task categories where RCTs have documented the clearest productivity gains. The METR −19% RCT and AISI +25% quality RCT were testing developer/analytical tasks. The population-level data shows most workers are using AI for tasks where productivity evidence is weakest.
  • 27% of employed AI work users report AI replaced existing tasks; only 21% report AI created new tasks. Task displacement is outpacing task creation — but neither the RCTs nor the enterprise surveys track this displacement-to-creation ratio at scale.
  • 76% of employer-provided subscribers use AI primarily for work vs. 38% of free-tier users. The quality of enterprise AI deployment (employer provision + integration) determines whether workers even attempt the high-leverage use cases that RCTs test. Most enterprise AI programs are measuring adoption rates, not use case distribution — so they cannot tell whether their workforce is in the 37% (data analysis) or the 80% (information lookup) band.
  • Cross-calibration with RCT corpus: The persistent gap between RCT findings (mixed) and self-report surveys (positive) partially resolves here: most self-report surveys sample employed AI users, who skew toward the 76% employer-provided subscriber segment. Most RCTs test specific high-leverage tasks. The population data shows most workers are not doing those tasks with AI — they are doing lookups and editing. Productivity impact expectations should be calibrated accordingly.

Dillon, Jaffe, Immorlica & Stanton (Microsoft Research / HBS) — Largest Cross-Firm RCT: Tool Access Doesn’t Shift Tasks (n=7,137, Nov 2025)

The largest enterprise-deployment RCT on Microsoft 365 Copilot. 66 large firms, 7,137 knowledge workers, 6-month randomized experiment (Sep 2023–Oct 2024). Three of four authors are Microsoft Research employees; Christopher Stanton is Harvard Business School/NBER.

  • Email time reduction: −2 hours/week (LATE = −17%). Baseline was 11.7 hours/week in Outlook. The effect is statistically significant (q<0.01), persistent across all six post-period months, and driven primarily by Outlook Copilot use. Workers read 6% fewer emails with no increase in time-per-email.
  • No task composition shift. Same number of email threads replied to (rules out >3.4% change). Same Teams meetings attended (rules out >4.6% meeting count increase). Same Word documents completed. Workers freed up time from email but did not convert it into measurable new output.
  • After-hours work reduced by 9%. Some of the recovered time went to shortening the workday, not to new work.
  • Firm-level variation is the dominant factor. Firm fixed effects explain nearly twice as much variation in individual Copilot usage as individual work patterns. Weekly usage ranged from 6% to 70% across firms — a 10-fold gap unexplained by industry or job type. Managerial practice and organizational culture are the implied drivers.
  • Co-invention is needed but didn’t happen. Point estimates suggest 50% more email time savings when a close coworker also had Copilot, but this is imprecisely estimated (cannot rule out zero). Larger task shifts require coordinating with colleagues on new norms — which this individual-license rollout did not generate.
  • Limitation: Study period is Sep 2023–Oct 2024 — 2023-era Copilot, not current generation. Control workers had access to ChatGPT throughout. The study measures marginal benefit of integrated Copilot over and above general AI availability.
  • Calibration note: This is a TIER 2 study on a prior-generation tool. Current Copilot (2025+) may produce different results. The “no task shift” finding is robust to the study design but may not hold with more capable models and explicit workflow redesign programs.

Source: research/07-adoption-challenges/dillon-jaffe-shifting-work-patterns-copilot-rct-2025.md · Dillon, Jaffe, Immorlica & Stanton · arXiv 2504.11436v4, November 2025 · MEDIUM-HIGH / TIER 2


Ni, Wang, Feng, Lu et al. (Fudan/Zhejiang/Dartmouth Tuck/Alibaba) — The Skill Inversion in Customer Service (November 2025)

The largest published GenAI customer service RCT. Alibaba’s Taobao after-sales division, 5,940 new agents, 2.56 million chats, four-week A/B test (January 2024). AI system: Alibaba Qwen LLMs fine-tuned for Taobao service scenarios. Independent academic team with full process-level data access (messaging logs, timing, linguistic patterns, multitasking behavior).

  • Average results are positive but hide a reversal. ITT: issue identification time −8.2%, dissatisfaction −3.4%, ratings +0.042 (all significant). At 100% LATE usage, effects scale 4×.
  • Bottom quintile (Q1) gains are large. At full usage: chat duration −14.8%, customer rating +2.144 points, dissatisfaction −0.567. Pre-post absolute: rating rises from 2.184 to 2.981.
  • Top quintile (Q5) gets measurably worse. Rating −1.514 pts, dissatisfaction +0.374, retrial rate +4.1 pp — the only quintile where objective quality (actual issue resolution) declines.
  • Mechanism identified: NOT overreliance. Top performers use AI less (usage rate 0.18 vs. 0.39 for Q1). Linguistic analysis confirms their message quality did not deteriorate. The driver is task-switching overhead — AI adds verification burden that disrupts the workflow of agents managing multiple concurrent chats, increasing shift-away time by 26.3% for Q5 vs. −31.2% for Q1.
  • Satisfaction ≠ resolution. Customer ratings improved across Q1–Q4; retrial rates (objective measure) did not improve for any group, and worsened for Q5. Organizations measuring only CSAT will over-claim success.
  • Design implication: Broad rollouts produce a performance averaging effect — raising the floor while depressing the ceiling. Role differentiation by skill tier is not an optional refinement; it is the difference between AI that improves the organization and AI that flattens it.
  • Comparison with Brynjolfsson et al. (QJE 2025): Both studies find top performers decline. Brynjolfsson conjectured overreliance as the mechanism. This study directly tests and refutes that hypothesis — identifying task-switching as the mechanism in B2C concurrent-chat environments.

Source: research/12-agent-workers/ni-wang-alibaba-genai-customer-service-rct-2025.md · Ni, Wang, Feng, Lu, Wang, Zhou · arXiv:2603.29888, November 2025 · HIGH / TIER 1


Hidden Costs of AI Coding Tools — Faros AI / Veracode / IBM (2024–2025)

Multi-source analysis of the real cost structure of AI coding tool adoption. Synthesis of Faros AI engineering telemetry (10,000+ developers), Veracode code quality analysis (100+ LLMs, 80 tasks), IBM Cost of a Data Breach (n=600), DX Research practitioner survey, and GitClear 211M-line analysis (2020–2024). TIER 2–3.

  • License fees represent 40–60% of actual first-year costs. Implementation costs exceed licensing by 30–40%; organizations routinely exceed initial AI tool budgets by 30–50% in Year 1.
  • The code review bottleneck erases speed gains at the team level. High-AI-adoption teams merge 98% more PRs, but review time increases 91% — with no net improvement in throughput metrics. AI moved the bottleneck; it did not remove it.
  • AI-generated code carries a quality tax. 1.7× more issues per AI-co-authored PR (CodeRabbit, 470 PRs); 45% of AI-generated code chooses the insecure implementation path (Veracode, 100+ LLMs). Quality debt compounds — maintenance costs reach 4× traditional levels by Year 2 in the GitClear data.
  • Shadow AI adds $670,000 to breach costs. One in five organizations experienced a breach due to unsanctioned AI tool use; 63% lack any AI governance policy (IBM, n=600, 2025).
  • Directionally consistent with METR RCT: the mechanism for cost overrun is the same as the mechanism for productivity shortfall — hidden verification and governance overhead, not tool capability.

Source: research/02-corporate-tools/hidden-costs-beyond-licensing.md · Faros AI, Veracode, IBM, DX Research, GitClear · 2024–2025 · HIGH–MEDIUM / TIER 2–3


Gartner — Code Quality Governance Framework (March 2026)

Gartner “Predicts 2026: AI Potential and Risks Emerge in Software Engineering Technologies.” Independent analyst forecast; directional, not empirical.

  • 25% of production defects will stem from inadequate human oversight of AI-generated code by 2027, up from <1% in 2023. The 2,500% figure applies only to prompt-to-app citizen-developer scenarios without governance.
  • Automation bias is the mechanism, not AI capability failure. Developers implicitly trust AI suggestions based on surface correctness without architectural scrutiny.
  • Four-layer governance model prescribed: pre-commit quality gates → PR-level architectural review → CI/CD pipeline enforcement → post-deployment monitoring. Organizations implementing all four layers retain the productivity gains; those skipping them face defect backlogs.
  • Review bottleneck is the hidden cost. AI accelerates code generation 5–10×; human review capacity stays flat. QA pipelines built for human-paced change become the new constraint.
  • Directionally consistent with METR RCT finding (experienced developers 19% slower) — both point to review/verification overhead as a primary drag on AI coding productivity.

Source: research/05-analyst-firms/gartner-ai-code-quality-governance.md · Gartner Predicts 2026 · MEDIUM / TIER 2


Humlum & Vestergaard (NBER WP 33777) — Still Waters, Rapid Currents: The Largest Administrative-Linked AI Labor Market Study (March 2026)

The largest study linking AI adoption survey data to administrative payroll records. University of Chicago Booth and University of Copenhagen. n=25,000 workers, 7,000 workplaces, 11 highly exposed occupations in Denmark, tracked through December 2024 — two full years after ChatGPT’s November 2022 launch. HIGH credibility / TIER 1. Revised March 2026.

The null result:

  • Difference-in-differences estimates on earnings and recorded hours are precisely zero — confidence intervals rule out effects larger than 2% at both worker and workplace levels.
  • Null results hold for daily users, workers reporting >1 hr/day time savings, workers with full employer investment (enterprise chatbot + training + encouragement), and workers across all 11 occupations.
  • Early-career employment has declined in exposed occupations, but adopting firms are not driving those declines.

The adoption reality:

  • 43% of employers explicitly encourage chatbot use; only 6% prohibit it.
  • Full-package workplaces (encouraged + enterprise + training): 93% adoption, 28% daily use, 19% saving >1 hr/day.
  • Baseline adoption without any employer initiative: 41%.
  • Employer encouragement — not tool access — is the primary driver of intensive use and reported benefits.

New task creation (the reorganization layer):

  • 8%–18% of users take on entirely new tasks depending on employer initiative level.
  • New tasks split: 42% content generation, 35% AI quality oversight and compliance, 26% AI integration into workflows.
  • 4% of non-users also report new AI-related workloads — spillover to the entire team, not just adopters.
  • 85% of chatbot users reallocate time savings to more work, not leisure.

The one variable that moves:

  • Chatbot adopters are more likely to switch occupations. By December 2024: +4% of FTE in most recent occupation (statistically significant; flat pre-trend confirms causality).
  • Job switchers move into roles with higher wage premia and higher AI relevance.
  • Job switchers see earnings grow 12 percentage points faster than average Danish workers.
  • Effect concentrated in occupations with individual tool discretion (IT support, office clerks); absent in licensed professions requiring credentials (teachers, accountants).

Calibration against the rest of the corpus:

  • Consistent with Suh & Oh (Korea, n=5,512): real time savings, not converting to measured output gains.
  • Consistent with Dillon et al. (n=7,137 Copilot RCT): no task composition shift from tool access alone; firm management is the primary adoption driver.
  • Explains METR paradox differently: METR finds developers 19% slower on open-ended tasks; Humlum/Vestergaard find zero earnings effect on software developers. Both are consistent with the “J-curve trough” — real work is changing; economic statistics have not yet caught up.
  • Extends the productivity J-curve framework (Brynjolfsson, Rock, Syverson, 2021) empirically: workplaces are in the organizational investment phase that precedes measurable aggregate gains.

Limitation: Denmark-specific wage institutions (collective bargaining in some occupations, administrative payroll records) may not generalize directly to U.S. markets. Authors note Danish AI adoption rates are comparable to U.S., and labor market flexibility is similar.

Source: research/07-adoption-challenges/humlum-vestergaard-nber-w33777-still-waters-rapid-currents-2026.md · Humlum & Vestergaard · NBER WP 33777 · May 2025, rev. March 2026 · HIGH / TIER 1


White House CEA ERP 2026 — Macroeconomic Labor Picture

Source: Council of Economic Advisers, White House, April 2026. Synthesis of Brynjolfsson et al. (2025), Johnston & Makridis (2025), METR (2025), Mousa (2025). MEDIUM-HIGH.

The ERP Chapter 5 synthesizes the contradictory labor RCT evidence and offers a macroeconomic framing:

  • Mixed signals are real: some studies find early-career employment declining in AI-exposed occupations (Brynjolfsson, Chandar, Chen 2025); others find no correlation with current unemployment; Johnston & Makridis find employment increasing in sectors where AI complements human tasks
  • Jevons’ Paradox is the CEA’s organizing frame: efficiency gains reduce per-unit labor needs but can increase total labor demand by lowering prices and expanding markets (confirmed historically in energy, agriculture, transportation)
  • Radiologist case: once the canonical AI-displacement example; now at historically high employment rates — the canonical Jevons’ Paradox AI example
  • U.S. unemployment as of December 2025: 4.4% — no macroeconomic displacement signal in aggregate data yet
  • METR task-length data (cited in ERP): AI task completion length (50% success rate for software engineering) is doubling every 7 months for 6 years — SWE-bench went from 4% (2023) to 72% (2024) success rate

The CEA does not resolve the debate — it documents it. Both the “AI is displacing work” and “AI is expanding work” narratives have credible evidence at the sectoral level.

Source: research/01-ai-native-landscape/cea-erp2026-ch5-ai-revolution.md


Beyond the Pilot — Practitioner voices on the jagged frontier (April 2026)

Corroborating practitioner perspectives from VentureBeat’s Beyond the Pilot series:

  • Amjad Masad, CEO, Replit (April 14, 2026): “There are two things that are absolutely working. Two roles that are getting automated or augmented… those are support… and then software. Outside of that, there’s really nothing that is working. There are a lot of toys, a lot of experiments.” This maps directly onto the RCT corpus: the two use cases where Goldman Sachs finds a median 30% gain (customer support + software development) are the same two Masad identifies as the only reliable production categories. Every other category is in the “toys and experiments” zone the METR −19% and Suh & Oh 0.008 correlation findings describe.

  • Andrew Ng, Founder, DeepLearning.AI (April 13, 2026): Frames agentic AI as the “next enterprise S-curve” — compound productivity gains require multi-step workflow decomposition, not single-model queries. Tasks that previously required three months and six engineers can now be prototyped over a weekend. Consistent with the INSEAD/HBS mapping RCT finding: revenue gains are concentrated in firms that reorganized full production chains (1.9× revenue, +18% customer acquisition), not those that added AI to individual steps.

Sources: research/13-multimodal-sources/beyond-the-pilot/2026-04-14-most-enterprise-ai-agents-are-slop-heres-why-they-fail.md · research/13-multimodal-sources/beyond-the-pilot/2026-04-13-from-models-to-agentic-systems-the-next-enterprise-ai-s-curv.md


See also (wiki): ai-deskilling-risk — dedicated page on the skill atrophy, sycophancy, and practice-displacement patterns that run counter to short-run productivity gains


METR AI Usage Survey — Self-Reported Value Gains (n=349, May 2026)

Source: research/07-adoption-challenges/metr-ai-usage-survey-technical-workers-2026.md — METR, n=349 technical workers, convenience sample, Feb–Apr 2026 · HIGH with caveats / TIER 1

METR’s May 2026 survey provides the first large-sample self-report benchmark from a credible independent institution that explicitly accounts for the gap between self-report and controlled experiment:

  • Median 1.4–2x value gain from AI tools (early 2026), up from 1.3x in early 2025. Median speed gain is 3x — but METR explicitly flags that speed overstates value due to task substitution (doing more low-value fast tasks instead of fewer high-value ones).
  • Trajectory is the most durable finding: 1.3x value (early 2025) → ~2x (early 2026) → ~2.5x forecast (early 2027). Each data point is individually inflated; the direction and rate of change are more reliable.
  • METR’s own staff report the lowest gains of any subgroup. The authors’ preferred explanation: familiarity with prior RCT findings on overestimation bias causes them to anchor more conservatively. This is a calibration baseline: sophisticated, AI-literate users see ~1.4–1.7x; general employees report higher numbers.
  • Known overestimation bias confirmed: METR’s 2025 RCT found workers overestimated AI time savings by 40 percentage points vs. controlled experiment results. This discount should be applied to all self-report surveys in this space.
  • Sample caveat: Convenience sample, heavily skewed toward technical workers (engineers, researchers, academics), 50% Claude Code users. Not generalizable to enterprise knowledge workers in non-technical roles.

The METR survey and METR’s own RCTs tell a consistent story: self-report measures systematically exceed controlled experiment results, but the underlying gains are real. The correct enterprise conclusion is that AI productivity gains exist at 1.4–2x value range, not the 3x–10x commonly cited in vendor materials.

See also: [[ai-productivity-measurement-gap]] for the full framework on how to measure AI productivity rigorously inside an organization.

AI Daily Brief — METR Moore’s Law Update: Capability Frontier vs. Production Reliability (April 2026)

Source: research/13-multimodal-sources/ai-daily-brief/2026-04-xx-metr-moores-law-agents-capability-acceleration.md · AI Daily Brief / METR benchmark, April 2026 · MEDIUM / TIER 1

The METR benchmark is not a productivity RCT but provides the capability-ceiling context that makes RCT results interpretable:

  • The benchmark measures task complexity (human-hours equivalent) at 50% and 80% AI success rates. The 50% threshold is the capability frontier; the 80% threshold is where production deployments need to operate.
  • Opus 4.6 exceeded 14.5 hours at 50% success. At 80% success, results are substantially lower — the gap between these two numbers is where unrealistic enterprise deployment expectations originate.
  • The benchmark has now saturated: METR researcher: “If the task distribution was just a tiny bit different, we could have measured 8 hours or 20 hours.” Directional signal (continued acceleration) is credible; specific number is noisy.
  • Connects to the core RCT finding pattern: capability measures (benchmarks) consistently exceed deployment outcomes (RCTs) because real-world tasks require reliability, integration, and organizational context that benchmarks do not capture.

See Also

  • wiki/healthcare-ai-deployment.md — healthcare-specific RCT evidence including MASAI, Afshar, Lukac, and Holmgren studies with deployment context for health system buyers
  • wiki/ai-vendor-evaluation.md — procurement framework and head-to-head vendor comparison methodology; uses Lukac et al. as the primary model for product-level evaluation
  • wiki/ai-productivity-measurement-gap.md — how to build internal measurement systems that avoid the self-report inflation documented across this corpus
  • wiki/workflow-redesign.md — the primary mechanism by which individual task-level gains (documented in these RCTs) aggregate to firm-level returns