A cross-industry tracker of AI return-on-investment claims, organized by source credibility tier. Distinct from the roi-evidence.md page (which covers methodology and frameworks for evaluating claims) — this page tracks specific, named deployments with attributed metrics.
Credibility Tier Framework
TIER 1 — Board-audited regulatory disclosure Investor presentations filed with securities regulators, audited annual reports, earnings call disclosures. Figures subject to legal liability. Highest credibility for citing.
TIER 2 — Executive-attributed, company-published Press releases, company blog posts, executive quotes in named interviews. Self-reported but named and attributable. Use with appropriate caveats.
TIER 3 — Vendor-published case studies Vendor writes about their own customer’s results. No control group. Self-selected metrics. Directional signal only — not primary evidence.
TIER 4 — Commissioned research Research firm paid by vendor to study vendor’s customers. Structural bias. Note sponsor.
IBM IBV “From AI Projects to Profits” (n=2,500, Jun 2025) — The ROI Collapse Benchmark
- 31% → 7% ROI collapse as GenAI pilots scaled to enterprise deployment. Early returns were real but measured on easy, peripheral tasks. The 7% average now sits below the ~10% cost-of-capital hurdle most CFOs apply.
- Top decile: 18% ROI — above cost of capital — confirming the high-return path exists but requires structural differentiators, not just more investment.
- Only 25% of AI initiatives delivered expected ROI over the past three years (CEO self-report). Three-quarters of post-2022 AI spend has missed its business case.
- 64% of AI budgets now target core functions vs. 36% noncore — the difficult strategic pivot that explains why average ROI is declining while top-decile ROI is rising: core-function deployment is harder but higher-value.
- 6% pursue AI ad hoc (down from 19% a year earlier) — the unsupervised-pilot era is ending.
Source: research/12-agent-workers/ibm-ibv-ai-projects-to-profits-2025.md — MEDIUM-HIGH / TIER 2 (Oxford Economics fieldwork; IBM commercial interest in agentic positioning)
Tracked Deployments
Retail: Carrefour Group
- Source: FY2025 Investor Presentation (Feb 17, 2026) — TIER 1
- E-commerce GMV: €7.0bn FY2025 (+21% YoY) — HIGH credibility
- 2026 target (2022 plan): €9.5B GMV — MEDIUM credibility (stated target, not audited outcome)
- Digital investment: €3B by 2026 — MEDIUM credibility (stated commitment)
- CSR index: 113% achievement rate FY2025 — HIGH credibility
- Supply chain AI (SymphonyAI, 2019–2020): Order accuracy 20–40% to 70%+ — LOW-MEDIUM (TIER 3 vendor case study)
- Food waste: 50% reduction vs. 2016 baseline cited as “ahead of plan” — MEDIUM (strategic plan target, not audited as standalone in investor docs)
- Full research: research/02-corporate-tools/carrefour-ai-retail-supply-chain-2026.md
QSR: Papa Johns + Google Cloud
- Source: Dual vendor announcement (NRF 2026, January 2026) — TIER 3
- Deployment: First agentic omnichannel ordering (Gemini Food Ordering agent); nationwide rollout target end 2026
- Q3 2025 metrics: Directional CRM + conversion gains, no audited figures published
- Credibility: LOW-MEDIUM — no independent verification, no control group
- Full research: research/02-corporate-tools/papa-johns-google-cloud-agentic-ordering-2026.md
Marketplace/Deal Platform: SlickDeals
- Source: AWS re:Invent 2025 session — Mike Lively (SVP Engineering, Slickdeals) + Deepti Venuturumilli (AWS) — TIER 2/3 (named internal speaker at vendor-hosted conference)
- Deal scoring latency: 3 hours → 30 seconds (360x) — MEDIUM-HIGH credibility (operational infrastructure metric, precisely stated, named speaker)
- Important context: Operational latency, not user-facing experience. Relevant when inventory expires within hours (flash sales, time-sensitive deals). Less relevant for stable catalog e-commerce.
- Merchant outbound clicks + revenue: +7% — MEDIUM credibility (direction plausible, no methodology disclosure — A/B vs. pre/post unknown)
- Test execution time: 6 weeks → under 1 week — MEDIUM credibility (engineering velocity metric, not directly revenue-linked)
- Recombee homepage personalization: 70%+ higher product detail page views, 30%+ higher CTR — LOW-MEDIUM (vendor case study, no date, no control group; named PM quote from Daniel Uhm)
- Stack: Databricks + EKS + Amazon SageMaker + Kafka + Elasticsearch; Siamese retrieval + XGBoost ranking
- Key insight: Infrastructure modernization (SQL Server → Databricks, LAMP → EKS) preceded and enabled the AI results. Community signal (votes, comments) as real-time ranking input is the proprietary moat — not replicable with generic vendor personalization tools.
- Full research: research/02-corporate-tools/slickdeals-ai-deal-scoring-personalization-2025.md
Telecom: TELUS Digital / Fuel iX (2025–2026)
- Source: Company press release + MWC 2025/2026 presentations — TIER 2
- Tokens processed (2025): 2 trillion via Fuel iX platform — MEDIUM credibility (company-published, no audit)
- Custom employee assistants: 6,000+ created by TELUS employees (no-code builds) across 50,000+ global employees — MEDIUM credibility (MWC 2025 case study vintage, not independently verified)
- Hours saved: 500,000+ (~40 min/interaction) — LOW-MEDIUM (self-reported, no control group or methodology)
- $100M+ value claim: Session title framing at MWC 2026 — not broken down by use case or period. LOW credibility as a citable ROI figure.
- Platform: Fuel iX — multi-LLM (20+ models), no-code assistant builder, ISO Privacy by Design certified, per-token cost tracking by use case
- Applicability caveat: TELUS Digital is a telecom BPO/digital services arm serving 30+ global carriers. Scale is structurally atypical for mid-market enterprises. The platform governance model (central control plane before broad self-service) is the applicable lesson; volume figures are not benchmarks.
- Cross-reference: Anthropic Agentic Coding Trends 2026 report cites 13,000 custom AI solutions vs. 6,000 in Fuel iX framing — different time period or counting scope. Neither independently audited.
- Full research: research/02-corporate-tools/telus-digital-ai-transformation-mwc-2026.md
Telecom: Vodafone VOXI + SuperTOBi (2024–2025)
- Source: Accenture case study (vendor-published, Oct 2024) + Microsoft case study (vendor-published, 2025) — TIER 3
- VOXI GenAI chatbot (2024): First customer-facing GenAI chatbot in UK telecoms; Azure OpenAI via Azure AI Studio; built in ~3 months; Accenture as implementation partner. Qualitative outcomes only in public disclosures: higher containment rates, lower AHT, improved CX. No cost-per-chat figures published.
- SuperTOBi/TOBi (2025): 70% first-contact resolution rate (digital channels); 45 million monthly interactions in 13 countries, 15 languages; 50% improvement in complex customer journey resolution; 1-minute call time reduction. Portugal pilot: 15% → 60% first-time resolution; online NPS +14 (reaching 64).
- Credibility flag: The widely-cited “70% cost-per-chat reduction” traces to a December 2019 IBM THINK blog about Watson/TOBi — a different product, different technology stack, different era. Do not use as GenAI evidence.
- Credibility: MEDIUM (Microsoft case study, named executive Beverley Bartlett on record, deployment scale verifiable) / LOW-MEDIUM (Accenture VOXI case study, qualitative only, Accenture is both vendor and author)
- Full research: research/02-corporate-tools/vodafone-voxi-genai-customer-service-2025.md
Healthcare: AI Scribe Financial Impact at UCSF Health (Holmgren et al., JAMA Network Open, January 2026)
- Source: Peer-reviewed cohort study, JAMA Network Open (January 9, 2026) — TIER 1 / HIGH credibility. Difference-in-differences design using actual EHR billing records. n=1,202,734 encounters, n=1,565 physicians.
- +1.81 RVUs per week per adopting physician (95% CI 0.86–2.75; P<.001)
- +0.80 encounters per week (P=.04)
- ~$3,044 additional annual revenue per physician at 2025 Medicare rates. Private payer rates (120–150% of Medicare) imply $3,600–$4,600 in commercially-mixed practices.
- Zero increase in claim denial rates — AI-assisted documentation did not trigger payer pushback.
- Mechanism: May reflect more patients seen (freed documentation capacity) or improved coding of existing complexity, or both. Either is financially legitimate.
- ROI range: At $300–$600/physician/month subscription cost and ~$3,000–$4,600 annual gain, simple break-even requires commercially-mixed rates; non-billing benefits (burnout reduction, coder workload) are additive.
- Limitation: Single academic medical center (UCSF); early adopters may self-select; cannot fully isolate capacity vs. coding mechanism. Multi-site replication needed.
- Full research: research/06-industry-verticals/holmgren-jama-ambient-ai-scribe-physician-productivity-2026.md
Healthcare Revenue Cycle: Oliver Wyman AI-RCM Survey (May 2026)
- Source: Oliver Wyman Healthcare Perspectives, n=200+ decision-makers + 90 end users, US provider organizations — MEDIUM-HIGH credibility (independent consulting firm, disclosed scope, some statistics cross-sourced from HFMA/Frontiers in AI)
- Enterprise-wide deployment: 20–40% of organizations have moved beyond pilots to broad or enterprise-wide AI-RCM deployment
- No-regret consensus: 92% of respondents agree that “no-regret AI investments” exist in revenue cycle — the highest consensus on any AI investment category found in any 2026 survey
- Spend acceleration: 70–90% of decision-makers plan to increase AI-RCM spending over the next three years; 80% actively exploring/piloting/implementing GenAI for RCM (up 38 pp in under two years)
- Performance metrics: AI coding accuracy ≥90% in specific clinical domains; up to 46% reduction in coding time for complex cases; “millions of dollars annually” in recovered revenue per organization
- Four no-regret applications: Ambient documentation, CDI, coding automation, electronic prior authorization — all integrate into existing workflows without requiring redesign
- Competitive divergence: Smaller/community providers under-investing vs. large health systems; early adopters realizing compounding benefits
- Full research: research/06-industry-verticals/oliver-wyman-ai-revenue-cycle-healthcare-2026.md
Vendor-Commissioned TEI: Forrester / Microsoft Agentic AI Solutions (Jan 2026)
- Source: Forrester Total Economic Impact, commissioned by Microsoft, January 2026. 8 interviews across 6 organizations; 420 survey respondents; composite $2.5B-revenue, 10,000-employee enterprise. LOW-MEDIUM credibility for ROI headline; MEDIUM for cost structure and operational improvement categories.
- 3-year risk-adjusted ROI: 120%; NPV $24.2M; payback 15 months — upper-bound estimates from self-selected reference customers, no control group.
- Largest benefit category: External spend reduction ($16.2M) — marketing agency, outsourcing, external counsel, procurement. More credible than revenue-side projections because these costs are concrete; survey respondents report finance outsourcing down 11.9%, legal cost per review down 20.7%.
- Development cost dominance: Agent development ($15.1M) is 75% of total 3-year costs. Subscriptions are 12% ($2.5M). Organizations budgeting agentic AI at subscription cost are underestimating total investment by 7x.
- Operational improvements (anecdotal, self-selected): RFP response from weeks to 1 day; SQL queries from 3 days to minutes; new-hire onboarding from 60 to 30 days.
- Cross-reference vs. M365 Copilot TEI (Mar 2025): Smaller composite org ($2.5B vs. $6.25B) generates higher absolute NPV ($24.2M vs. $19.7M); longer payback (15 vs. 10 months) because development costs scale faster than license costs.
- Full research: research/05-analyst-firms/forrester-tei-microsoft-agentic-ai-2026.md
Enterprise Survey: Futurum Group 1H 2026 AI ROI Survey
- Source: Futurum Group press release (Feb 17, 2026) — survey of 830 global IT decision-makers — MEDIUM-HIGH credibility (independent analyst firm, survey self-report; methodology partially disclosed; full report behind subscription paywall)
- ROI metric shift: Productivity gains as primary metric fell from 23.8% to 18.0% (–5.8 pts); combined hard financial ROI (revenue growth 10.6% + profitability 11.1%) nearly doubled to 21.7%
- Agentic AI prioritization: 17.1% cite as #1 technology priority (up from 13.0% in 2H 2025) — 31.5% YoY increase; combined top-two: 39.3% (up from 32.0%)
- Platform consolidation: 65.9% on integrated platforms (up from 60.0%); 41.0% actively reducing application count; best-of-breed fell to 20.7%
- Build culture: 56.0% prefer in-house build (unchanged) — AI coding tools reinforcing, not eroding, build preference
- Analyst caveat: Futurum has commercial interests in AI vendor ecosystem; “nearly doubled” financial ROI metric reflects partly a survey redesign (splitting “overall financial performance” into revenue + profitability sub-measures); directional shift is credible, magnitude uncertain
- Full research: research/01-ai-native-landscape/futurum-enterprise-ai-roi-1h-2026.md
Enterprise Survey: Dun & Bradstreet Global AI Momentum Survey (n=10,000, May 2026)
- Source: D&B quarterly global panel — 10,000 businesses, 32 countries, Q1–Q2 2026. Vendor-commissioned. MEDIUM credibility — scale is the asset; D&B commercial interest in data-identity solutions is the liability.
- Adoption-to-ROI collapse: 97% active AI initiatives → 10% strong ROI. The gap is the finding.
- ROI distribution: 60% report some measurable ROI; 24% broad or strong returns; 10% strong ROI; 30% scaling to production
- Data readiness: Only 5% say data fully ready for AI — converges with Cloudera/HBR 7% from independent methodology
- Top obstacle: Limited data access (50%), privacy/compliance (44%), data quality (40%), system integration (38%)
- Risk management gap: Only 10% express high confidence in AI risk identification and mitigation
- Key implication for ROI framing: The 97→10% collapse is the most concise articulation of the deployment-to-value gap available in a single-survey data point; use alongside McKinsey 6% and BCG 5% to triangulate the high-performer share
- Full research: research/01-ai-native-landscape/dnb-global-ai-momentum-survey-2026.md
Enterprise Survey: Anthropic Economic Index — Learning Curves (March 2026)
- Source: Anthropic — Massenkoff, Lyubich, McCrory, Appel, Heller. n=1,000,000 conversations, February 5–12, 2026. Vendor-primary. MEDIUM credibility.
- Learning curve: Users with 6+ months Claude tenure show +3–4 percentage points higher task success rate after controlling for task type — the gain is genuine skill accumulation, not task selection.
- Value concentration: Computer and Mathematical occupations account for 35% of Claude usage. Opus (most capable tier) selected for 34% of software developer tasks vs. 12% of tutor tasks — model tier selection tracks task value.
- Capability scaling signal: For every $10 increase in task hourly wage, Opus selection rises 1.5pp on Claude.ai and 2.8pp on the API — API users are 2x more intentional about matching model to task value.
- Automation surface shift: Business sales outreach automation and automated trading both grew >2x in three months; coding has migrated from chat UI to API/agent layer. Seat counts on chat interfaces undercount actual economic activity.
- Vendor caveat: Anthropic-published, no independent replication. Direction consistent with independent evidence; magnitudes are upper bounds.
- Full research: research/01-ai-native-landscape/anthropic-economic-index-learning-curves-2026.md
NBER WP 34984 — The Cost-Reduction ROI Myth Disproven (March 2026)
748 CFOs (Duke/Fed Atlanta/Fed Richmond) provided both perceived and revenue-attributable AI productivity data. The mechanism analysis directly challenges the standard enterprise AI business case:
- Cost reduction motivations (reducing labor costs, reducing non-labor costs) show no significant correlation with actual measured revenue productivity gains — and frequently negative coefficients
- Innovation and demand motivations (developing/improving products, reaching customers more effectively) are the only factors statistically correlated with real gains in both 2025 and 2026 specifications
- The mean productivity gain CFOs perceive (1.8%) is approximately 3x the implied revenue-based gain (0.6%) — the “AI productivity paradox”
- Finance sector shows the largest measured gains (~0.8% implied LP in 2025, >2% projected 2026)
- Practical implication: If the ROI justification for an AI initiative is primarily cost reduction or headcount, this data suggests the returns are unlikely to show up in the income statement. The deals that do show up are aimed at revenue expansion and product improvement.
Source: research/01-ai-native-landscape/nber-w34984-ai-productivity-corporate-executives-2026.md
Human-AI Teams: Ju & Aral MIT/Johns Hopkins RCT (February 2026)
- Source: Randomized field experiment — MIT Sloan (Sinan Aral) + Johns Hopkins; n=2,234 workers; real advertising production task; February 2026 final version — TIER 1 / HIGH credibility
- Volume gain: +50% output per worker vs. human-human teams — the largest productivity gain documented in any human-AI collaboration RCT
- Real-world outcome paradox: Despite higher volume and text quality scores, H-AI teams showed no statistically significant performance advantage against 4.9 million live ad impressions. Volume gain did not convert to market outcome gain.
- Homogenization risk: AI-assisted teams produce output that converges toward a mean. Creative variance — the signal that distinguishes top-performing from average work — decreases.
- Substitution mechanism: H-AI teams made 62% fewer direct edits and delegated 17% more to AI — the result was substitution of human judgment, not augmentation.
- Implication for ROI measurement: Agentic deployments in judgment-intensive functions that measure only output volume will over-report ROI. The relevant metric is outcome quality, not output count.
- Full research: research/01-ai-native-landscape/ju-aral-collaborating-ai-agents-rct-2026.md
HBR Analytic Services — The Workflow Integration Gap (April 2026)
Two companion surveys (n=325 + n=385, December 2025 and March 2026 fieldwork; HBR Analytic Services, published April 29, 2026) sponsored by Hyland and Appian respectively. Apply MEDIUM-HIGH credibility — vendor-sponsored but directionally consistent with BCG, McKinsey, and Deloitte independent findings.
- Only 16% of AI-deploying organizations report high measurable value — 84% leave significant value unrealized despite active deployment
- The gap is not adoption but integration: 59% have AI in production; only 18% have it embedded in actual work processes
- Workflow embedding is the decisive variable: 71% of organizations that embedded AI in processes achieved substantial or moderate value, versus far lower rates for standalone deployments
- Data readiness is the prior constraint: 94% say connected data, processes, and applications are critical; only 27% have that connectivity in place
- Governance lags agentic deployment: 25% already use AI agents; only 48% have defined guardrails
Source: research/07-adoption-challenges/hbr-analytic-services-ai-readiness-workflow-2026.md · Dec 2025–Mar 2026 fieldwork, April 2026 published · MEDIUM-HIGH · TIER 1
KPMG Global Tech Report 2026 — The 4.5x High-Performer ROI Differential (n=2,500, Mar 2026)
KPMG’s largest tech-executive survey provides the most granular ROI-by-cohort breakdown in the 2026 corpus, and the finding is structural, not statistical noise.
- High performers generate 4.5x ROI on digital investments versus the industry average of 2x — a 2.25x performance gap explained by governance structure, not technology selection or investment level.
- The 74%/24% ambition-execution split: 74% of tech executives say AI delivers business value; only 24% achieve ROI across multiple use cases. This gap has worsened by 7 percentage points from the prior survey period — organizations are reporting more perceived value while achieving less measurable return.
- The governance fragmentation differential is the mechanism: only 2% of high performers report disconnected AI projects (vs. 34% of average performers); only 8% say tech debt prevents new investment (vs. 45% of average performers). These are not marginal differences — they are categorical separations between two operating models.
- ROI by cohort: high performers 4.5x; smaller organizations 3.6x; transformation-focused organizations 3.2x; industry average 2x; organizations with fewer cost pressures 2.6x. Smaller firms outperform average despite lower absolute investment — suggesting governance and focus matter more than scale.
- Corroborated by Roland Berger 2026 (~90% returns lag spend; ~10% capture consistent value), Oliver Wyman 2026 (27% of all CEOs say AI ROI met expectations), and Deloitte 2026 (34% deep transformation vs. 66% efficiency-only).
Source: research/04-consulting-firms/kpmg-global-tech-report-2026.md · n=2,500, 27 countries, Mar 2026 · MEDIUM-HIGH · TIER 1
Stanford HAI AI Index 2026 — Economy Chapter (April 2026, HIGH credibility)
The most comprehensively sourced annual AI economy report, drawing on Quid investment data, McKinsey surveys, Lightcast job postings, and Brynjolfsson payroll studies. Key ROI-relevant findings:
- 88% organizational adoption, single-digit EBIT impact share. McKinsey’s 2025 survey (cited in the Stanford synthesis) finds 88% of surveyed organizations use AI in at least one function (up from 78% in 2024). Fewer than 6% of organizations capture significant financial value. This convergence — 88% adoption, ~6% high performer share — is the clearest single-number statement of the deployment-to-value gap.
- 14%–50% task-level productivity gains in structured work (customer support, software development, marketing tasks). Zero or negative in judgment-intensive work. The gains are real but bounded by task type — aggregate enterprise ROI depends on whether the organization has redesigned enough workflows to capture the gains at scale.
- US macro labor productivity: +2.7% in 2025, nearly double the prior-decade average (BLS). This is the first macro signal that AI task-level gains are beginning to aggregate, though the distribution across firms is highly uneven.
- Consumer surplus $172B in early 2026 (Brynjolfsson et al., choice experiment, N=2,000), up from $112B in 2025. Users are capturing more value than AI companies are monetizing — consistent with Nordhaus (2004): innovators historically capture only ~3% of total social returns.
Source: research/01-ai-native-landscape/stanford-hai-economy-2026.md — HIGH credibility, April 2026
BCG AI Radar 2025–2026 — The ROI Commitment/Return Gap (n=1,803 / n=2,360)
- 94% of CEOs will continue investing regardless of near-term payoff (BCG AI Radar 2026, n=2,360). This is the definitive data point on the end of conditional AI budgeting — and the reason the deployment-to-value gap is not self-correcting via normal capital allocation discipline.
- 75% rank AI as top-3 strategic priority; 25% report significant value (BCG AI Radar 2025, n=1,803). The 3:1 ambition-to-return ratio has held through two consecutive annual surveys.
- 60% track no financial AI KPIs. Without financial accountability, there is no mechanism to distinguish working investments from those that are not.
- 2.1x ROI from focused (3.5 use cases) vs. broad (6.1 use cases) deployment. Concentration of effort into end-to-end redesign, not breadth, is the differentiator.
- BCG’s companion Build for the Future 2025 (n=1,250): 5% of companies generate substantial, scalable AI value; 60% generate minimal returns despite active use.
Source: research/04-consulting-firms/bcg-ai-radar-2025-2026.md — TIER 1, MEDIUM-HIGH
McKinsey “Rewiring” (March 2025) — The EBIT Impact Ceiling
- >80% of respondents report no tangible enterprise-wide EBIT impact from gen AI use. 17% attribute ≥5% of EBIT to gen AI — a significant but still small minority (n=1,491, July 2024).
- Workforce reductions are among the attributes most correlated with bottom-line value — and larger organizations act on this more than smaller ones. Cost reduction is a more reliable short-term ROI mechanism than revenue expansion, but the NBER CFO study (WP 34984) complicates this: cost-reduction motivations show no significant correlation with measured revenue-based productivity gains.
Source: research/04-consulting-firms/mckinsey-rewiring-2026.md — TIER 1–2
Patterns Across the Evidence Base
What holds up under scrutiny:
- Scale metrics (GMV, store counts, net sales) from investor disclosures — these are audited
- Deployment scope (warehouse count, supplier count) from case studies — structurally verifiable even if results are not
- ESG/sustainability targets in investor documents — increasingly subject to third-party validation (SBTi, CDP)
What requires caveats:
- Productivity percentage claims from vendor case studies (always directional, never primary evidence)
- “X times faster” or “Y% reduction” from press releases without methodology disclosure
- Before/after accuracy claims without naming the measurement methodology
What to avoid citing:
- Any unattributed “studies” or unnamed research firms
- Percentage claims with no named organization, year, and measurement definition
- Figures that cannot be traced to a named document with a date
Common Gaps in Retail AI ROI Claims
- No baseline defined. “Improved accuracy” without a starting point is not a metric.
- No time horizon. Results claimed without specifying when measurement was taken post-deployment.
- No scope definition. “Pilot results” ≠ full deployment results.
- Conflated causation. E-commerce GMV growth attributed to AI when infrastructure, pricing, and market conditions also changed.
Carrefour’s investor disclosure avoids most of these gaps for the financial metrics. The SymphonyAI case study has gaps 1 and 2 addressed (baseline stated, deployment timeline documented) but not 3 (pilot vs. full deployment performance not distinguished).
Economist Enterprise AI Survey 2026 — The Measurement-Optimism Gap (TIER 1 / MEDIUM-HIGH)
n=1,221 senior technology executives, 296 CIOs, 18 countries; fieldwork November 2025–January 2026; commissioned by Databricks; published May 2026.
- 80% say AI programs exceed expectations — fewer than half have any formal mechanism to verify whether that’s true. The confidence-without-measurement gap is the defining accountability failure of enterprise AI in 2026.
- Only 40% of organizations require teams to track AI business impact. Among self-identified “AI leaders,” 84% claim returns surpass expectations but only 43% mandate measurement. Optimism and accountability infrastructure are running on separate tracks.
- 97% of organizations with unified data architecture report ahead-of-schedule ROI vs. 77% without — a 20-point structural predictor. Treat with caution: Databricks (the commissioning vendor) sells data infrastructure. Finding is directionally consistent with HBR Analytic Services 2026 and BCG 10/20/70 data, but vendor-favorable.
- This study is the most direct measurement of the expectation-vs.-evidence gap in the 2026 corpus. Pair it with NBER WP 34984 (CFOs report AI not showing up in income statement) and McKinsey State of Organizations 2026 (81% report no bottom-line impact) for the full picture.
Source: research/01-ai-native-landscape/economist-enterprise-ai-survey-2026.md · Nov 2025–Jan 2026 fieldwork, May 2026 published · MEDIUM-HIGH · TIER 1
NVIDIA State of AI 2026 — Vendor Survey (TIER 3 / MEDIUM-LOW)
n=3,200+ respondents, Aug–Dec 2025, published March 9, 2026. Five industries: financial services, retail/CPG, healthcare, telecom, manufacturing.
- 88% report AI increased annual revenue (30% by >10%) — self-reported, not independently audited. Practitioner-heavy sample (40% AI practitioners) and opt-in from NVIDIA ecosystem inflate positive responses.
- 87% report reduced annual costs (25% by >10%; retail/CPG: 37% by >10%) — same caveat.
- 64% actively using AI — deployment signal credible; consistent with Stanford HAI (78% org) and McKinsey (88% any function). NVIDIA’s lower figure reflects its “active use” definition.
- 44% deploying or assessing AI agents — highest agentic figure in the 2026 multi-industry corpus. Telecom leads at 48%.
- 86% increasing AI budgets — directional signal, consistent with Gartner, KPMG, Deloitte.
Why the 88% and McKinsey’s 6% can both be true: Self-selection, practitioner overrepresentation, and a low bar for “contributed to revenue” explain most of the gap with McKinsey’s >5% EBIT threshold. Not contradictory; measuring different things.
Source: research/01-ai-native-landscape/nvidia-state-of-ai-2026.md
Menlo Ventures State of GenAI in the Enterprise 2025 (TIER 1 / MEDIUM-HIGH)
n=495 U.S. enterprise decision-makers, November 7–25, 2025, published December 9, 2025. Sponsored by Menlo Ventures (investor in Anthropic — flag for LLM share figures).
- Enterprise AI deals convert to production at 47% — nearly double traditional SaaS’s 25%. Production conversion is not the rate-limiting step; value capture post-production is where most deployments stall.
- 76% of AI use cases in 2025 were purchased rather than built (up from 53% in 2024). The economics of building custom systems on foundation model APIs have shifted against internal builds at all but the largest engineering organizations.
- At least 10 AI products now generate over $1B in ARR; roughly 50 have crossed $100M. Enterprise AI has moved from speculative pilots to durable product revenue at scale.
- Developer teams self-report 15%+ velocity gains from AI coding tools — directional signal, not RCT-measured. Consistent with academic literature (Peng et al., 2022: 26% faster; METR 2025 RCT: experienced developers 19% slower on complex tasks). The gap between studies likely reflects task type and developer experience level.
- Only 16% of enterprise AI deployments qualify as true agents. The ROI narrative around agentic AI applies to a small fraction of what is currently in production.
Source: research/01-ai-native-landscape/menlo-ventures-state-genai-enterprise-2025.md
Gartner CSO Survey 2026 — The Sales AI Reinvestment Gap (TIER 1 / MEDIUM-HIGH)
n=210 CSOs and senior sales leaders; fieldwork Jan–Feb 2026; published May 19, 2026.
- AI saves sellers 4.8 hours/week — measurable, real, and largely wasted. 72% of sales organizations fail to reinvest those hours into high-value selling activities.
- Bimodal ROI distribution: 25% of orgs report 50%+ positive return; 20% report 50%+ negative return. The average masks a performance split driven entirely by whether management actively redirects time savings.
- Reinvestment multiplier: orgs that actively reinvest AI time savings are 2.2x more likely to exceed customer growth goals and 3.1x more likely to exceed lead-to-opportunity conversion targets.
- 31% of CSOs cite “difficulty proving ROI of AI-driven tools” as a top 2026 challenge — consistent with the NBER WP 34984 finding that AI is not showing up in income statements at most firms.
- Companion buyer survey (n=645, Aug–Sep 2025): 69% of B2B buyers turn to sales reps to validate AI-generated insights — confirming that buyer-side AI adoption strengthens, not replaces, the human sales role.
Source: research/05-analyst-firms/gartner-cso-sales-ai-reinvestment-gap-2026.md · Jan–Feb 2026 fieldwork, May 2026 published · MEDIUM-HIGH · TIER 1
Futurum Group — Enterprise AI ROI Metric Shift (TIER 1 / MEDIUM-HIGH)
n=830 global IT decision-makers; published February 17, 2026. Futurum Group independent research; no disclosed vendor sponsor.
- Direct financial impact nearly doubled as the primary ROI metric — from ~11% to 21.7% in a single survey cycle. Productivity gains fell from 23.8% to 18.0%. Enterprise AI buyers are now accountable to CFOs and boards, not department heads.
- Agentic AI surged 31.5% YoY as a technology priority; 17.1% rank it their #1 investment priority. The assistive-to-agentic transition has moved from prediction to procurement reality.
- 65.9% prefer integrated platforms over best-of-breed stacks — 5.9-point increase. Governance complexity across fragmented AI portfolios is driving vendor consolidation.
- Consumption-based pricing is now the dominant GenAI commercial model (42.9%). AI cost is variable and usage-linked, not seat-licensed — a direct CFO planning implication.
Source: research/09-ai-adoption-cycle/futurum-enterprise-ai-roi-shift-2026.md · Feb 2026 · MEDIUM-HIGH · TIER 1
Practitioner voices (pillar 13)
Named F500 practitioners on-record with production deployment metrics. Each quote traces to a specific ingested source file.
“We process on our network 160 billion transactions. If you think about a one-in-a-million occurrence, that happens to us 160,000 times a year.”
Johan Gerber, EVP, Head of Security Solutions, Mastercard — Beyond the Pilot / VentureBeat, April 2026 · Source: research/13-multimodal-sources/beyond-the-pilot/2026-04-13-mastercard-vs-fraud-ai-tech-unpacked.md
“Right now, we are handling about 95% of the orders that come through. It’s about 150,000 orders every day, and we’re continuing to add sites.”
Will Crouchhorn, Product Manager, Wendy’s — Me, Myself, and AI / MIT SMR + BCG, December 2025 · Source: research/13-multimodal-sources/me-myself-and-ai/2025-12-16-hungry-for-learning-wendys-will-croushorn.md
Wendy’s Fresh AI Initiative handles 95% of drive-through orders autonomously at 150,000 orders/day. This is production-scale named-company evidence from a Fortune 500 quick-service restaurant — not a vendor case study, not a pilot. The “60% reduction in sorry utterances” metric from the same episode is a production quality proxy: reduction in agent failure acknowledgments is a direct behavioral signal of model improvement, not a lagging CSAT score.
The 160 billion figure is the production denominator behind every Mastercard AI fraud-model ROI claim. Any percentage improvement on that base is a TIER 1 metric — the underlying transaction volume is independently verifiable from Mastercard’s investor filings.
“Over the course of those first few months there, we went from zero users to 250,000 users. Today, one in two JPMorgan employees around the world use it nearly every day.”
Derek Waldron, Chief Analytics Officer, JPMorgan Chase — Beyond the Pilot / VentureBeat, December 2025 · Source: research/13-multimodal-sources/beyond-the-pilot/2025-12-17-how-jpmorgan-engineered-a-30k-ai-agent-economy.md
Adoption scale (not just deployment announcement) from a named CAO. 50% daily active use among 250,000 users at a regulated financial institution is a production evidence benchmark — not a pilot result and not vendor-published.
“We’re seeing after LLMs, a 2x increase in topic detection there. And if we can get the topic right, we can get the tooling right. We can increase the number of self-service that happens on the platform, which means we can actually make more of our CS agents available to have actual conversations with customers.”
Pranav Pathak, Director of Product Machine Learning, Booking.com — Beyond the Pilot / VentureBeat, April 2026 · Source: research/13-multimodal-sources/beyond-the-pilot/2026-04-13-how-bookingcom-built-ai-that-converts-millions.md
Illustrates the ROI chain that survey data obscures — accuracy improvement (2x topic detection) flows into tooling improvement, then into self-service rate, then into human agent capacity reallocation. This is the causal structure enterprise AI ROI claims should be required to show.
Mid-market practitioner evidence
- research/07-adoption-challenges/ai-success-at-12-months-realistic-outcomes.md — what realistic year-one AI outcomes look like for a 200–2,000 person company; calibrates board expectations against vendor-ROI narratives
- research/07-adoption-challenges/mid-market-ai-case-studies-measured-value.md — P&L-linked case studies from mid-market AI deployments; where the quantified evidence actually exists outside the Fortune 500
Snowflake “ROI of Gen AI and Agents 2026” — Industry-Segmented ROI (n=2,050, MEDIUM)
Snowflake commissioned Omdia/Informa TechTarget to survey 2,050 enterprise leaders across 9 countries and 6 industries (Aug–Sep 2025, published March 2026). All respondents are active AI users at organizations with 500+ employees — selection bias toward positive outcomes.
- $1.49 per $1 invested (49% average ROI) — up from 41% the prior year
- Industry spread: 38%–69% — manufacturing trails advertising & media by 31 points; the gap is data readiness, not model quality
- 92% of early adopters report positive returns — but only 49% formally measured ROI before reporting it
- 20% of unstructured data is AI-ready; 32% of structured data meets AI-readiness standards — independent corroboration of Cloudera/HBR’s 7% “completely ready” finding
- 57% use unauthorized AI tools; 66% of C-level leaders do — governance gap runs top-down, not bottom-up
- Agentic AI: 31% in production; expected 47% ROI in next 12 months
Reconciliation with McKinsey 6% EBIT: Not contradictory. Snowflake measures self-reported returns among active users who chose to answer a productivity survey; McKinsey measures bottom-line EBIT impact across all companies. Both are accurate. They describe different populations.
Source: research/05-analyst-firms/snowflake-roi-gen-ai-agents-2026.md
Return on AI Institute: What 1,006 Executives Got Wrong About Their Programs (March 2026)
Davenport and Srinivasan (HBR, March 17, 2026, n=1,006, 11 countries, 32 industries) provide the primary-survey answer to “what specifically do high-value AI programs do differently”:
- The differentiator is measurement discipline, not tool selection. The seven factors: data quality, AI talent and expertise, clear business objectives, effective change management, scalable infrastructure, continuous monitoring, and ERP/CRM integration. All seven are operational disciplines, not technology choices.
- Data quality multiplier: 2x. Organizations with clean/complete/current data are 2x as likely to achieve strong AI value. This is the upstream investment most programs skip.
- Change management gap: 30+ points. Structured change management (training + stakeholder engagement + communication) produces >80% adoption. “Set-and-forget” deployments produce <50%. Same technology, 30-point outcome gap.
- Static deployment decay: −15% in 6 months. Deployments without real-time monitoring and automated retraining lose effectiveness as data patterns shift. The programs that sustain ROI budget for operational maintenance, not just deployment.
- Integration beats point solutions on adoption. AI that surfaces insights inside existing ERP/CRM platforms users already open daily outperforms standalone AI tools on utilization — even when the standalone tool has superior model performance.
Note: 90% reporting value vs. McKinsey’s 6% EBIT-impact is not a contradiction. Davenport/Srinivasan include “small” benefit (9%); McKinsey requires >5% EBIT impact. The studies measure different outcome levels. Both are credible.
Source: research/04-consulting-firms/davenport-return-on-ai-institute-hbr-2026.md
Deloitte 2026: The Transformation-Efficiency Divide (n=3,235)
Deloitte surveyed 3,235 senior leaders across 24 countries (Aug–Sep 2025) on AI deployment status, ROI, and governance. The central finding: access has outpaced value creation.
- Worker access to enterprise AI rose 50% in 2025 (under 40% → ~60%), yet fewer than 60% of workers with approved tools use them regularly.
- Only 34% of organizations are using AI to genuinely reimagine their business. The remaining 66% capture efficiency gains but not structural competitive advantage.
- Only 25% of organizations report that 40%+ of AI pilots have reached production. Fifty-four percent expect to cross that threshold within 3–6 months.
- Automation timeline: 36% expect 10%+ of jobs fully automated within one year; 82% within three years.
- Governance lag: only 21% have a mature governance model for autonomous AI agents, yet 74% plan significant agentic deployments within two years.
Credibility note: Deloitte is a consulting firm with commercial interest in AI transformation services. The deployment and ROI figures are self-reported by leaders. The survey is TIER 2 (Aug–Sep 2025). Use as benchmarking data for board conversations, not as causal evidence of AI value.
Source: research/04-consulting-firms/deloitte-state-of-ai-enterprise-2026.md
Consulting Firm ROI Studies — Pillar 04 Index
Key Pillar 04 files with quantified ROI claims. All carry vendor/consulting commercial interest caveats — treat as directional benchmarks, not primary causal evidence.
| File | Key Finding | Credibility |
|---|---|---|
| research/04-consulting-firms/deloitte-ai-roi-paradox-2025.md | n=1,854, ROI horizon 2–4 years (not 7–12 months); 2–4x productivity for task-level work vs. 0–1x at program level | TIER 2 |
| research/04-consulting-firms/bcg-ai-first-cost-advantage-2026.md | 10/20/70 rule: 70% of value in workflow redesign; AI leaders 3x cost reduction, 2.7x ROIC vs. peers | MEDIUM-HIGH |
| research/04-consulting-firms/mckinsey-ai-transformation-manifesto-2026.md | 20% EBITDA uplift for leaders; workflow redesign is #1 EBIT predictor; n=20 leading companies | MEDIUM |
| research/04-consulting-firms/pwc-ai-performance-study-2026.md | 8x revenue growth and 5x profitability for “AI front-runners” vs. peers | MEDIUM (PwC advisory interest) |
| research/04-consulting-firms/randy-bean-ai-data-leadership-benchmark-2026.md | n=~110 Fortune 1000 CDOs; only 21.3% generating measurable revenue from AI | MEDIUM-HIGH (invitation-only F1000 panel) |
| research/07-adoption-challenges/gartner-io-ai-roi-stall-2026.md | n=782 I&O leaders (Q4 2025): only 28% of AI use cases fully succeed; 20% fail; 52% stall; IT service management is the one category with consistent success | MEDIUM-HIGH |
| research/07-adoption-challenges/coastal-oxford-economics-ai-operations-report-2026.md | n=800 production deployments (May 2026): 46% not meeting expectations; 73% have persistent data problems; only 1 in 6 has a dedicated AI team | MEDIUM (vendor-commissioned, Oxford Economics independent) |
Gallagher 2026: The 28-Month ROI Horizon and the Confidence-Governance Paradox (n=1,200+)
Gallagher’s third annual AI Adoption and Risk Benchmarking Survey (n=1,200+ global businesses, February 2026) is the only corpus source that quantifies the expected ROI payback period directly.
- 28 months average estimated timeline to realize ROI on AI deployment — nearly 2.5 years from investment to value realization.
- 63% of organizations actively measure ROI; technology and financial services sectors lead on measurement frameworks.
- 63% have fully operationalized or implemented AI within at least part of operations (up from 45% in 2025).
- 82% report positive business revenue impact — but this is the self-reported perception figure. McKinsey’s 6% EBIT-impact data and the Workday 14% net-positive-outcome rate provide the objective counterweight.
- Governance paradox: 93% say they understand AI risks “quite well” or “very well” — up from 77% in 2024. Yet fewer than 50% have formal AI risk management frameworks, ethical impact assessments, or AI-specific incident response plans. Confidence has increased faster than infrastructure.
Credibility note: MEDIUM-HIGH. Gallagher is a Fortune 500 insurance broker with commercial interest in highlighting AI risk exposure. Third annual tracking survey enables YoY comparison (45%→63% adoption). The confidence-governance gap finding is self-damning relative to Gallagher’s commercial interest, which adds credibility. Fieldwork dates not disclosed in public materials; treat percentages as directional.
Source: research/04-consulting-firms/gallagher-ai-adoption-risk-benchmarking-2026.md
Thomson Reuters Institute 2026: The Professional Services Measurement Gap (n=1,514)
TRI’s fourth annual survey (n=1,514, fieldwork Oct–Nov 2025, 27 countries, legal/tax/accounting/risk/government professionals) finds the most severe ROI measurement failure of any corpus survey:
- Only 18% of professional services organizations collect any ROI metrics from AI — roughly unchanged from 2025 despite nearly doubling adoption rates.
- 40% of professionals don’t know whether their organization measures AI ROI at all.
- 82% total are either not measuring or unaware whether measurement exists.
- Among those measuring (the 18%), metrics are internally focused: employee usage (77%), internal cost savings (73%), employee satisfaction (64%). Client satisfaction (54%) and revenue impact (42%) lag far behind.
- The formal strategy link: TRI’s 2025 Future of Professionals Report finds organizations with a formal AI strategy are 3x more likely to achieve positive ROI than those without one. The 82% measurement gap is the mechanism by which organizations remain outside the 3x cohort.
Why this matters for the measurement debate: The Gallagher finding (63% measure ROI, 28-month payback) appears inconsistent with the TRI finding (18% measure ROI). Resolution: Gallagher surveyed cross-industry with financial services overrepresented; TRI surveyed professional services (legal, tax, accounting) specifically. Professional services firms — with hourly billing models and partner-level resistance to AI disclosure — are structurally less likely to have ROI measurement infrastructure than financial services.
Credibility note: MEDIUM-HIGH. Thomson Reuters is a professional services data vendor with commercial interest in AI adoption. Sample drawn from TR lists — may oversample TR platform users. But the self-critical findings (50% threat perception, 82% no measurement) do not serve a vendor narrative, which adds credibility.
Source: research/10-client-analysis/thomson-reuters-ai-professional-services-2026.md
Gartner: Foundations Beat Models — The 4x Investment Differential (n=353, April 2026)
Gartner’s survey of 353 D&A and AI leaders (Nov–Dec 2025) provides the investment-ratio evidence that links foundation spending to AI outcomes:
- Successful AI organizations invest 4x more as a percentage of revenue in data quality, governance, AI-ready talent, and change management vs. organizations reporting poor outcomes.
- Only 39% of technology leaders are confident their AI investments will produce positive financial impact — meaning the 4x investment gap is measurable against the confidence gap.
- Organizations at highest AI-ready D&A maturity achieve up to 65% better business outcomes (revenue growth + cost optimization) vs. lowest-maturity peers. Self-reported; directionally consistent with MIT CISR (17pp profit performance gap Stage 4 vs. Stage 1).
- Corroboration from independent sources: Davenport/Return on AI Institute (n=1,006, March 2026) finds data quality is a 2x ROI multiplier; Cloudera/HBR (n=230, March 2026) finds only 7% of enterprises are completely data-ready.
Implication for ROI evidence calibration: The 39% confidence figure is the financial-outcome side of the confidence gap that also appears in the enablement illusion (Gartner GLMS, n=12,004: 27% have comprehensive AI strategy) and the production-rate gap (McKinsey: 6% EBIT-impact cohort). These are all measuring the same bifurcation from different angles.
Source: research/05-analyst-firms/gartner-data-foundations-ai-success-2026.md
Cross-reference — Gartner AI Maturity Longevity Study (Jun 2025, n=432): High-maturity organizations keep AI in production for 3+ years at 2.25x the rate of low-maturity peers (45% vs. 20%). 91% of high-maturity orgs have a dedicated AI leader; only 14% of low-maturity business units report being ready to use AI. The longevity finding reframes ROI: sustainable value capture requires operational infrastructure, not just launch capability.
Source: research/05-analyst-firms/gartner-ai-maturity-enterprise-2025.md — MEDIUM-HIGH / TIER 1 (n=432, Q4 2024 fieldwork, Jun 2025)
G-P AI at Work 2026: The ROI Reckoning — 73% Underwhelmed, 88% Concerned About Performative Use (n=2,850 executives, May 2026)
G-P’s third annual AI at Work survey (n=2,850 VP-level-and-above executives, 6 countries, May 2026) adds three findings not yet quantified in the independent corpus:
ROI disappointment rate:
- 73% of executives report AI investments fell short of ROI expectations in the past 12 months
- 16% experienced outright negative ROI
- ~70% prepared to cut AI budgets if 2026 goals unmet
- Share of orgs “aggressively innovating with AI” fell from 60% → 42% year-over-year
Performative productivity — first large-sample quantification:
- 88% of executives are concerned employees are using AI performatively: generating outputs, hitting usage metrics, and meeting mandates without producing real business value
- 47% say they are “very or extremely concerned” that this is already occurring
- This is the largest-sample confirmation yet of the mechanism the corpus elsewhere calls “performative compliance” — explaining why adoption curves rise while productivity curves stay flat
Review overhead — unmeasured drag on ROI:
- 69% report teams spending extra time monitoring and reviewing AI outputs before they reach decision-makers
- Only 23% have total confidence in AI accuracy
- This review time does not appear in standard productivity metrics — systematically overstating net productivity gains
Cross-validation against independent sources:
- BCG (n=10,635): 5% substantial gains → same gap
- McKinsey (n=1,993): 6% high performers → same gap
- NBER W34836 (n=5,867): 9-in-10 executives report no productivity/employment impact → same gap
- G-P gives a mechanism (performative use) where the others give only the outcome
Credibility note: MEDIUM. G-P is an employer-of-record platform with commercial interest in findings about workforce value and global hiring. Workforce-planning sections carry that bias. The performative-productivity and ROI-disappointment findings are directionally consistent with independent sources and do not serve an obvious vendor narrative.
Source: research/07-adoption-challenges/gp-ai-at-work-2026-reckoning.md
Gallagher AI Adoption and Risk Benchmarking 2026 (n=1,250, US/UK/CA/AU, Feb 2026)
ROI timeline — 28-month payback horizon:
- Organizations across all size bands report an average of 28 months expected before AI transformation value outweighs upfront cost
- Nearly two-thirds are actively measuring ROI — one-third are deploying without tracking returns
- Contextualizes the widespread ROI disappointment in G-P (73%), Deloitte (47% cost savings satisfied), and Gartner GLMS data: expectations may have been miscalibrated to 12 months, not 28
Governance gap — confidence without formalization:
- 93% say they understand AI risks “quite/very well” — up from 77% in 2024
- Fewer than 50% have adopted formal AI risk management frameworks
- Fewer than 50% have conducted ethical impact assessments
- Fewer than 50% have developed AI-specific incident response plans
- Only 56% have communicated their AI adoption strategy to their own workforce
Insurance coverage gap — active losses, partial coverage:
- 1 in 5 insurance professionals report clients experienced AI-related losses in the past year
- Just over half of those losses were fully covered by existing policies
- Most impacted classes: cyber liability, product liability, employment practices liability
- 200+ active AI/ML legal cases identified in insurance professionals’ books
Headcount cross-validation:
- 59% report headcount reduction done or planned — fourth independent data point alongside Morgan Stanley (4% net), Gartner (80%), Gallup (27% disruption report)
Credibility note: MEDIUM-HIGH. Gallagher is an insurance broker with commercial interest in AI risk advisory services. Insurance professional data on client losses is primary observation (not survey self-report) — highest-credibility finding in the dataset. ROI and governance figures are directional; specific percentages carry vendor-adjacent framing risk.
Source: research/07-adoption-challenges/gallagher-ai-adoption-risk-benchmarking-2026.md
WEF/Accenture MINDS “Proof over Promise” (Jan 2026)
Named case study evidence from 32 organizations, 30+ countries, curated high-performer cohort (selection bias toward success). Released at Davos.
- Foxconn & BCG: 80% of decision-making automated, ~$800M in value unlocked
- ICBC (financial services): €61M (¥500M) profit increase from AI-powered decision support
- Fujitsu (supply chain): $15M warehousing cost reduction, 50% staffing reduction
- Ant Group (healthcare): 90%+ diagnostic accuracy across 5,000 facilities
- UCSF & SandboxAQ: 36x acceleration in Parkinson’s drug discovery
- Horizon Power: 50,000-fold efficiency gain in energy market forecasting
- Pattern across all: double-digit productivity/revenue gains are real — but require operational redesign, not technology installation
- ~75% of leading organizations reinvest AI returns into new domains (compounding loop mechanism)
Credibility note: MEDIUM-HIGH. WEF + Accenture joint; independent assessment council. MINDS is a curated cohort — not representative of average enterprise. Named case studies are most credible elements. Accenture has commercial interest in AI services. Consistent with BCG/McKinsey high-performer cluster pattern.
Source: research/07-adoption-challenges/wef-accenture-proof-over-promise-2026.md
Oliver Wyman Forum CEO Agenda 2026 (April 2026)
n=415 CEOs (266 public, 149 private), co-published with NYSE, representing ~10% of global market cap. Third annual survey. CEO-specific sample — distinct from broader senior-executive surveys in the corpus.
- 27% of all CEOs say AI ROI met or exceeded expectations — down from 38% a year prior. Sharpest single-year drop in CEO AI satisfaction in this survey series.
- 49% of deployment leaders (scaling AI in 2+ business categories) report ROI meeting/exceeding expectations, vs. 17% of laggards — a 32-point gap driven entirely by workflow redesign and organizational design, not investment level.
- 53% say it’s too early to assess (up from 41%) — capital is being deployed faster than accounting systems can measure outcomes; this is not a negative signal for AI, but it signals CFO scrutiny pressure in 2026 planning cycles.
- ~25% report zero revenue impact — the most concerning cohort; concentrated among organizations that deployed AI onto existing workflows rather than redesigning first.
- Deployment leaders redesign workflows at 49% vs. 38% average — the same leader/laggard pattern found in PwC (2x), Deloitte (34% transformation vs. 66% efficiency-only), and BCG AI Radar 2026 (5% value-generating).
Credibility note: MEDIUM-HIGH. CEO-only panel at this scale is rare; NYSE co-authorship adds independence from Oliver Wyman’s commercial interest. Self-reported expectations, no independent financial verification.
Source: research/04-consulting-firms/oliver-wyman-ceo-agenda-2026.md
Capgemini Research Institute “AI Perspectives 2026” (Jan 2026)
Two companion studies: n=1,505 executives at >$1B companies (15 countries, Nov 2025) and n=500 CXOs including 100 CEOs (>$10B companies, Aug–Sep 2025).
- 38% of organizations have operationalized GenAI beyond pilots — the threshold Capgemini uses is any use case in production, not financial impact. Brackets the BCG/McKinsey 5–6% (financial-impact threshold): 38% have crossed from pilot to production; 5–8% have crossed production to measurable ROI.
- AI budgets rising to 5% of business spend (from 3%) — a 67% YoY increase; organizations shifting to 5-year investment horizons rather than 12-month payback expectations
- Budget composition: infrastructure + data + governance + workforce upskilling — not primarily model licensing; consistent with Forrester TCO finding that total cost is ~3x license cost
- 63% pruning low-value AI projects — the pilot proliferation era is ending; CFO scrutiny is reshaping AI portfolio decisions
- Top ROI enablers named by organizations: executive sponsorship (67%), workforce upskilling (60%), governance frameworks (53%) — not model selection or compute
Credibility note: MEDIUM-HIGH. Capgemini Research Institute — commercial consulting interest in enterprise AI. Large disclosed n=, transparent methodology. 38% operationalization figure directionally consistent with BCG (5% substantial), McKinsey (6% EBIT impact), Deloitte (34% transforming core processes) when threshold differences are accounted for.
Source: research/04-consulting-firms/capgemini-ai-perspectives-2026.md
KPMG Global AI Pulse 2026 (n=2,110, 20 markets, March 2026)
- 64% of global leaders report meaningful AI business value — AI leaders: 82%; average: 64%; laggards: 20%. The 18-point gap between leaders and average companies is the maturity premium (KPMG, n=2,110, March 2026).
- Organizations investing in talent alongside AI are 4x more likely to see meaningful value (77% vs. 20%) — the single most important sequencing finding from this survey. Talent investment is the multiplier, not model selection.
- 74% will prioritize AI investment even during a recession — AI spend is now non-discretionary infrastructure for most large-company C-suite leaders.
- US CEO finding: 64% of US CEOs (n=100, revenues >$500M) report GenAI returns at or above expectations — directional corroboration for the “satisfaction is higher than the narrative suggests” pattern.
Credibility note: MEDIUM-HIGH. KPMG is a professional services firm with direct commercial interest in AI advisory. n=2,110 with rigorous revenue and seniority filtering across 20 markets — one of the larger C-suite surveys available. The 4x talent multiplier is a segmented correlation, not an experiment; organizations that invest in talent may differ in other ways. Consistent with Deloitte n=3,235 and IBM IBV n=2,007 on talent-investment findings.
Source: research/07-adoption-challenges/kpmg-global-ai-pulse-2026.md
EY CEO Outlook 2026 — The AI Accountability Gap (n=1,200, FT Longitude, Jan + May 2026)
- Only 11% of global CEOs tie AI impact to financial reporting reviewed regularly by senior management — despite 80% of the same CEOs committing to increased AI investment. This 11% threshold is the CEO-level version of the governance gap that McKinsey’s 6% EBIT cohort and Davenport’s 32% profit-attribution finding describe from the deployment side.
- AI investment held at 80% through a geopolitical shock: Between September 2025 and April 2026, CEO citations of geopolitical risk as their top concern doubled (28%→56%). AI investment intention stayed flat at 80%. For CFOs modeling AI budget risk, this is the clearest available evidence that AI spend has crossed from discretionary to structural.
- Only 20% of CEOs say AI has significantly exceeded expectations (Wave 1, Jan 2026). In financial services specifically, 82% of FS CEOs (n=240) report AI at or above expectations — a gap that reflects FS’s 5–10 year head start on AI infrastructure.
- AI delivering enterprise-level impact by function (Wave 2, May 2026): customer value creation 42%, operations 41%, innovation 40%, strategy 41%.
- 48% of CEOs are pursuing acquisitions or divestments specifically to access AI capabilities — the clearest signal that organic AI capability development is failing to keep pace with strategic need.
Credibility note: MEDIUM-HIGH. FT Longitude (Financial Times subsidiary) conducted independent fieldwork; EY-Parthenon commissioned the survey and has advisory commercial interest. n=1,200 global CEOs, 21 countries, 5 industries, quarterly waves. The 11% accountability gap stat is self-damaging relative to EY’s advisory positioning, which adds credibility. Consistent with NBER (90% no past measured impact), Davenport (32% tie AI to revenue), Writer (75% strategy theater).
Source: research/04-consulting-firms/ey-ceo-outlook-2026.md · Jan + May 2026 · TIER 1
AlixPartners Enterprise Software Predictions 2026 — AI Maturity Valuation Differential
AlixPartners’ analysis of 58 publicly traded SaaS companies (SaaS Capital Index + AlixPartners) quantifies the market’s pricing of AI maturity at the firm level — the clearest public-market evidence for the proposition that AI program depth translates to financial value:
- AI-mature firms (n=13, proven AI revenue generation): 10.4x EV/NTM Revenue | 20% median revenue growth | +11% median operating income
- AI-advanced firms (n=23, AI in core products, measurable adoption): 6.7x | 13% growth | −4% operating income
- AI-emerging firms (n=22, AI in development/pilot): 5.4x | 12% growth | −1% operating income
The 10.4x vs. 5.4x gap represents a 93% valuation premium for demonstrated AI revenue generation over AI feature development. Investors are pricing AI maturity in three components: AI leverage ratios (revenue/margin growth relative to AI cost base), outcome-based performance benchmarks (customer impact, task completion speed, output per employee), and data asset quality. This is the same 4-5x differential that KPMG’s governance-as-accelerator finding (3-6x higher improvement rates) and Roland Berger’s “Industrializer” vs. “majority” split approach from different measurement angles.
Also notable: 20-30% coding productivity gains from AI tools are not converting to reduced R&D spending or faster product delivery in the majority of software companies — directly corroborating the Goldman Sachs macro finding and BCG’s 5%-substantial-gains cohort from the software vendor side.
Source: research/07-adoption-challenges/alixpartners-enterprise-software-predictions-2026.md
Goldman Sachs AI Economic Research Cluster (Feb–May 2026)
Macroeconomic analysis by Goldman Sachs economics division (Hatzius, Walker, Briggs, Dong) — the highest-authority investment bank macro check on enterprise AI ROI claims.
- No economy-wide productivity signal: Goldman finds “no meaningful relationship between productivity and AI adoption at the economy-wide level” as of early 2026. This directly triangulates the NBER w34836 finding (90% no past impact, n=6,000 executives) from the financial-markets angle.
- GDP contribution “basically zero” in 2025: Hardware imports offset domestic investment multiplier. $667B in 2026 hyperscaler capex produces only 0.1–0.2 pp GDP contribution (Hatzius, Feb 2026).
- The 30% is real and narrow: Organizations that measured AI impact reported a median 30% productivity gain in customer support and software development specifically. These are use cases with structured inputs, defined outputs, and natural baselines.
- Only 1% of S&P 500 companies discussing AI quantified earnings impact. 70% discussed AI; 10% quantified use-case impact; 1% quantified earnings impact. This gap is the ROI-gap mechanism.
- 19% real organizational adoption (Census BTOS, March 2026) vs. 70–90% in employee-tool surveys. The difference: organizational deployment vs. individual access.
- FOMO > ROI as investment driver: Competitive anxiety is the primary AI investment motivator in Goldman’s May 2026 research — not demonstrated returns. Matches NBER w34836 perception-gap and Writer 75% strategy-theater findings.
Cross-reference: See also [[productivity-rcts]] Goldman Sachs section for the full methodology detail.
Source: research/01-ai-native-landscape/goldman-sachs-ai-economic-research-2026.md · Feb–May 2026 · HIGH · TIER 1
Macquarie Bank + General Mills — Named Enterprise AI Deployments with Primary-Sourced Metrics (2025–2026)
Two of the most credibly-sourced enterprise AI/agentic deployment cases in the public domain, both from non-vendor primary disclosure contexts.
Macquarie Bank — SRE Agentic Platform (Dynatrace, Feb 2026):
- 99.98% platform availability (6,000+ microservices, 8,000+ releases/year)
- 79% faster incident detection; 59% fewer critical incidents; 80% reduction in defects
- Named executive: Phillip Grasso-Nguyen (Head of Engineering Excellence and Reliability), Dynatrace Perform presentation
- Agentic framing: “AI agents that are doing automatic diagnostics and running runbooks” — narrow scope, defined permissions, bounded action
Macquarie Bank — Workforce AI (Google Cloud / Gemini Enterprise, Oct 2025):
- 38% more customers to self-service; 40% fewer fraud false positives; 100,000+ employee hours reclaimed
- 99% of employees trained; all-staff rollout target in 6 months
- Named: Ashwin Sinha (Chief Data and AI Officer) — Google Cloud press release
General Mills — Project ELF logistics AI (CFO investor disclosure, Feb 2025):
- $20M+ logistics savings since FY2024 (CFO Kofi Bruce at investor conference — primary financial disclosure)
- Order planning: 18 hours → under 30 minutes; 70% of AI recommendations auto-accepted; 97% data accuracy
- Vendor: Palantir; $300M total savings over 3 years from broader data program
- CDTO Montemayor distinguished current generative AI (ELF) from “real agentic AI architectures” as next phase — the $20M is pre-agentic
Credibility notes: Macquarie SRE metrics are MEDIUM-HIGH (named executive, technical event, vendor context). General Mills $20M is HIGH (CFO at investor conference, primary financial disclosure). Both represent selected wins; no control group; no independent audit.
Source: research/12-agent-workers/macquarie-general-mills-named-agentic-deployments-2025-2026.md · 2025–2026 · MEDIUM-HIGH/HIGH · TIER 1–2
Roland Berger “Profitless Prosperity in AI” (n=203, Mar 2026)
Strategy consulting firm survey of 203 senior executives. No AI platform commercial interest. Cross-regional, multiple industries.
- ~90% of firms report AI returns lagging spending — consistent with Goldman Sachs (no economy-wide productivity signal), McKinsey (6% substantial gains), Gartner (72% CIOs breaking even or losing), and Bain (23% value attribution).
- Only ~10% of organizations (“Industrializers”) consistently capture meaningful financial value from AI — the strategy-consulting-firm confirmation of the McKinsey 6% / BCG 5% / PwC top-20% pattern, from an independent source with no AI product stake.
- 99% report formal leadership involvement — the gap between executive attention and financial return is now the central diagnosis. Having the right people in the room is table stakes, not a differentiator.
- ~40% rely primarily on off-the-shelf solutions — buy-heavy deployment patterns limit internal capability compounding; not the right-or-wrong question, but a mechanism that explains why wrappers plateau.
- Two root causes of the gap: (1) no continuous value-steering metrics (one-off assessments, intuition, sprawling KPIs); (2) shallow integration (“bought the Ferrari, running it on a go-kart engine” — wrappers that collapse under operational load). Both are operating model failures, not technology failures.
- Roland Berger’s framing of the decisive constraint: “leadership is now the constraint. Technology readiness is no longer the binding limitation.”
Source: research/04-consulting-firms/roland-berger-profitless-prosperity-ai-2026.md · Mar 2026 · MEDIUM-HIGH · TIER 1
Vanguard AI Program: $500M Across 5 Use Cases (Oct 2025)
Davenport/Bean case study in MIT Sloan Management Review. Named sources: Nitin Tandon (CIO), Ryan Swann (CDAO). Single-firm case study; no independent audit. TIER 2 (Oct 2025 fieldwork).
- ~$500M measurable business value from a portfolio of five use cases: contact center (Crew Assist / Azure OpenAI), adviser intelligence, retail digital advisor, investment analytics (dividend prediction LLM), and developer productivity.
- 25% developer productivity gain from AI-assisted code generation — consistent with CMU/Stanford benchmarks; at the top of the measured range, suggesting strong workflow redesign accompanying tool deployment.
- 50% workforce AI Academy completion (out of ~20,000 employees) — training-first sequencing predates broad tool deployment; directly corroborates Conference Board / Gartner training-multiplier evidence.
- 5x dividend-cut prediction accuracy: LLM trained on 22,000 earnings call transcripts; companies flagged as “negative” were nearly 5x more likely to cut dividends within one month — a verifiable, auditable alpha signal, unlike productivity claims relying on self-report.
- ~50% call admin time reduction from call preparation and summarization tools for 150,000+ external advisers.
- Governance stance: named tools “Crew Assist” (not “AI replacement”), applied AI where decisions are made (not where easy to deploy), treated human-machine collaboration as the design constraint. Ryan Swann: “The best outcomes when humans and machines collaborate, not compete.”
Source: research/01-ai-native-landscape/vanguard-ai-roi-500m-case-study-2025.md · Oct 2025 · MEDIUM-HIGH · TIER 2
Grant Thornton 2026 AI Impact Survey: The Integration Premium (n=950, Feb–Mar 2026)
Cross-industry survey of 950 U.S. C-suite and senior business leaders, fielded February 23 – March 18, 2026. Published April 2026. Grant Thornton Advisors LLC — no direct AI product commercial interest. MEDIUM-HIGH credibility. TIER 1.
The central ROI finding: integration stage, not AI capability level, predicts whether revenue growth materializes.
- 58% vs. 15%: Organizations with fully integrated AI report AI-driven revenue growth vs. those still piloting — a ~4x gap. This is a population-level outcome difference, not a case study claim.
- Corroboration: McKinsey State of AI Nov 2025 — 6% of companies capture >5% EBIT impact; BCG AI at Work 2025 (n=10,635) — 5% of organizations report substantial financial gains. Grant Thornton adds the mechanism: pilot status versus integration status, not AI sophistication, predicts which side of the outcome distribution an organization falls on.
- The pilot trap: 85% of organizations still in pilot mode report no AI-driven revenue growth. Piloting is not a safe holding position — it is the default path to no measurable return.
- 78% governance failure rate: 78% of executives lack confidence they could pass an independent AI governance audit within 90 days. Boards approved 74% of major AI investments while 48% have not set governance expectations. The spend is decoupled from the oversight.
- Strategy vs. execution gap: 51% name strategy as the biggest ROI driver. Only 22% of operations leaders have a fully developed, implemented AI strategy. The variable most cited as predictive of success is the one most absent from actual operations.
Source: research/04-consulting-firms/grant-thornton-ai-impact-survey-2026.md · Feb–Mar 2026 fieldwork, April 2026 published · MEDIUM-HIGH · TIER 1
Developer Conference Case Studies — What Enterprises Actually Showed on Stage (2025–2026)
Source: research/07-adoption-challenges/developer-conference-enterprise-ai-case-studies.md
A structured inventory of enterprise AI case studies from GitHub Universe, Google Cloud Next, AWS re:Invent, and Microsoft Ignite/Build — with credibility analysis of each.
- Conference case studies are existence proofs, not base rates. BCG AI at Work 2025 (n=10,635) finds only 5% of organizations capture substantial financial gains; conference stages present selected successes with no denominator.
- Pattern of credible case studies: narrow scope + measurable workflow + existing data advantage. Danfoss (80% email-order automation), Macquarie Bank (40% fewer fraud false positives), Super-Pharm (50%→90% inventory accuracy) succeed because they targeted repetitive, data-rich processes — not “AI transformation.”
- Vendor-funded studies dominate. GitHub/Accenture RCT (8.69% pull request increase — sample size and duration undisclosed). Forrester independent analysis shows Microsoft Copilot workplace conversion of 35.8% — nearly two-thirds of access holders are non-users. Google ROI of AI 2025 (n=3,466, Google-commissioned, National Research Group) reports 74% first-year ROI from a pre-screened gen-AI deployment population.
- Agentic AI was the dominant 2026 conference theme (GitHub Agent HQ, 30+ AWS agentic sessions, Google/Microsoft agent governance) — but no conference presented production metrics for autonomous enterprise agents. Treat as roadmap signal, not deployment evidence.
Credibility: MEDIUM — vendor-funded studies flagged per source; DORA/Forrester independent analyses rated separately HIGH/MEDIUM-HIGH. TIER 1 (2025–2026 conference cycle).
Cisco AI Readiness Index 2025 — ROI Measurement Gap (n=8,039, Aug 2025)
n=8,039, 30 markets, August 2025. Source credibility: MEDIUM-HIGH (Cisco vendor; infrastructure commercial interest; independent Satori analysis). TIER 1.
- Only 32% of organizations have a formal process to measure AI investment impact — yet 69% rank AI as their top IT budget priority. The majority are investing without measurement infrastructure, which makes rational ROI assessment structurally impossible.
- 80% report that urgency to demonstrate ROI has risen sharply in the past 6 months. Board and CFO pressure is intensifying while measurement capability lags.
- Pacesetters (13%) close the measurement gap: 95% have AI impact measurement processes (vs. 32% overall). The measurement discipline is a precondition for Pacesetter outcomes, not a consequence.
- Financial returns are real for those who measure and deploy rigorously: 91% of Pacesetters met/exceeded expectations for profitability (vs. 64% all companies); 92% for revenue (vs. 63%); 89% for new revenue streams (vs. 61%).
- 30% of all companies expect 50–100% ROI within the next year (vs. 48% of Pacesetters). Expectation inflation is industry-wide; delivery at that rate is not.
Source: research/07-adoption-challenges/cisco-ai-readiness-index-2025.md · Aug 2025 · MEDIUM-HIGH · TIER 1
Morgan Stanley AI Adoption Survey 2026 — Enterprise ROI Evidence (Survivorship-Selected Cohort)
Survey of n=935 corporate executives (US, Germany, Japan, Australia) in 5 AI-exposed sectors, restricted to firms using AI for 12+ months. Published February 5, 2026. Source credibility: MEDIUM (survivorship-selected, self-reported; named analysts Michelle Weaver and Stephen Byrd). TIER 1.
- 11.5% average net productivity gain in the active-deployer cohort — highest sector: healthcare (+20%+), lowest: real estate. Distribution not a single average: 14% of companies exceed 20% gains.
- 4% net headcount reduction is the aggregate employment signal. The mechanism is vacancy non-backfill, not mass layoffs.
- UK at −8% net jobs is the most severe country-level displacement effect in the corpus from a named financial institution source.
- The survivorship caveat is critical: This sample excludes companies that tried AI and stopped, companies that never deployed, and companies whose deployments failed. Comparing the Morgan Stanley 11.5% to BCG’s 5% substantial-gains finding (general enterprise population): the same economics viewed through a survivorship vs. general-population lens, not a contradiction.
Position in the ROI evidence stack: Morgan Stanley fills the “what do 12-month active deployers actually report” gap. The general-population studies (BCG 5%, McKinsey 6%, Deloitte 34% transforming core processes) capture the full enterprise distribution including non-deployers. Morgan Stanley captures the leading edge. The leading edge matters for forecasting where the median will be in 18–24 months.
Source: research/05-analyst-firms/morgan-stanley-ai-adoption-survey-2026.md · Feb 2026 · MEDIUM · TIER 1
Atlassian State of Teams 2026 — The Executive ROI Confirmation Gap (n=12,035, Jan–Feb 2026)
n=12,035 knowledge workers + 173 Fortune 1000 executives, double-blind, Jan–Feb 2026. Source credibility: MEDIUM-HIGH (Atlassian commercial interest; consistent with BCG/McKinsey/Deloitte ROI findings). TIER 1. The Atlassian dataset contributes two unique data points to the ROI evidence stack: the executive self-assessment angle (can you confirm ROI?) and the fragmentation cost calculation ($161B/year).
- Only 6% of executives can confirm clear, organization-wide AI ROI — despite 89% reporting speed gains. This is the most granular measure of the gap between perceived productivity and realized financial impact in the 2026 corpus. Corroborates BCG 5% substantial gains (financial performance) and McKinsey 6% EBIT impact cohort from an independent executive-confirmation angle.
- $161 billion annual fragmentation tax across the Fortune 500 — coordination overhead that AI tools have not resolved because they were deployed on top of unchanged structural fragmentation. The largest single dollar figure for AI-adjacent waste in the corpus.
- The 14% team-level ROI cohort achieves returns through planning/prioritization integration (5.6x), collaboration improvement (9.4x), and worker trust (2.3x) — behavioral markers of workflow-embedded AI, not tool-access AI.
Source: research/07-adoption-challenges/atlassian-state-of-teams-2026.md · Jan–Feb 2026 · MEDIUM-HIGH · TIER 1
INSEAD/HBS RCT 2026 — The Mapping Problem: Task Gains ≠ Business Results (n=515)
n=515 high-growth startups; randomized controlled trial; INSEAD AI Founder Sprint; March 2026. Source credibility: HIGH (independent academic RCT, no vendor commercial interest; INSEAD + Harvard Business School). TIER 1.
- Task-level AI gains do not automatically produce firm-level ROI. Treatment (structured production-mapping case studies) vs. control (equal tool access + equal technical training): treatment firms achieved 1.9× revenue, +18% paying customers, +12% tasks completed, +44% AI use cases discovered, −$224K (−39.5%) external capital required — with no headcount change.
- The bottleneck is organizational, not technical. Treatment effect showed zero variation by founder engineering background or baseline performance. The constraint is knowing which workflows to reorganize end-to-end, not skill with the tools.
- Revenue gains are upper-tail concentrated. Benefits concentrated in the 90th–95th percentile — firms that rebuilt complete production chains, not those adding AI to individual tasks. The median treated firm showed operational gains without equivalent structural revenue shifts.
- Implication for ROI programs: Enterprise AI programs that subsidize tool access and prompting training address the wrong constraint. The investment that moves the needle is structured workflow-mapping before deployment.
Source: research/01-ai-native-landscape/insead-hbs-mapping-ai-production-rct-2026.md · March 2026 · HIGH · TIER 1
VentureBeat Q1 2026 + LinkedIn CTO — The GPU-to-ROI Transition
VentureBeat analyst Rob Streche Q1 2026 enterprise AI infrastructure survey; Iran Berger (CTO, LinkedIn) on-record practitioner account. May 2026. Source credibility: MEDIUM-HIGH (VentureBeat; n not disclosed; Cisco/OutShift sponsor). TIER 1.
- The ROI pressure point has shifted from access to economics. GPU availability concern fell from 20.8% to 15.4% in Q1 2026; cost-per-inference/TCO concern jumped from 34% to 41% in the same period. Organizations that spent two years securing compute are now being asked what it returned.
- 72% lack sufficient infrastructure control — meaning most organizations cannot answer basic questions about per-feature AI cost at current or future scale. The measurement gap documented by Cisco (only 32% have formal ROI measurement) extends into the infrastructure layer.
- LinkedIn’s production-gate model: ROI analysis required before any AI feature reaches 100% of production. Three components: cost instrumentation at projected scale, business metric linkage to revenue, and opportunity-cost modeling against alternative compute uses. Most enterprises complete step 1 partially and skip steps 2 and 3.
- 60–80% of hyperscaler AI workloads are now inference, not training. Enterprises that built governance and cost models for training-era economics need to reconfigure for inference-era unit economics.
Source: research/02-corporate-tools/linkedin-ai-infrastructure-roi-2026.md · May 2026 · MEDIUM-HIGH · TIER 1
KPMG Global AI in Finance 2026 — Finance ROI Satisfaction vs. Actual Value (n=1,013, March 2026)
KPMG’s global finance-function primary survey adds three findings to the ROI evidence stack that no existing corpus source covers:
- 71% of finance leaders report AI meeting or exceeding ROI expectations — but only 23% say exceeding. The satisfaction rate is directionally high; the “exceeding” rate is low. The pattern matches the broader corpus: Goldman Sachs finds “basically zero” GDP contribution; NBER w34984 finds a 3x gap between CFO-perceived and revenue-implied productivity gains; McKinsey finds only 6% achieve >5% EBIT impact. KPMG’s finance-specific data confirms that satisfaction-level metrics are not evidence of financial impact.
- Governance as ROI multiplier — 3–6x: Organizations producing AI audit evidence efficiently show 3–6x higher improvement rates (error reduction: 33% vs. 6%; confidence in scaling: 42% vs. 14%). This is the first finance-function quantification of the governance-ROI link that McKinsey’s RAI benchmark ($25M+ → EBIT >5%) and Grant Thornton’s integration premium (4x revenue growth) document at the enterprise level. The mechanism: governance infrastructure that enables audit evidence generation is the same infrastructure that catches errors before they reach financial statements.
- Agentic AI deployers show 32pp average performance gap; nearly 40pp on forecast accuracy and ROI. This is the sharpest performance differential in the finance-function AI corpus and the most direct evidence that agentic deployment — not just AI deployment generally — creates measurable competitive separation. The gap is consistent with the broader corpus pattern (PwC: 74% of economic value captured by top 20%; BCG: 5% of organizations report substantial financial gains) but finance-specific.
- Cross-validation with NBER w34984: The CFO productivity paradox (perceived 1.8% vs. revenue-implied 0.6% gains) and the KPMG 71%/23% satisfaction gap are measuring the same phenomenon from different angles. Both confirm that satisfaction is not a proxy for financial impact, and that the organizations actually capturing value are a distinct minority.
Source caveat: KPMG has commercial interest in finance AI transformation engagements. Survey is self-reported; no control group. The 71% satisfaction and 32pp agentic gap figures require independent corroboration — they are not available from non-vendor sources as of May 2026. Apply MEDIUM-HIGH credibility; treat as directional, not audited.
Source: research/04-consulting-firms/kpmg-global-ai-in-finance-2026.md · March 2026 fieldwork, May 2026 published · MEDIUM-HIGH · TIER 1
MIT Technology Review / EDB 2026 — Sovereignty Architecture as ROI Predictor (n=2,050+, May 2026)
MIT Technology Review Insights / EnterpriseDB survey (n=2,050+ senior executives, 13 countries, May 14, 2026). Vendor-sponsored (EDB); directionally corroborated by Gartner 4x data-foundations investment differential (n=353). TIER 1.
- Organizations deeply committed to AI and data sovereignty achieve 5x higher ROI from generative and agentic AI — with a 0.93 correlation coefficient between sovereignty posture and AI success outcomes. This is the highest single-predictor correlation in the 2026 corpus.
- The 0.93 correlation does not establish causation. A simpler interpretation: organizations with sovereignty architecture are also organizations with data governance, measurement discipline, and workflow redesign maturity. These are not separable from the sovereignty finding — they are the same organizations that consistently outperform on every AI ROI measure in the corpus (McKinsey 6%, BCG 5%, Grant Thornton 4x, Gartner 4x foundations differential).
- 70% of executives believe a sovereign data and AI platform is necessary to succeed — but intent and architecture diverge. The gap between intent and implementation maps to the adoption-outcome gap documented throughout the corpus.
- Governance-ROI link: The sovereignty data adds an architectural dimension to the existing evidence chain. The mechanism is: control architecture → measurable outcomes → governance credibility → faster procurement and scaling. This sequence is consistent with Cisco’s finding that Pacesetters (mature governance) are 4x more likely to move pilots to production.
Source: research/06-security-frontier/mit-tr-edb-ai-data-sovereignty-2026.md · MIT TR / EDB May 2026 · MEDIUM-HIGH · TIER 1
McKinsey State of Organizations 2026 — 88% Deploy, 81% See Nothing (n=10,018, March 2026)
McKinsey’s largest annual organizational survey (n=10,018 senior executives, 15 countries, 16 industries, fieldwork June–September 2025) delivers the starkest single-survey quantification of the deployment-impact gap in the 2026 corpus. TIER 1.
- 88% of organizations are deploying AI; 81% report no meaningful bottom-line impact. This is not a niche finding — n=10,018 is the largest survey sample yet to confirm the pattern that BCG (5% substantial gains, n=10,635), McKinsey QuantumBlack (6% EBIT impact, n=1,993), Deloitte (34% transforming core processes, n=3,235), and Writer (29% significant ROI, n=2,400) have all documented from different angles.
- Only 1% of US C-suite leaders describe their AI rollouts as mature. The maturity gap is the clearest single-number proxy for why 81% see no impact: deployment is not maturity.
- $5:$1 people-to-technology investment ratio — organizations applying this ratio are 4.3x more likely to sustain top-tier financial performance over the next decade. This ratio directly quantifies what the corpus has documented qualitatively: BCG’s 10-20-70 (10% technology, 20% data, 70% people), McKinsey’s workflow-redesign finding (21% fundamentally redesigning = the ROI cohort), PwC’s finding that top 20% capture 74% of economic value.
- Leadership ownership gap: Only 14% of leaders consistently champion AI with a clear strategy; 1 in 6 organizations has no C-suite AI owner. The C-suite clarity gap (56% → 27% at middle management, a 29-point silo gap) is the operational mechanism that turns technology deployment into organizational theater.
- Cross-validation: The 88%/81% finding is independently corroborated by BCG 5%, McKinsey QuantumBlack 6%, Deloitte 34%, Writer 29%, Capgemini 38% operationalized, and Futurum 21.7% reporting direct financial ROI as primary metric. These are five independent datasets from different firms, different samples, and different methodologies all converging on the same range: 5–19% of deploying organizations are capturing meaningful value.
Source: research/01-ai-native-landscape/mckinsey-state-of-organizations-2026.md · McKinsey State of Organizations, n=10,018, March 2026 · MEDIUM-HIGH · TIER 1
Wharton / GBK Collective — Three-Year Longitudinal ROI Evidence (n=801, October 2025)
The only repeated cross-sectional academic series tracking enterprise AI ROI measurement at consistent sample criteria (1,000+ employees, >$50M revenue, U.S.) over three years. TIER 1 (October 2025). MEDIUM-HIGH credibility.
- 72% of enterprises now formally measure Gen AI ROI using business-linked metrics (profitability, throughput, workforce productivity). This is a structural change from prior waves — not a marginal shift. HR (84%) and Finance (80%) lead; Legal lags.
- 74% report positive ROI to date. By sector: Tech/Telecom 88%, Banking/Finance ~83%, Professional Services ~83%, Manufacturing 75%, Retail 54%. Negative ROI is rare at under 7%.
- 80% expect positive returns within 2–3 years. This is the most useful single figure for CFO expectation-setting. It directly contradicts the widespread assumption of 6–12 month payback. Consistent with Deloitte 2025 (n=1,854) finding of a 2–4 year ROI horizon.
- The seniority gap is real and consistent: VP+ report 81% positive ROI; mid-managers report 69%. VP+ are twice as likely to report “significantly positive” ROI (45% vs. 27%). This pattern appears in McKinsey State of Organizations (2026) and IBM IBV (2026) — a consistent signal that C-suite optimism and ground-level reality diverge.
- Large enterprises ($2B+ revenue) are 34% “neutral/too early” on ROI vs. ~9% for mid-market firms — scaling complexity is measurable.
- Budget reallocation is beginning: 11% of enterprises now fund AI by cutting legacy IT and HR/Workforce programs (+7pp year-over-year). Still minority behavior; net-new budget dominates.
Cross-validation: Corroborates McKinsey State of Organizations 2026 (88%/81% deployment-impact gap), Futurum 1H 2026 (hard financial ROI as primary metric doubles to 21.7%), and Fed Atlanta/NBER w34984 (CFO productivity perception gap). All converge on the same picture: positive ROI is real for a majority but concentrated in specific sectors, firm sizes, and adoption cohorts.
Source: research/01-ai-native-landscape/wharton-gbk-enterprise-ai-year3-2025.md · Wharton/GBK Collective, n=801, October 2025 · MEDIUM-HIGH · TIER 1
BEA Working Paper WP2026-3: The First U.S. Government Measurement of AI Productivity Gains (February 2026)
Source credibility: HIGH. TIER 1. U.S. Bureau of Economic Analysis — government economists, national accounts data, no commercial interest, commissioned by the White House AI Action Plan. Methodology: difference-in-difference using Census Annual Business Survey (~230,000 businesses) and BEA-BLS Integrated Industry-Level Production Account, n=1,586 industry-year observations. Caveats: early estimates, alternative specification yields less robust results, AI not yet a line item in national accounts.
The first official government-data finding that AI produces statistically significant TFP gains at the industry level:
- AI-intensive industries (top quartile of organizational AI adoption) experienced TFP growth approximately 2% per year higher than non-AI-intensive industries after 2021 (p<0.01)
- Average Labor Productivity (ALP) was approximately 1% per year higher in AI-intensive industries post-2021 (p<0.1)
- AI appears labor-saving and input-saving: AI-intensive industries show lower labor contributions from both college and non-college workers, and lower intermediate-input use (particularly services)
- Only 4.4% of U.S. private companies had AI in production processes as of 2023 (Census BTOS); 5.2% at peak in 2022. These are organizational adoption rates, not individual usage rates.
- AI-intensive industries consistently: Information (NAICS 51), Professional/Scientific/Technical Services (54), Management of Companies (55), Real Estate (53). Sector divergence widening: top-quartile AI threshold nearly doubled (3.8%→7.1%) while bottom-quartile stayed flat.
Reconciliation with Goldman Sachs “no economy-wide productivity”: Both findings are correct. Goldman finds no economy-wide relationship because only 4-5% of companies have reached embedded production AI. BEA finds significant gains within that 4-5% top-quartile cohort. The macro signal is diluted by the base; the cohort signal is real.
Source: research/01-ai-native-landscape/bea-wp2026-3-ai-industry-accounts.md · BEA WP2026-3, Highfill & Samuels, February 2026 · HIGH · TIER 1
Forrester TEI: Microsoft 365 Copilot — Upper-Bound ROI Estimates and True Cost Structure (March 2025)
Source credibility: LOW-MEDIUM. TIER 3. Commissioned by Microsoft; Forrester TEI format using 12 self-selected Microsoft reference customers. No control group. Selection bias inherent. Do not cite 116% ROI without the vendor-commissioned caveat.
- 116% 3-year risk-adjusted ROI, $19.7M NPV, 10-month payback for a modeled composite enterprise ($6.25B revenue, 25,000 employees). These are model outputs from a composite fictional organization — treat as upper-bound estimates.
- 9 hours/month user time savings — the single most citable figure; user-reported and consistent across multiple independent M365 Copilot studies.
- True cost structure: $17.1M total costs vs. $5.8M licensing over 3 years. Training ($6.9M) and implementation ($4.4M) represent 66% of total investment — organizations budgeting only license cost will underestimate TCO by nearly 3x.
- Use for: establishing implementation/training cost categories (real regardless of ROI model), and the 9-hours-per-user productivity benchmark.
- Do not use for: citing 116% ROI as a reliable benchmark, or projecting revenue-side gains without independent validation.
Source: research/05-analyst-firms/forrester-tei-m365-copilot-2025.md · Forrester/Microsoft, March 2025 · LOW-MEDIUM · TIER 3
NBER WP35046: Why Expert Economists Expect Modest GDP Gains Despite AI Progress (April 2026)
Source credibility: HIGH. TIER 1. NBER Working Paper 35046, multi-group forecasting tournament: 69 economists (AI/growth specialists), 38 superforecasters, 27 AI industry professionals, 25 AI policy researchers, 401 general public. Survey October 2025–February 2026. Philip Tetlock (Wharton), Jason Abaluck (Yale/NBER), Federal Reserve economists. Open Philanthropy funded — disclosed.
The canonical “why AI won’t show up quickly in GDP” finding, from the people best positioned to assess it:
- 61.4% of economists expect moderate or rapid AI progress by 2030 — yet their unconditional GDP forecast is just 2.5% annualized, barely above government baselines of 1.9–2.1%
- The most cited reason in written rationales: diffusion lag — economists drew direct analogies to electrification, automobiles, and personal computers; multi-decade lags routinely separate GPT arrival from measurable productivity impact
- Rapid-scenario GDP forecast by 2045-2049: 3.5% (economists), 5.3% (AI experts) — “not historically unprecedented” per paper; comparable to post-WWII growth
- Primary source of expert disagreement: not “will transformative AI arrive?” — it’s “if it arrives, will the economic gains be broadly distributed?” — a policy question, not a technical one
- Top-10% wealth share rises from 71.2% (2023) to ~80% by 2050 under rapid AI (economist forecast) — inequality risk is the macro growth constraint, not capability limits
For CFO/board conversations: This is the authoritative answer to “why don’t we see AI in productivity numbers yet?” — because the people most equipped to forecast it expect diffusion, structural headwinds, and energy bottlenecks to delay the signal for years, even if capability arrives on schedule.
Source: research/01-ai-native-landscape/nber-karger-forecasting-economic-effects-ai-2026.md · NBER WP35046, Karger/Kuusela/Abaluck et al., April 2026 · HIGH · TIER 1
EY-Parthenon: The 63/14/7 AI Use Distribution — Efficiency vs. Competitive Strategy (April 2026)
Source credibility: MEDIUM-HIGH. TIER 1. EY-Parthenon (EY strategy consulting) survey of 271 US growth leaders at $500M+ companies. Fieldwork February 19–March 12, 2026. Published April 28, 2026. U.S.-only; ±6pp margin of error. Corroborated by PwC 74/20 value concentration and BCG 5%/McKinsey 6% high-performer cluster.
- 63% using AI for efficiency/productivity only — defensive posture; table-stakes benefits with no competitive differentiation
- 14% using AI to stay ahead of competitors — the cohort translating AI capability into market positioning
- 7% using AI to diversify revenue streams — smallest group, creating new lines of business
- 78% believe AI will accelerate growth yet trust AI for pricing (34%), new product development (28%), M&A evaluation (27%) — the strategy-level decisions where AI advantage would actually matter
- 41% fear AI enables new competitive market entrants — the organizations using AI defensively are facing offensive pressure from AI-enabled entrants
- 49% “strongly agree” organization knows how to leverage data and AI for growth — more than half of growth leaders lack foundational AI capability their own strategy depends on
- Cross-reference: PwC (74% of AI economic value captured by 20%), BCG (5% substantial gains), McKinsey (6% >5% EBIT impact) — the EY-Parthenon data provides the mechanism for that concentration: it’s not random, it’s the 14%/7% who deploy AI strategically
Source: research/04-consulting-firms/ey-parthenon-growth-strategy-ai-2026.md · EY-Parthenon, n=271, April 2026 · MEDIUM-HIGH · TIER 1
Oliver Wyman Forum / NYSE — CEO Agenda 2026 (n=415, $13T market cap, April 2026)
The most capital-weighted CEO AI ROI survey available. The year-over-year trend is the signal: confidence in AI returns declined, not increased.
- 27% report AI ROI meeting or exceeding expectations — down from 38% in 2025; a 29% relative decline in one year
- 12% qualify as “AI ROI leaders” (>10% enterprise-wide impact) — down from 17% in 2025
- 53% say it’s too early to assess AI ROI — up from 41% in 2025; the “too early” cohort is growing as expectations exceed measured returns
- 67% remain in planning or pilot stages — at $1B+ revenue threshold; two-thirds of major companies have not scaled AI
- Workflow redesign gap reappears: 49% of deployment leaders redesign workflows vs. 32% average — the same mechanism BCG, PwC, and McKinsey identify as the differentiator
- Size amplifier: nearly 5x more mega-size firms report >10% AI cost savings vs. midsize companies
- Geography gradient: Asia-Pacific 47% deployment leaders; North America 37%; Europe 30%
- Year-over-year deterioration is directionally consistent with Roland Berger (~90% returns lagging spend), Gartner (72% CIOs break-even or losing), McKinsey (6% substantial gains), PwC (74% value concentrated in top 20%)
Source: research/04-consulting-firms/oliver-wyman-ceo-agenda-ai-workforce-2026.md · Oliver Wyman Forum / NYSE, n=415, April 2026 · HIGH · TIER 1
Futurum Group 1H 2026 Enterprise Software Survey — ROI Metric Shift (n=830, Feb 2026)
The first large-sample primary survey to document the enterprise AI ROI metric inflection: direct financial impact replacing productivity as the leading success measure.
- Direct financial impact (revenue + profitability) nearly doubled as the primary ROI metric — rising to 21.7% from roughly 11% in the prior period. Productivity gains fell from 23.8% to 18.0% (−5.8 percentage points).
- CFO accountability pressure is the mechanism. Productivity metrics (hours saved, task speed) are owned by department heads. P&L metrics are owned by CFOs and boards. The metric shift reflects who is now asking the accountability question in budget reviews.
- Corroboration from Forrester “Three Years Into GenAI” (n=1,500, Apr 2026): Only 15% of AI decision-makers report an EBITDA lift after three years of deployment — confirming the gap between the new metric demand and actual outcomes achieved.
- The survey also documents the agentic priority surge (17.1% rank agentic AI as #1, up from 13.0%; +31.5% YoY) — consistent with the ROI metric shift: assistive AI produces task-level productivity metrics; agentic AI can produce P&L-connected outcomes.
Source: research/09-ai-adoption-cycle/futurum-enterprise-ai-roi-shift-2026.md · Futurum Group, n=830 global IT decision-makers, February 17, 2026 · MEDIUM-HIGH / TIER 1
Gartner I&O AI ROI Survey — Infrastructure-Level Failure Rates (n=782, Nov–Dec 2025)
The only primary dataset that measures AI project success rates specifically among infrastructure and operations practitioners — distinct from enterprise-wide executive surveys.
- Only 28% of AI use cases in I&O fully succeed and meet ROI expectations. 20% fail outright. The remaining 52% stall: investment committed, value unrealized.
- ITSM and cloud operations account for 53% of all successes — the most structured, highest-data-density environments in I&O. Auto-remediation, self-healing infrastructure, and agent-led workflow management fail at the highest rates.
- Root causes of failure: 57% of those who failed cited unrealistic scoping (“expected too much too fast”); 38% cited persistent skill gaps; 38% cited poor data quality.
- Success factors are governance factors: embedding AI into existing workflows (rather than running parallel experiments), executive ownership before deployment, and centralized AI product management rather than BU-funded project management.
- The 72% failure-or-stall rate is consistent with Oliver Wyman/NYSE CEO data (only 27% report AI ROI meeting expectations), Atlassian (6% org-wide ROI confirmation), and BCG AI at Work 2025 (5% substantial gains).
Source: research/07-adoption-challenges/gartner-io-ai-roi-stall-2026.md · Gartner, n=782 I&O leaders, November–December 2025 · MEDIUM-HIGH / TIER 1
IBM IBV / Oxford Economics — The 79/24 Revenue Gap (n=2,007, Q3–Q4 2025)
The largest multi-industry executive survey on AI ROI expectations vs. roadmap clarity, with independent methodology (Oxford Economics).
- 79% of executives expect AI to significantly contribute to revenue by 2030 — but only 24% can articulate where that revenue will come from. The 55-point expectation-to-roadmap gap is the most cited number from this dataset; it is consistent with Oliver Wyman/NYSE CEO data (27% reporting ROI meeting expectations) and BCG AI at Work 2025 (5% achieving substantial gains).
- 68% of executives worry their AI efforts will fail specifically due to lack of integration with core business activities. The tool is not the problem; the workflow architecture is.
- Organizations that scale AI across multiple workflows anticipate 24% greater productivity gains and 55% higher operating margins than single-use-case peers by 2030 — the multi-workflow premium is the clearest ROI lever in this dataset.
- AI investment projected to surge 150% (as % of revenue) 2025–2030. Executives expect 42% aggregate productivity gains; 67% believe they’ll capture most of those gains by 2030. The other third will invest heavily and still be chasing the return.
- IBM IBV is IBM Consulting’s research arm — vendor context applies; the integration-failure finding aligns with IBM’s commercial thesis. Oxford Economics partnership provides methodological independence. Apply vendor discount to specific percentages; directional findings are corroborated across independent sources.
Source: research/07-adoption-challenges/ibm-ibv-enterprise-2030-ai-ambition-gap-2026.md · IBM IBV / Oxford Economics, n=2,007 executives, Q3–Q4 2025 · MEDIUM-HIGH / TIER 1
Datadog State of AI Engineering 2026 — Production ROI Visibility Gap
Behavioral/telemetry data from thousands of organizations running AI in production, February–March 2026. Unique in the corpus: measures what actually happens at scale, not what executives report.
- 5% of all AI model requests failed in February 2026; 2% in March. These production failure rates are absent from most AI ROI conversations — but if 1–5% of AI interactions fail, the ROI models built on uptime assumptions are systematically overstated.
- 60% of failures in February were rate limit errors — infrastructure capacity limits, not model errors. Remediation requires architectural decisions and vendor contract changes, not prompt engineering. Organizations that cannot distinguish these failure modes cannot measure actual ROI accurately.
- Only 28% of LLM calls use prompt caching despite 60–90% cost savings on repeated-context calls. Organizations paying full price on every API call for unchanged system prompts have a direct, measurable efficiency gap that reduces ROI without appearing in model-level metrics.
- Median token/request doubled YoY; 90th percentile quadrupled. Organizations that modeled AI cost at 2024 consumption levels are running material infrastructure budget variance. CFOs lacking token consumption trend data cannot benchmark current AI ROI against initial business cases.
- The governance implication: organizations without AI monitoring infrastructure cannot distinguish infrastructure failures from model failures, cannot measure actual vs. projected costs, and cannot benchmark ROI against deployment promises.
Source: research/01-ai-native-landscape/datadog-state-of-ai-engineering-2026.md · Datadog LLM Observability telemetry, thousands of orgs, Feb–Mar 2026 · MEDIUM-HIGH / TIER 1
Beyond the Pilot — Intuit (Richard Clark, SVP/CDO, Oct 2025)
Named executive on-record with specific financial outcomes from AI deployment in invoicing/payments workflows.
- “We’re getting paid five days faster from these clients on average, 10% more likely to be paid fully on time.” — Richard Clark, SVP and Chief Data Officer, Intuit (Oct 2025, unscripted interview)
- Credibility: HIGH — named C-suite executive, specific quantified metric, production deployment, no vendor framing
- Applicability: mid-market CFOs. Five-day DSO improvement is directly measurable on any balance sheet; 10% on-time payment lift has direct cash flow implications. Neither metric requires AI attribution modeling — the outcome is verifiable in AR aging reports.
- Topic tags:
ai-roi-evidence·cfo-ai-workflows
Source: research/13-multimodal-sources/beyond-the-pilot/2025-10-08-ai-chat-70-of-enterprises-adopt-ai-agents-the-real-world-imp.md · Beyond the Pilot / VentureBeat, Oct 2025 · HIGH / TIER 2
See also
- ai-productivity-measurement-gap — why measured gains differ from perceived gains; METR RCT on AI coding; BEA national accounts; behavioral evidence on the 14% net-positive outcome rate
- productivity-rcts — the controlled experiment evidence base for task-level and firm-level gains
- workflow-redesign — the production-chain redesign that converts individual task gains to firm-level income-statement outcomes
- research/01-ai-native-landscape/real-roi-by-function.md — function-by-function ROI ranking synthesis; RCT evidence hierarchy from Brynjolfsson/Dell’Acqua/Noy-Zhang/METR; where enterprise AI is actually delivering measurable returns
The Innovation Tax: Banking-Sector Causal ROI Evidence (Kikuchi, U Tokyo, Feb 2026)
First causal study isolating AI adoption impact from selection effects in financial services. n=126 banks (41 treated, 85 controls), 2018–2025. Synthetic DiD using ChatGPT launch (Nov 2022) as exogenous shock.
- Cross-section shows +42 bps ROE for AI adopters vs. non-adopters — but this is selection, not causation. AI-adopting banks are already stronger performers.
- Causal DiD shows -428 bps ROE in adoption window — the actual financial cost of implementation before gains materialize.
- J-curve structure: performance recovers post-adoption, consistent with costs preceding benefits. Banks completing the J-curve are positioned as long-term outperformers.
- Size gap: small banks pay 517 bps vs. 129 bps for large banks. Scale economies in AI implementation are substantial.
- Policy implication: ROI timelines cited by vendors (often 12–18 months) likely reflect large-bank J-curve recovery; small/mid-size organizations should plan for a longer and steeper dip.
Source: research/06-industry-verticals/kikuchi-innovation-tax-banking-genai-2026.md · arXiv:2602.02607 · U Tokyo · Feb 2026 · MEDIUM-HIGH · TIER 1
EY AI Pulse Wave 4: Productivity Gains Are Real — But Mostly Recycled Back Into AI (Dec 2025)
The fourth wave of EY’s annual US AI Pulse Survey (n=500 SVP+, fieldwork Sept–Oct 2025) puts specific numbers on what organizations actually do with AI productivity gains.
- 96% of AI-investing organizations report some productivity gains; 57% report significant gains. Cross-reference: Deloitte SoAI n=3,235 (66% productivity/efficiency gains) and Workday n=3,200 (consistent reinvestment pattern) broadly corroborate the 96% any-gain figure.
- Only 17% reduced headcount. The other 83% reinvested in more AI (47% expanded existing capabilities; 42% developed new; 38% upskilled employees). Counterpoint: Gartner autonomous business survey (n=350, in corpus) found 41% planning workforce reductions to fund AI ROI — suggesting self-reporting to a vendor-sponsored survey may understate actual headcount pressure.
- Investment threshold effect: 71% vs. 52% significant gains at ≥$10M vs. <$10M. Consistent with the Atlan 200-deployment finding that AI ROI requires workflow redesign (fixed cost) before gains compound.
- Budget structural shift: 27%→52% expect to allocate ≥25% of IT budget to AI within 12 months. AI moving from line item to infrastructure-grade spend.
- Limitation: US-only, senior leaders only, self-reported productivity, EY commercial interest in framing that encourages deployment. TIER 2 (Sept–Oct 2025 fieldwork; results may differ with current models).
Source: research/04-consulting-firms/ey-ai-pulse-survey-wave4-2025.md · EY US AI Pulse Survey Wave 4, n=500 SVP+, Sept–Oct 2025, published Dec 9, 2025 · MEDIUM / TIER 2
Everyday AI Ep 755–760: The 97%/29% Gap and the McKinsey EBIT Finding (April 2026)
Everyday AI podcast (Jordan Wilson) episodes 755 and 760 (April 14 and 21, 2026) surface two ROI-relevant data points corroborated by named sources:
- 97% of executives report personal AI wins; only 29% see enterprise-level ROI. (Episode claim, Ep 760 — MEDIUM credibility as podcast/media source.) Consistent with McKinsey’s 6% meaningful-profit finding (Ep 755, citing McKinsey directly) and Atlassian’s 6% executive ROI confirmation rate already in this corpus.
- McKinsey: workflow redesign showed the largest EBIT impact among 25 attributes tested. (Ep 760, citing McKinsey — HIGH credibility for the underlying McKinsey finding.) Consistent with the corpus-wide finding that tools without redesign produce Stage 1 outcomes only.
- BCG: 70% of AI’s value derives from people and processes, not technology. (Ep 760, citing BCG — HIGH credibility.)
- Gallup: employees with manager support are 9x more likely to report AI transformed their work. (Ep 760, citing Gallup — HIGH credibility.)
- Budget allocation pattern observed in top performers: 5–10x more on training and process redesign than on tool licenses. (Ep 760, Jordan Wilson synthesis — MEDIUM credibility.)
Enterprise case studies cited in Ep 760: Moderna (80% internal AI tool adoption via 2,000-person weekly forum), BBVA (83% bank-wide weekly AI activity starting from 250 senior leaders), JPMorgan (flat headcount plus role reshaping).
Source: research/13-multimodal-sources/everyday-ai/2026-05-22-ep755-780-enterprise-ai-mining.md · Everyday AI podcast, Jordan Wilson · Ep 755 (Apr 14, 2026) + Ep 760 (Apr 21, 2026) · MEDIUM (podcast) / TIER 1
McKinsey Rewired (2nd ed., Lamarre/Smaje/Levin 2024) — Six-Capability Framework for AI Value
McKinsey practitioners’ book (2nd edition, 2024 — TIER 3, prior model generation). Framework remains structurally applicable; specific case evidence (E.ON Next) should be labeled directional only. Cross-referenced against TIER 1–2 independent evidence in corpus.
- 70–90% of AI value resides in people, operating model, and workflow design — not in model choice. Consistent across BCG (10/20/70 rule: 10% technology, 20% algorithm, 70% adoption/change), MIT CISR (n=721), and Stanford HAI 2026 findings.
- Six mutually reinforcing capabilities: roadmap, talent, operating model, technology, data, adoption. Missing any one capability constrains ROI across all six — this is the mechanism behind the “partial deployment” ROI failure mode.
- ROI compounds only when all six capabilities are present. Organizations that deploy the technology layer without the operating model and adoption layers achieve what the corpus categorizes as Stage 1 outcomes — tool adoption without workflow redesign — which correlates with the 66% productivity gains that don’t translate to measurable P&L (Deloitte n=3,235).
- The 2nd edition’s additions (Ch. 5 agentic workflows, Ch. 11 agentic talent model, Ch. 22 agentic engineering) extend the framework to agentic AI deployment — consistent with the corpus trend from assistive to agentic transition requiring governance infrastructure, not just tool access.
Source: research/04-consulting-firms/mckinsey-rewired-2nd-edition-synthesis.md · McKinsey Rewired 2nd ed. (Lamarre, Smaje, Levin) · 2024 · MEDIUM / TIER 3 (prior model generation — framework applicable, specific case evidence directional only)
MASAI RCT (n=105,934): The Attention Reallocation Mechanism in Clinical AI (Jan 2026)
The MASAI trial (The Lancet, Jan 31, 2026) is the largest completed RCT of AI in clinical screening — and provides the clearest mechanism evidence in the corpus for why AI ROI requires workflow redesign, not just tool deployment.
- AI-supported mammography reduced interval cancers by 12% and improved radiologist sensitivity from 73.8% to 80.5% — while radiologist workload fell 44%. The ROI mechanism: AI triaged low-risk cases to single reading, routing cognitive effort to high-signal cases.
- False positive rate unchanged — the common objection to AI in diagnostic medicine (more false alarms, more cost) did not materialize. Specificity was statistically equivalent between arms.
- Aggressive cancer rate fell 27% — meaning AI caught the cancers that mattered, not just more cancers overall.
- The mechanism — AI absorbs high-volume triage, human expertise concentrates on high-signal decisions — mirrors the ROI pattern Microsoft’s security copilot RCTs (n=167, n=162, Nov 2025) documented in phishing triage. The pattern generalizes across domains with structured, high-volume, signal-noisy inputs.
- Limitation: single country (Sweden), one vendor (Transpara Detection/ScreenPoint Medical), one domain. The ROI quantum is domain-specific; the workflow redesign mechanism is broadly applicable.
Source: research/06-industry-verticals/masai-lancet-ai-mammography-rct-2026.md · MASAI RCT, The Lancet, n=105,934, Jan 31, 2026 · HIGH / TIER 1
Gartner (n=350, May 2026): Headcount Cuts Don’t Drive ROI — Investment in Remaining Workers Does
Gartner May 2026 survey of 350 global executives ($1B+ revenue) who had already deployed autonomous AI or intelligent automation. Q3 2025 fieldwork. Named analyst: Helen Poitevin, Distinguished VP.
- 80% of deploying organizations reduced their workforce — but headcount reduction rates were statistically identical between high-ROI and low-ROI organizations. Layoffs create budget room, not return.
- High-ROI organizations shared a different pattern: they invested aggressively in the workers who remained — retraining for AI oversight, creating new roles (workflow architects, AI oversight specialists), and redesigning operating models to reflect AI-handles-execution / human-handles-judgment.
- Gartner forecasts autonomous AI will be a net-positive job creator by 2028–2029 — driven by new categories of work that do not exist today. This is a forward projection (model assumptions not disclosed); treat as directional.
- The mechanism corroborates the corpus-wide pattern: ROI compounds through workflow redesign and human amplification, not through headcount reduction alone.
Source: research/05-analyst-firms/gartner-autonomous-business-layoffs-roi-2026.md · Gartner, n=350 global executives ($1B+), May 2026 · MEDIUM-HIGH / TIER 1
Wharton WHAIR — Self-Reported ROI vs. Measured ROI: The 75%/10% Gap (n=800+, Oct 2025)
Source: Wharton Human-AI Research (WHAIR) + GBK Collective, “Accountable Acceleration,” October 28, 2025. n=800+ U.S. enterprise leaders (>1,000 employees, >$50M revenue). Self-reported. MEDIUM / TIER 1.
The most important finding in the Wharton longitudinal data is not the headline 75% positive ROI figure — it is the gap between self-reported and verified returns.
- 75% of enterprise leaders report positive ROI on Gen AI investments — the headline number from the study
- Researchers explicitly note this is self-reported rather than verified against financial data, and that optimistic skew is a known risk
- 72%+ have ROI measurement frameworks — meaning roughly 1 in 4 expressing positive ROI has no defined metric against which they are measuring it
- Calibration benchmark: D&B’s 10,000-organization survey (same period) finds only 10% report “strong ROI.” The 75% vs. 10% gap is the self-report inflation premium.
- 43% of leaders citing positive ROI simultaneously warn of skill atrophy — the contradiction at the center of 2025 enterprise AI adoption
- The study does not break down ROI by function, company size, or deployment model — the aggregate positive sentiment masks substantial variation in actual returns
Source: research/07-adoption-challenges/wharton-whair-genai-enterprise-accountable-acceleration-2025.md
Stanford SIEPR — Consumer GenAI Productivity (n=200,000+ households, 2026)
Source: research/01-ai-native-landscape/stanford-siepr-consumer-ai-productivity-2026.md
- GenAI users complete consumer digital tasks 76–176% more efficiently than non-adopters (behavioral data, not self-reported)
- For specific bounded tasks (product comparisons, tax questions, troubleshooting): 500–1,000%+ efficiency gains
- Critical qualifier: gains apply to information-seeking, bounded tasks at home — not open-ended enterprise knowledge work
- Freed time goes to leisure (+31 pp leisure browsing), not skill development (−21 pp productive browsing)
- Digital divide is widening: younger, higher-income households adopt substantially faster; the gap is not converging
McKinsey State of Organizations 2026 — The Adoption-to-Value Gap (n=10,018)
Source: McKinsey & Company, “The State of Organizations 2026.” n=10,018, 15 countries, 16 industries. Jun–Sep 2025 fieldwork, Mar 14, 2026 publication. MEDIUM / TIER 1.
- 88% of organizations experimenting with AI; 81% report no meaningful bottom-line impact — the adoption-to-value gap stated at the largest sample size in the corpus
- Only 19% report AI-accelerated revenue above 5%; only 1% describe rollouts as “mature” (U.S. C-suite)
- 23% “AI Pioneers”: the segment capturing value shares one defining pattern — workflow redesign alongside (or before) tooling deployment
- Organizations focusing on people AND performance: 4.3x more likely to sustain top-tier financial results
- McKinsey’s investment prescription: $5 in people for every $1 in technology — the ratio most organizations invert
Source: research/04-consulting-firms/mckinsey-state-of-organizations-2026.md
PwC 29th Global CEO Survey (n=4,454, January 2026)
Source: research/04-consulting-firms/pwc-ai-research-2026.md · September–November 2025 fieldwork, 95 countries · MEDIUM-HIGH / TIER 1
- 56% of CEOs report zero financial return from AI — neither revenue gains nor cost savings. Only 12% have achieved both; 22% say AI has increased their costs.
- Only 14% of workers use generative AI daily. Fewer than 25% of CEOs say AI is applied to a “large or very large extent” in any major business area.
- The vanguard (12%) shares one profile: AI deployed across more business areas (44% applied to products vs. 17% for others), investment in technology integration, defined AI roadmaps, and formalized responsible AI governance.
- 9pp TSR gap: organizations with fewest AI trust/safety concerns delivered total shareholder returns 9 percentage points higher over 12 months than those with the most concerns — a measurable governance ROI signal for board-level conversations.
- PwC chairman Mohamed Kande on the root cause: “People forgot that the adoption of technology, you have to go to the basics.” The finding replicates McKinsey’s workflow-redesign pattern and BCG’s 10-20-70 rule in a different methodology.