Data readiness is the degree to which an organization’s existing data can feed AI workloads without a remediation pass. It is the most under-budgeted component of enterprise AI deployment — data preparation consumes 40–60% of total AI project cost and 60–80% of project time (O’Reilly), yet is routinely scoped as a minor line item.

The Innoflexion Data Readiness Index defines operational thresholds: above 0.75 supports autonomous AI; 0.50–0.74 requires human review checkpoints; below 0.50 demands structured remediation before production use.

Data Reset Decision Framework (Corpus Synthesis, April 2026)

The most expensive pre-deployment decision is whether a workflow requires a full data foundation investment (Reset) or can deploy on existing data structures (Proceed/Prepare). Domain count is the primary decision variable.

  • Proceed (1 domain, structured, coherent): Infrastructure modernization only. SlickDeals 360x latency improvement + 7% revenue gain with no semantic integration layer. Timeline: 6–16 weeks. Budget: $70K–$400K.
  • Prepare (2 domains or mixed schema): Targeted entity resolution and schema normalization before deployment. 4–12 weeks of data work. Budget adds $80K–$400K to deployment costs.
  • Reset (3+ domains or DRI <0.50): Semantic integration layer (Ontology or equivalent) required. Nebraska Medicine built a revenue cycle workflow in 10 hours — on top of a 6-month Ontology that cost a full enterprise implementation. Timeline: 9–24 months. Budget: $450K–$2.3M total.
  • Reset is waste when the workflow can be scoped to one domain and deliver 80%+ of value, or when the Reset timeline exceeds the executive patience horizon. BCG finding: concentrating AI on 1–3 domains produces results; spreading across 100 use cases fails.
  • Reset compounds when multiple future workflows will share the integration layer. Palantir’s 139% net dollar retention (Q4 2025) reflects customers expanding because each new use case after the Ontology is faster and cheaper than the last.
  • Four-variable diagnostic: domain count (primary), data structure type (secondary), decision-point density (tertiary), data maturity score (moderating).
  • Four workflow archetypes succeed repeatedly: high-volume structured decisions (fraud/scoring), knowledge retrieval bottlenecks (legal/support), document synthesis before human decision (revenue cycle/contracts), monitoring at inhuman scale (supply chain/fraud volume).
  • 26% of organizations can use unstructured data in a way that delivers business value (IBM CDO Survey, 2025) — the unstructured-data capability gap is the most acute data-readiness blocker for high-decision-point workflows.

Source: research/07-adoption-challenges/ai-data-reset-decision-framework.md

Data Cleaning Real Timelines & Case Studies (Gartner, Qlik, Caylent, Anaconda, 2025)

  • Gartner (Feb 26, 2025, n=248 data management leaders): 60% of AI projects unsupported by AI-ready data will be abandoned through 2026; 63% of orgs lack or are unsure of AI-ready data management practices.
  • Qlik / Wakefield Research (Feb 2025, n=500 U.S. data pros at $500M+ firms): 81% report significant data quality problems; 85% say leadership isn’t addressing; 96% expect widespread crises; 47% worry their company overinvested in AI.
  • Caylent / Teamfront / Arborgold (2025): 4 SQL clusters, 2,500+ stored procedures compressed from 40-week manual estimate to 10 weeks — 70% AI-automated, 20% AI-assisted, 10% hand-coded.
  • Anaconda State of Data Science (recurring, 2024 wave): data scientists spend 39–45% of time on data preparation (not the 80% figure commonly cited, which includes collection/labeling).
  • Enterprise migrations typically 6–18 months end-to-end; silver-tier cleaning (entity resolution, business rules) is the binding constraint, not raw ingestion.
  • Dominant mid-market execution model: SI + AI tooling, 6–12 months, $300K–$1.5M for a single-domain, single-use-case initiative.

Source: research/07-adoption-challenges/data-cleaning-real-timelines-case-studies.md

Legacy Data Remediation TCO by Industry (April 2026)

  • Data prep = 40–60% of AI project cost, 60–80% of time (O’Reilly, 2025)
  • 96% of companies begin AI projects without sufficient quality training data; unplanned remediation $10K–$90K per project (Optimus AI, 2025)
  • Healthcare: Epic implementations $1M–$500M; 42.3% acute-care EHR market share
  • Financial services: 0.65–0.74 DRI score range — below the autonomous-AI threshold; multi-core-banking sprawl is the primary coherence problem
  • Manufacturing: ERP/MES integration gate; 25% of ERP implementations classified catastrophic (Gartner 2025)
  • Integration platforms: Fivetran $500/mo starter with connector-level pricing escalation; dbt Cloud $100/dev/mo; MuleSoft/Informatica $100K–$2M ARR at enterprise scope
  • 83% of data migration projects fail or overrun (Gartner)

Source: research/07-adoption-challenges/legacy-data-remediation-tco-by-industry.md

AI Failure Pattern Library — Data Mirage Archetype (Mar 2026)

  • Pattern 2 (“Data Mirage”) from the six root-cause failure archetypes: AI pilots succeed on clean, curated sample data; production deployment exposes the real data landscape. 60–70% of project time shifts to data preparation. The business case collapses.
  • Data quality issues present in 71% of AI project failures; 38% of formally abandoned projects cite “insurmountable data quality” as primary reason (Pertama Partners, n=2,400+, 2025).
  • Only 7% of enterprises say their data is completely ready for AI (Cloudera/HBR, Mar 2026).
  • Organizations conducting formal data readiness assessments before AI project approval achieve 47% success rate vs. 14% without — a 2.6x improvement.

Source: research/07-adoption-challenges/ai-failure-pattern-library.md

Enterprise RAG Data Quality Dependency (2025–2026)

RAG deployments fail most often on data quality, not retrieval architecture. Legal AI tools built on RAG hallucinate 17–33% of the time even with retrieval grounding (Stanford/Magesh, JELS, 2025). 70%+ of 2025 RAG deployments launched without systematic evaluation. The highest-ROI RAG use cases (customer support: +14% productivity, n=5,179; internal knowledge search) require well-maintained, permissioned knowledge bases — the same data-readiness prerequisites that gate all enterprise AI. Organizations that tried context-stuffing without retrieval added vector layers within 12 months in 71% of cases (Gartner Q4 2025, n=800).

Source: research/01-ai-native-landscape/enterprise-rag-guide.md

Internal knowledge management is the highest-volume, lowest-risk RAG use case. A 300-person company loses $2.2M/year in search friction (McKinsey Global Institute; Slite, n=100+, 2025); AI-powered knowledge retrieval delivers 116–353% ROI (Forrester TEI of Glean and Microsoft 365 Copilot, 2025). The same permissioned-knowledge-base data-readiness requirements that gate customer-support RAG apply here.

Source: research/01-ai-native-landscape/ai-internal-knowledge-management.md


Data Readiness Investment ROI (Gartner, Precisely, Deloitte, 2025–2026)

  • Gartner (Feb 2025): 60% of AI projects will be abandoned through 2026 at organizations without AI-ready data
  • Precisely / Drexel LeBow (Jan 2026, n=500+ data leaders): 88% confident in readiness; 43% simultaneously name data readiness as #1 AI barrier — the confidence-readiness gap
  • Data-strategy cohort: 32% expect positive ROI in 6–11 months; 71% report high data trust vs. 50% without a formal program
  • Deloitte 2025: “future-built” companies reach production with 62% of AI initiatives vs. 12% for laggards; time-to-impact 9–12 months vs. 12–18 months
  • 1-10-100 rule holds for AI-era data: $1 to prevent at entry, $10 to remediate downstream, $100 when the error reaches a decision or customer
  • MIT Sloan / Redman: 47% of newly-created data records contain at least one critical error affecting downstream processes

Source: research/07-adoption-challenges/data-readiness-investment-roi.md

Medallion Architecture: Bronze/Silver/Gold Timelines (April 2026)

  • Pattern is authoritative but not mandatory — Microsoft Learn (March 2026): “a recommended best practice but not a requirement”
  • Bronze (raw, no validation) up across priority sources: 4–8 weeks
  • Silver (dedupe, entity resolution, schema enforcement, joins, business rules): 3–6 months for first domain in mid-market — the layer where 60% of AI projects die (Gartner Feb 2025)
  • Gold (dimensional model, aggregates, business-ready): 1–3 months after silver is stable
  • Full coverage across 4+ business domains: 18–24 months; most mid-market companies finish only 2 domains well
  • Vendor “48-hour medallion” claims (Nexla) cover the pipeline, not entity resolution or business-rule definition — those require CFO/sales ops/operations VP time
  • Storage 3x multiplier; 3 layers multiplies ETL jobs, monitoring, failure points
  • Survey signal: 61% cite data quality as top challenge; 75% don’t trust their data for decisions (Gartner n=248, 2024)

Source: research/07-adoption-challenges/bronze-silver-gold-data-tiering-timelines.md

Snowflake 2026: The 20%/32% Floor (n=2,050)

Snowflake’s enterprise survey (n=2,050, Aug–Sep 2025, Omdia/Informa TechTarget independent fieldwork) found across 6 industries that only 20% of unstructured data and 32% of structured data is considered AI-ready. This is the most specific structured/unstructured split available in a large-n enterprise survey and corroborates the Cloudera/HBR 7% “completely ready” finding from a different angle.

The industry ROI spread (38% manufacturing vs. 69% advertising & media) traces directly to this data readiness gap: ad/media companies cite data quality as a challenge at only 27% (vs. 40% average), explaining most of the 31-point ROI differential. Data readiness is the leading structural predictor of AI ROI across industries — not model quality, not tool selection.

65% of respondents cite data silo breakdown as the hardest implementation challenge. This aligns with the bronze/silver/gold timeline research: entity resolution across organizational silos is the step that consumes the most time and cannot be automated.

Source: research/05-analyst-firms/snowflake-roi-gen-ai-agents-2026.md

Data Architecture Options for Mid-Market AI (April 2026)

  • Five architectures dominate vendor pitches: warehouse, lake, lakehouse, mesh, fabric. Only warehouse and lakehouse are live decisions for a 200–2,000 person company.
  • Mesh is an operating model (Dehghani 2020 four principles: domain ownership, data-as-product, self-serve platform, federated governance) — requires platform team of 3+ and domain data product owners most mid-market orgs lack.
  • Fabric (Informatica, Denodo, IBM) is primarily sold to Fortune 500 with 50+ heterogeneous source systems; premature for <15-system mid-market environments.
  • 2025 pricing: Snowflake $2.00–$3.10/credit; BigQuery ~$6.25/TiB scanned; Redshift ra3.4xlarge ~$3.26/hr; Databricks DBU-based. Mid-market annual spend runs $30K–$300K depending on workload discipline.
  • Default pick follows the cloud you already run: Redshift (AWS), BigQuery (GCP), Fabric (Azure/M365), Snowflake (cross-cloud), Databricks (heavy ML/Spark).
  • Cost failure modes: no auto-suspend, unbounded scans, serverless defaults on provisioned-appropriate workloads. Flexera 2025: Snowflake customer costs up ~40% driven by workload growth, not rate changes.

Source: research/07-adoption-challenges/data-architecture-options-mid-market-ai.md

AI-Assisted Data Cleaning Tools 2025-2026

  • dbt Copilot GA in dbt Cloud Team ($100/dev/mo) and above; accelerates SQL model + test authoring but does not decide grain, business logic, or source semantics.
  • Monte Carlo added “Fix with AI” and Agent Observability (Sept 2025); typical mid-market ACV $25K–$50K for 30–100 tables. Soda AI converts NL prompts to DQ checks; Soda Team is $8/dataset/mo — the only DQ tool with published per-dataset pricing.
  • Enterprise catalogs are enterprise-priced: Alation ~$198K/year baseline (AWS Marketplace, July 2024); Collibra $170K (12mo) to $510K (36mo); Informatica IPU deals six-figure minimums. Buying before a CDO/governance function exists typically strands the tool.
  • What stays manual across all tools: entity resolution arbitration, business rule definition, source-system tribal knowledge, triage of false-positive anomalies, governance decisions. IBM and Dataversity note LLM-based DQ assistants hallucinate when source metadata is sparse — i.e. in most mid-market environments.
  • Realistic mid-market stack: dbt Cloud Team + Soda Team + Monte Carlo = $30K–$70K/year in tools, plus 0.5–1.0 FTE per tool to operationalize.

Source: research/07-adoption-challenges/ai-data-cleaning-tools-2025-2026.md

Dun & Bradstreet Global AI Momentum Survey — 5% Data Ready (n=10,000, May 2026)

The largest enterprise AI panel published to date (10,000 businesses, 32 countries, Q1–Q2 2026 fieldwork) finds the same structural signal as Cloudera/HBR from a different angle: only 5% of organizations say their data is fully ready for AI, while 97% report active AI initiatives. The 97→10% adoption-to-ROI collapse — only 10% report strong ROI — is the survey’s headline finding.

Primary data-readiness obstacles cited across 10,000 businesses:

  • Limited data access: 50%
  • Privacy and compliance risks: 44%
  • Data quality and integrity: 40%
  • Lack of system integration: 38%
  • Shortage of skilled AI professionals: 37%
  • High confidence in AI risk management: only 10%

The survey introduces a specific agentic-era failure mode: AI agents crossing system boundaries need verified, continuously refreshed business entity data (standardized company/contact/product identity) to avoid confident errors at machine speed. Without it, agents that simultaneously touch CRM, ERP, procurement, and finance systems produce conflicting outputs based on inconsistent entity resolution. Corroborated by Palantir’s Ontology-first model and MIT CISR’s digital-colleagues architectural requirements.

Source note: D&B sells the D-U-N-S® Number and D&B Commercial Graph™ — the survey’s central finding directly supports their product. Read with that bias. The 5% figure is directionally corroborated by Cloudera/HBR (7%, March 2026), Gartner 4x foundations differential (April 2026), and Snowflake 20%/32% structure/unstructured split.

Source: research/01-ai-native-landscape/dnb-global-ai-momentum-survey-2026.md


Cloudera / HBR Analytic Services — AI Data Readiness Primary Survey (Mar 2026)

Two companion surveys published March 5, 2026 quantify the data readiness gap with primary-survey evidence.

HBR Analytic Services survey (n=230 HBR audience members involved in AI data decisions, October 2025 fieldwork, Cloudera-sponsored):

  • Only 7% of enterprises say their data is completely ready for AI adoption
  • 27% say their data is not very or not at all ready; 66% fall in between
  • 73% believe their organization should prioritize data quality more than it currently does
  • 73% found processing and preparing data for AI to be challenging
  • Only 23% have an established AI data strategy; 53% are developing one; 24% have neither
  • Top barrier: siloed data and integration difficulty (56%)
  • 47% believe agentic AI will resolve their data quality issues — a misdiagnosis; agentic systems amplify whatever data quality exists

Cloudera Data Readiness Index (n=~1,300 global IT leaders):

  • 96% claim AI integration into core business processes
  • ~80% admit AI and data initiatives are constrained by limited data access — the AI-readiness illusion
  • Only 18% describe their data as fully governed
  • Sector gap: Telecom 54% full data visibility vs. Financial Services and Public Sector at 30–31%

The 7%/80%/18% figures corroborate independent surveys: Gartner (60% abandon AI projects without AI-ready data), Qlik (81% significant quality problems), Deloitte (data-ready cohort 62% production rate vs. 12% laggards). The magnitude is consistent even though the Cloudera commercial interest is noted.

Source: research/07-adoption-challenges/cloudera-hbr-ai-data-readiness-2026.md

Gartner AI Research 2026 — Data Readiness Signals

  • 57% of companies acknowledge their data is only enterprise-grade at best (Gartner Hype Cycle 2025)
  • Finance AI adoption essentially flat (58% → 59% YoY) with primary barriers being data literacy/technical skills and inadequate data quality (Gartner CFO Survey, n=183, Nov 2025)
  • 91% of finance functions deploying AI experienced only low or moderate initial impact; organizations that persist are 2–3x more likely to achieve moderate-to-high impact
  • Gartner positions “AI-ready data” at the Peak of Inflated Expectations — correction ahead

Source: research/04-consulting-firms/gartner-ai-research-2026.md

IDC FutureScape 2026 (Oct 2025)

  • Companies that do not prioritize high-quality, AI-ready data will experience a 15% productivity loss when scaling GenAI and agentic solutions (IDC FutureScape prediction, by 2027)
  • Only 1% of organizations have achieved an optimized, AI-fueled enterprise state; over 50% remain in early transformation stages
  • The spending-maturity gap: AI IT spending grows at 31.9% CAGR toward $1.3T by 2029, but organizational readiness lags far behind — data readiness is the primary bottleneck IDC identifies

Source: research/04-consulting-firms/idc-ai-research-2026.md

AWS re:Invent 2024–2025 Customer Sessions

  • Confluent’s Andrew Sellers: “90% of any new application is data modeling” — stated on-record at re:Invent. The model is the last 10%; the first 90% is getting data into a shape the application can use.
  • IBM’s Armand Ruiz: 90% of 1,000+ GenAI pilots Ruiz reviewed were RAG use cases — the bottleneck in every case was data access and retrieval quality, not model capability or novel training requirements.
  • Capital One’s Prem Natarajan: “Your data advantage is your AI advantage” — stated explicitly at re:Invent. Capital One’s 10,000-agent production deployment is built on proprietary data assets the model cannot replicate.
  • MongoDB panel: data siloing identified as the single most common production barrier. Organizations that could not federate data across silos could not move pilots to production regardless of model choice.
  • These four on-record signals from named practitioners corroborate the Gartner, Qlik, and Precisely survey findings above: data readiness blocks production, and the bottleneck is structural, not technical.

Source: research/13-multimodal-sources/aws-reinvent/aws-reinvent-2024-2025-enterprise-ai-sessions.md

Databricks Data+AI Summit 2025

  • Jamie Dimon (JPMorgan Chase CEO): “The hardest part is the data… It isn’t the AI. Getting the data in the form that it’s usable is a hard thing to crack.” JPMorgan spends ~$2B/year on AI — the world’s largest bank identifies data preparation, not model capability, as the binding constraint.
  • U.S. Navy: reviewed $40B in transactions, freed $1.1B for other priorities, saved 218K work hours — outcome depended on data platform consolidation before AI deployment.
  • Insulet: 97% lower TCO and 12x faster processing after replacing legacy ETL — data infrastructure modernization was the prerequisite for AI gains, not the AI models themselves.
  • Databricks platform metrics: 97% Unity Catalog governance adoption, 530% AI BI user growth YoY — the data governance tooling exists; the constraint is organizational readiness to use it.
  • Corroborates the Gartner 60%-abandonment finding and Capital One “data advantage = AI advantage” thesis from re:Invent.

Source: research/13-multimodal-sources/databricks-data-ai-summit/databricks-data-ai-summit-2025-enterprise-ai-sessions.md

Snowflake Summit / BUILD 2025–2026

  • Discover Financial: Estimates achieving equivalent data quality coverage with traditional platforms would require 25 years of full-time effort — strongest anchor for the automated data quality investment case (claim relayed through Anomalo partnership, MEDIUM credibility).
  • NZHL (New Zealand Home Loans): “Biggest impact is trust. Trust in their data.” — data trust as the prerequisite for AI readiness and agentic AI adoption.
  • Northwest London Healthcare: Data ingestion compressed from weeks to days after Snowflake migration; cross-org data sharing (weather → patient demand prediction) as next-tier AI use case.
  • University of Auckland: 9-person developer team shifted from data maintenance to AI model building after Fivetran + Snowflake migration.
  • Pattern across all Snowflake customer stories: data platform quality — not model selection — is the binding constraint on AI value.

Source: research/13-multimodal-sources/snowflake-summit/snowflake-summit-build-2025-2026-enterprise-ai-sessions.md

Gartner I&O AI Use Case Survey (Apr 2026, n=782)

  • 38% of I&O leaders who faced AI setbacks cite poor data quality or limited data availability as a direct cause of failure — tied with skills gaps as the top organizational blocker.
  • Data readiness appears alongside skills gaps as the persistent structural barrier: no model upgrade fixes an underlying data quality problem.
  • Reinforces Gartner’s own Feb 2025 prediction (n=248) that 60% of AI projects unsupported by AI-ready data will be abandoned through 2026.

Source: research/05-analyst-firms/gartner-ai-io-stalling-2026.md

IBM IBV + Adobe “Win the Moment” Customer Intent Study (Apr 15, 2026, n=1,000)

  • Only 34% of customer data organizations collect today is used for customer experience decisions — the remaining 66% sits fragmented across platforms and silos, never informing an offer, intervention, or experience.
  • 75% of executives say most organizations are too slow to respond to changing customer expectations; the diagnosed root cause is data utilization, not data volume.
  • Top operational friction points: cross-channel identity resolution (54%), lifecycle adaptability (48%), dynamic content generation at scale (45%) — identity resolution is the single largest pattern-breaker that kills real-time orchestration pilots.
  • Detection-to-action window (hours from intent signal to action) gap between top 20% and industry average ranges from 2.2x (industrial products) to 23x (media, entertainment, telecoms) — a data-utilization metric, not a technology-speed metric.
  • Executives expect AI-enabled signal share to nearly double from 35% to 63% by 2028, doubling the volume of signals and decisions the data layer must reconcile — a multiplier on the existing utilization gap.
  • Vendor caveat: IBM Consulting + Adobe co-publishers with commercial interest in customer data platform and agentic orchestration engagements.

Source: research/04-consulting-firms/ibm-ibv-customer-intent-2026.md

Re-Ranking Architectures for Enterprise RAG: Why Embedding Similarity Fails (May 2026)

The retrieval layer in enterprise RAG pipelines has a precision problem that embedding similarity alone cannot solve. For agentic AI systems relying on live organizational data, retrieval quality is a direct determinant of output quality and agent reliability.

  • Why embedding similarity fails: cosine similarity between a query and a chunk measures approximate semantic proximity, not relevance. High similarity ≠ correct answer. Low similarity can mean highly relevant but phrased differently. At enterprise scale, the difference materializes as agents returning wrong document versions, stale policy text, or topically adjacent but factually incorrect content.
  • Cross-encoder reranking (most precise): re-encodes query + candidate chunk together, scoring true relevance. BEIR benchmark nDCG@10 ~0.52. Adds 200–400ms latency per batch. Right for high-stakes retrieval (compliance, legal, financial underwriting).
  • ColBERT late interaction (speed-precision balance): token-level embeddings compared across query and document. BEIR nDCG@10 ~0.48. Adds 40–80ms. Right for production-scale deployments where cross-encoder latency is prohibitive.
  • BM25 hybrid (keyword + semantic): sparse retrieval (exact keyword match) combined with dense retrieval (semantic similarity). BEIR nDCG@10 ~0.44 but lower latency than both. Particularly effective for enterprise terminology, product codes, and domain-specific jargon that embeddings blur.
  • When to skip reranking: document populations under ~500 chunks with consistent formatting and terminology — the additional latency exceeds the precision gain.
  • Cisco VAST InsightEngine: example of enterprise production RAG with cross-encoder reranking at network-operations scale; retrieval latency maintained under 200ms at 50TB+ document corpus.

Source: research/19-agent-frameworks/retrieval-reranking-enterprise-rag.md · May 2026 · HIGH · TIER 1

IBM IBV “The Tech Debt Reckoning” (Nov 2025, n=1,300)

  • Tech debt — aging infrastructure, fragmented data architectures, brittle code — is “the less visible but often larger cost” of AI deployment. Models, compute, and talent are the visible portion; preparing the existing tech estate is the larger absorber of budget.
  • Enterprises that fully price tech-debt remediation into AI business cases project +29% higher ROI than those that don’t. The mechanism is planning hygiene, not technology: pricing the debt surfaces the real cost so the business case is honest from day one.
  • 18–29% of total AI implementation cost through 2027 is expected to be absorbed by tech-debt remediation; 15–22% schedule extension (30-month programs become 36-month programs). Debt-blind business cases that omit these line items produce a -14% ROI post-mortem versus the +39% projected on paper.
  • Named the core data-readiness mechanism: “when a majority of an organization’s high-value data is locked behind legacy systems, contaminated by inconsistencies, or sitting on an aging, vulnerable infrastructure, the algorithmic horsepower of the newest model is irrelevant. The machine starves.”
  • Companion IBM IBV 2025 Chief Data Officer survey findings: 72% of CDOs agree realizing GenAI value depends on leveraging proprietary data, but only 26% say their organization can use unstructured data in a way that delivers business value — the unstructured-data capability gap is the most acute data-readiness blocker in 2025–2026.
  • Prescription: every AI initiative gets a line-item tech-debt estimate in its business case. Debt-adjusted ROI is the ranking tool. Concentrate initiatives in a few domains so one remediation compounds across related initiatives — 80% of executives agree that fixing debt in one AI initiative improves the ROI of related future initiatives.
  • Vendor caveat: IBM Consulting has direct commercial interest in application-modernization and tech-debt-remediation engagements; +29% figure is self-reported business-case projection, not measured post-deployment ROI.

Source: research/04-consulting-firms/ibm-ibv-tech-debt-reckoning-2026.md

Data Structure Readiness as a Workflow-Level Criterion (April 2026)

At the workflow level — as opposed to the enterprise level — data readiness reduces to one question: is the input data that drives this workflow structured, consistent, and accessible without manual extraction? This is Criterion 1 in the six-criterion workflow AI-readiness scorecard.

The five-point scale for data structure readiness:

  • Score 1: Data lives in 3+ systems; requires manual export or cleaning; formats vary by team or region. Production AI deployment on this data will reproduce the Data Mirage failure pattern (Pertama Partners, n=2,400+, 2025-2026): pilot succeeds on curated sample data; production fails 5.2 months in when the real data landscape is encountered.
  • Score 2: Data is accessible from 2 systems; some inconsistency in field definitions or formats.
  • Score 3: Data lives in one system; consistent schema; accessible via API or query without manual steps.
  • Score 4: Single system of record; documented schema; complete for the last 12+ months; quality has been audited.
  • Score 5: Single source; consistent; audited; real-time or near-real-time accessibility; lineage documented.

Corpus evidence for the workflow-data link: SlickDeals (AWS re:Invent 2025) — real-time structured data (community votes, deal metadata, user engagement) enabled 360x latency compression and 7% revenue gain. Nebraska Medicine (Palantir AIPCon 8, 2025) — revenue cycle automation built in 10 hours on a 6-month Ontology that structured the data first. JPMorgan (Jamie Dimon, Databricks Summit 2025): “The hardest part is the data… Getting the data in the form that it’s usable is a hard thing to crack.” JPMorgan spends ~$2B/year on AI; even at that scale, data structuring is the binding constraint.

Why Criterion 1 is the primary workflow gate: Only 7% of enterprises say their data is completely AI-ready (Cloudera/HBR, March 2026). Organizations that conduct formal data readiness assessments before AI project approval achieve 47% success rate vs. 14% without — a 2.6x improvement (Pertama Partners, 2025-2026). The assessment does not need to be enterprise-wide; it needs to cover the specific workflow’s input data, production sources, and access path.

The practical test: Before approving a workflow for AI deployment, answer: Can the AI application access the data it needs, in the format it needs, from the systems where it lives, without manual intervention? A “no” is not a project-killing condition — it is a scope-definition condition. The data remediation work must be in scope and budgeted before AI deployment begins.

Source: research/07-adoption-challenges/workflow-level-ai-readiness-checklist.md

Yale CELI Agentic AI Enterprise Series (Apr–May 2026, 12 sectors)

  • 80% of enterprises cite data limitations as the primary agentic AI scaling obstacle — the most direct large-sample corroboration of the Cloudera/HBR 7% “completely ready” finding, from an independent institutional source.
  • Only 7% describe their data as “completely ready” for AI deployment (Cloudera/HBR, cited by Yale CELI). 63% either lack AI-suitable data management or are unsure.
  • The K-shaped divergence mechanism: data-mature firms adopt agentic AI incrementally at low marginal cost, compounding their advantage. Data-immature firms face step-change infrastructure investment before deployment is viable — measured in years, not quarters.
  • Yale CELI identifies three data strategies in the field: Deepening the Substrate (migrate unstructured operational data to cloud; capital-intensive, durable); Wrapping with Protocol Layers (translation layers on legacy systems; lower cost, faster); Vendor-Supplied Connectivity (Anthropic Cowork, Microsoft Copilot Studio; fastest deployment, highest vendor concentration risk).
  • Banking-sector parallel: 77% of banking leaders cite data privacy as their top agentic scaling barrier; 62% of hospitals report EHR data silos blocking clinical AI deployment.

Source: research/01-ai-native-landscape/yale-celi-agentic-ai-enterprise-series-2026.md · Yale CELI (Sonnenfeld et al.) · 12 sectors + public sector · Apr–May 2026 · HIGH (institutional) / MEDIUM (case data) · TIER 1


Rewired (Lamarre et al., 2024) — Digital and AI Backbone Framing

Rewired dedicates its longest section (Section Four + Five, 13 chapters) to what it calls the “technology for speed and distributed innovation” and “embedding data everywhere.” Two chapters are directly applicable to data readiness decisions:

Ch. 25 — “No Data Architecture, No AI Advantage” (p.391): The book’s most actionable data claim. The argument: without a governed, reusable data architecture, AI investments produce “recurring bespoke-integration costs and bounded outcomes.” Each new use case requires its own integration work. The economic mechanism of compounding returns — each new AI use case cheaper and faster than the last — only activates when a shared data foundation exists. This is the corpus’s primary evidence for why data architecture investment is a strategic decision, not a technical cleanup project.

Ch. 26 — “Data Products as Reusable Building Blocks” (p.401): The operational expression of Ch. 25. A data product (customer 360, supplier-risk profile, equipment condition score) is built once and reused across multiple AI workloads. This is what drives Palantir’s 139% net dollar retention (Q4 2025, SEC-filed): after the foundational Ontology, each new AI use case draws on existing data products rather than rebuilding from raw sources.

Mid-market adaptation: Rewired’s Ch. 21–22 advocate building an AI platform — appropriate for Fortune 500. For the mid-market, the equivalent is purchasing a platform (Snowflake + Cortex, Databricks, or Palantir Foundry) and investing the saved build cost in the data product layer on top. The economic logic is the same; the make-vs-buy decision differs by scale. PwC AI Performance Study 2026 (n=1,217) finds median mid-market AI buy achieves 1.4–2.2x ROI vs. 0.7–1.3x for build over three years.

The Rewired data backbone in the context of this wiki: The decision frameworks above (Data Reset, Medallion Architecture, DRI scoring) operationalize what Ch. 25 describes at a strategic level. The corpus’s specific guidance (domain count as the primary variable, 3+ domains requiring full Reset, the Proceed/Prepare/Reset classification) can be handed to a CIO as the implementation layer beneath Rewired’s principle.

Source: research/04-consulting-firms/mckinsey-rewired-2nd-edition-synthesis.md


Practitioner voices (pillar 13)

  • Ronald den Elzen, Chief Technology and Digital Officer, Heineken Group (Me, Myself, and AI / MIT SMR + BCG, January 2025):

    “We really need to work on data quality. Our managing directors in markets start understanding that they need to invest in the quality of data, that they start to understand that unique data, so first-party data, is relevant to compete in the marketplace.”

    Context: Heineken’s CTDO describes a transition from treating data quality as a technical concern to treating it as a boardroom investment priority. The explicit link to first-party data and competitive differentiation maps directly to the Criterion 1 framework above — without clean, proprietary data, AI personalization and demand forecasting tools underperform published benchmarks. Credibility: HIGH (MIT SMR + BCG editorial, named exec, unscripted).

    Source: research/13-multimodal-sources/me-myself-and-ai/2025-01-07-how-a-160-year-old-startup-uses-ai-the-heineken-companys-ron.md

  • Reynolds Shin, Co-founder, Databricks (VentureBeat / Beyond the Pilot, September 2025):

    “For it to improve, you kind of need to know how well it’s doing. And for any sort of improvement you deploy, you also need to know, hey, is it actually doing better or is it doing worse? And then the other thing which is very important in the traditional applied machine learning is the data, which is do you have the right data? Are you getting the monitoring of your data into the right place? Are you getting the right input? Are they clean? So this is a very important data engineering problem.”

    Context: The Databricks co-founder frames data quality and monitoring as the non-negotiable foundation for any AI improvement loop. The framing is directly applicable to the data structure scorecard above: without Score 4+ data (single system, audited, complete), the feedback loop that lets AI systems improve over time cannot close. Credibility: HIGH (VentureBeat editorial, named exec and co-founder, unscripted).

    Source: research/13-multimodal-sources/beyond-the-pilot/2025-09-10-venturebeat-scaling-up-with-databricks-achieving-the-potenti.md

  • Marc Benioff, CEO, Salesforce (Salesforce Dreamforce Keynote, 2025):

    “You have got to get your data right. You have got to get the priorities right. You have to get the governance right.”

    Context: Benioff’s Dreamforce keynote framed data readiness as the prerequisite to agentic scale — not a technical detail but a CEO-level investment sequencing decision. The quote appears alongside deployment evidence from PepsiCo, Engine, and Grupo Falabella, all of which required data infrastructure work before agentic deployments produced measurable results. Credibility: MEDIUM (vendor keynote, named exec, directionally consistent with independent research).

    Source: research/13-multimodal-sources/salesforce-dreamforce/salesforce-dreamforce-2025-enterprise-ai-sessions.md


Mid-Market Data Governance: Why 93% of Companies Fail AI Before Writing a Line of Code

RSM’s Middle Market AI Survey (n=966) finds 41% of mid-market companies experiencing AI implementation issues cite data quality as their top barrier — ahead of security (39%) and talent gaps (35%). Yet 91% have adopted generative AI. Cloudera/HBR Analytic Services (n=230+, Oct 2025) finds only 7% of enterprises describe their data as completely AI-ready; 27% say theirs is “not very or not at all ready.” Roughly 55% of enterprise data sits “dark” — uncategorized, untagged, and invisible to AI systems (DataStackHub/Veritas, 2025).

Three mid-market-specific failure modes:

  1. Demo data vs. production data: Pilots succeed on curated sample data; production fails on the company’s actual fragmented records.
  2. Siloed data without integration: 56% of data decision-makers cite siloed data and integration difficulties as the top obstacle (Cloudera/HBR).
  3. Skipped data governance: Organizations with mature data governance reduce AI implementation costs 20–35% and accelerate time-to-value 40–60% (Atlan, 2025). The prerequisite is not expensive. Skipping it is.

The practical pre-deployment test: Can the AI application access the data it needs, in the format it needs, from the systems where it lives, without manual intervention? A “no” is a scope-definition condition — data remediation must be budgeted before AI deployment begins.

Source: research/07-adoption-challenges/mid-market-data-governance-prerequisite.md

Nasuni “State of Enterprise File Data 2026” — The Unstructured Data Agent Failure (n=1,000, Mar 2026)

The largest-n survey specifically measuring unstructured data as the AI agent failure mechanism:

  • 97% of 1,000+ employee organizations are deploying or piloting AI agents; only 18% have reached enterprise-wide scale
  • 57% of AI projects are not meeting their stated objectives — root cause traced to unstructured data, not model quality or governance frameworks
  • >90% of all organizational data is unstructured (documents, emails, images, recordings, design files, collaboration assets)
  • 94% of enterprises struggle to manage unstructured data; only 16% currently treat it as a core IT investment priority
  • 90% report obstacles to AI scaling: data security (43%), integration roadblocks (36%), lack of data trust (33%)
  • 46% say AI initiatives have already exposed data quality and governance gaps they did not previously know existed
  • Only 21% operate centrally managed file environments; average org uses 4 separate storage/backup/DR systems

The unstructured data angle is the gap in prior corpus coverage. The Gartner 4x differential, Cloudera/HBR 7% AI-ready finding, and Snowflake 20%/32% structured/unstructured floor all address the general data readiness deficit. Nasuni’s finding specifies the mechanism: AI agents retrieve from ungoverned file systems containing superseded versions, unclassified content, and inconsistently permissioned records — and produce unreliable outputs as a result.

Source: research/07-adoption-challenges/nasuni-unstructured-data-ai-failure-2026.md · Nasuni/Sapio Research · n=1,000 · March 2026 fieldwork · MEDIUM. TIER 1.

See also

PwC’s May 2026 survey of 767 U.S. operations and supply chain leaders ($100M+ revenue organizations) puts the most specific numbers yet on data quality as the bottleneck that precedes all others:

  • 87% say poor data quality has hampered their progress in achieving value for digital initiatives
  • 30% report significant improvement in data quality and reliability
  • 51% say their companies establish a clean, structured data foundation before scaling digital initiatives

The 51%/87% split is the diagnostic: roughly half of organizations are launching AI without a clean data foundation, and nearly nine in ten are experiencing value degradation as a result. The operations and supply chain context makes this particularly acute — agent systems in logistics, procurement, and inventory management operate at high transaction volumes where data errors compound.

The corroboration is consistent. Forrester’s 2026 GenAI Enterprise Value study identifies data readiness as the primary ROI barrier across enterprise functions. Deloitte’s survey (n=3,235) finds that the 34% of companies in “deep transformation” address data quality before scaling AI, not concurrently.

Source: research/04-consulting-firms/pwc-digital-trends-operations-2026.md

Gartner: The 4x Investment Differential (April 2026)

Gartner’s survey of 353 D&A and AI leaders (Nov–Dec 2025) quantifies the investment gap between AI winners and laggards as a data-foundation spending ratio:

  • 4x higher investment in data quality, governance, AI-ready talent, and change management — as a percentage of revenue — among organizations reporting successful AI outcomes vs. those reporting poor outcomes
  • Only 39% of technology leaders are confident their current AI investments will have positive financial impact — a confidence gap that maps directly onto the 61% of organizations not making the 4x foundation investment
  • Organizations at highest AI-ready D&A maturity report up to 65% better business outcomes (revenue growth + cost optimization) vs. lowest-maturity peers
  • Only 23% of IT leaders are very confident in their organization’s ability to manage GenAI security and governance (separate Gartner survey, n=360 IT leaders, Q2 2025)
  • AI deployment doubled from 40% to 80% of organizations between 2024 and 2025 — adoption is not the constraint; foundation quality is

Gartner’s three-pillar sequencing: Ambition → Foundations → People. Organizations that skip Foundations (data quality, governance) in pursuit of model experimentation land in the 61% uncertain cohort. Independent corroboration: Davenport/Return on AI Institute (n=1,006, March 2026) finds data quality produces a 2x ROI multiplier; Cloudera/HBR (n=230, March 2026) finds 7% of enterprises fully data-ready.

Source: research/05-analyst-firms/gartner-data-foundations-ai-success-2026.md


Salesforce CIO Survey 2025 — Governance Confidence Gap at the CIO Level

n=200 global CIOs, 24 countries, NewtonX double-blind methodology, November 2025.

  • Only 23% of CIOs completely confident they are investing in AI with built-in data governance — this is the CIO’s self-assessment, making it more conservative than external analyst benchmarks
  • Only 14% of IT budget dedicated to data security — the allocation gap confirms governance underinvestment is structural, not a perception error
  • Only 35% of CIOs working more closely with their CDO on agentic AI despite naming data security and trusted data as their top two AI fears — the governance gap is a cross-functional collaboration failure as much as a technology gap
  • Full AI implementation has increased to 42% of enterprises (up from 11%) while data governance confidence remains at 23% — the deployment-governance gap is widening, not closing, as scale increases

Source: research/01-ai-native-landscape/salesforce-csuite-agentic-ai-2026.md · November 2025 · MEDIUM-HIGH · TIER 1


Infor Enterprise AI Adoption Impact Index — Data Readiness as a Scaling Barrier (April 2026)

n=1,000 decision-makers (US/UK/Germany/France; vendor-commissioned). Adds practitioner-level self-assessment to the academic and analyst benchmarks above:

  • 27% are unsure or disagree that their data is mature enough for reliable AI — a direct self-assessment, not an inferred readiness gap; likely a conservative undercount for organizations that haven’t yet attempted structured AI data preparation
  • 36% cite data security, sovereignty, privacy, or compliance as the primary barrier to scaling AI beyond pilots — the leading barrier, above talent (25%) and unclear ROI (23%)
  • 49% of AI-generated outputs still require manual expert review before acting — not because of skepticism, but because the data pipelines and governance frameworks for autonomous execution are not yet in place
  • 37% cite enhanced data security and sovereignty as the top long-term AI priority — data governance is not being deferred; it is the primary investment target

In the context of the broader data-readiness corpus, the 27% figure aligns directionally with the Gartner 4x differential (organizations investing 4x more in data foundations are 4x more likely to capture AI value) and the Cloudera/HBR 7% fully-ready finding. Vendor context: Infor benefits from evidence that ERP-embedded data governance solves this barrier. Weight accordingly.

Source: research/07-adoption-challenges/infor-enterprise-ai-adoption-impact-index-2026.md · April 2026 · LOW-MEDIUM · TIER 1

AWS / HBR Analytic Services — “Agentic AI: Expectations, Readiness, and Results” (n=623, Jan 2026)

The clearest quantification of the agentic-specific data readiness gap in the current corpus:

  • 13% of organizations believe their data architecture is “well-equipped for agentic AI” — 64% “somewhat equipped,” ~23% not equipped
  • 11% very well-prepared on governance controls; 5% very well-prepared on workforce readiness
  • 95% lack clear success metrics for agentic AI initiatives — the measurement gap makes ROI demonstration impossible
  • 74% say AI use is very important; only 26% report being “very effective” at leveraging AI — a 3:1 intent-execution gap
  • Leading organizations (those that have closed readiness gaps) report 42% innovation gains and 39% customer experience improvements vs. 33–36% for average active users
  • The “somewhat equipped” category is where risk concentrates: adequate for experiments, insufficient for autonomous agentic systems operating at production scale

Cross-reference: Cloudera/HBR (7% fully data-ready, n=230); Gartner 4x investment differential (n=353); Databricks telemetry (19% at production scale); OutSystems (12% centralized governance, n=~1,900). Directional convergence across four methodologically distinct sources.

Source: research/12-agent-workers/aws-hbr-agentic-ai-data-foundations-2026.md · Jan 2026 · MEDIUM-HIGH · TIER 1

Coastal / Oxford Economics — “2026 AI Operations Report” (n=800, May 2026)

Production-stage confirmation that data problems do not resolve post-launch:

  • 70% report data access/quality issues during AI setup; 73% encounter the same problems in production — the gap widens rather than closes after deployment. Data problems migrate from project phase to operating system.
  • 47% of organizations incorrectly believe agentic AI will resolve data quality problems (Cloudera/HBR companion finding) — Coastal’s n=800 production-only sample shows what that misbelief produces: embedded data problems in live systems.
  • Corroboration chain: Gartner 4x investment differential (n=353) → Cloudera/HBR 7% data-ready (n=325) → Coastal 73% production-data-issues (n=800) — three methodologically distinct sources converging on the same structural failure.

Source: research/07-adoption-challenges/coastal-oxford-economics-ai-operations-report-2026.md · May 2026 · MEDIUM · TIER 1

Gartner Supply Chain — Data as the Binding Constraint on AI Scale (Apr–May 2026)

Source: research/05-analyst-firms/gartner-supply-chain-ai-adoption-barriers-2026.md

Supply chain domain confirmation from two Gartner surveys that data infrastructure is the structural bottleneck, not tool availability:

  • 83% of supply chain organizations are constrained by foundational data issues when applying AI — the most common reason AI stays incremental rather than transformational (n=509, Gartner, Jul–Oct 2025)
  • 56% of CSCOs name legacy system integration as a major challenge — the inability to access demand signals, inventory positions, and supplier data programmatically blocks AI from operating on reliable inputs
  • Gartner’s AI leaders in supply chain build a unified data layer above existing systems before activating any AI tools — the specific infrastructure step that separates the 17% from the 83%
  • Supply chain is the domain where data readiness is most directly observable: demand forecasting accuracy can be measured immediately, making the data-quality gap visible in ways that are harder to see in document-processing or decision-support use cases

Source: research/05-analyst-firms/gartner-supply-chain-ai-adoption-barriers-2026.md · Oct–Nov 2025 fieldwork · MEDIUM-HIGH (Gartner independent; TIER 1)

Nutanix Enterprise Cloud Index 2026 — Infrastructure Gap as Data Readiness Bottleneck

The 8th annual ECI (Wakefield Research, n=1,600, 14 countries, Nov 2025) measures infrastructure readiness as a proxy for AI data readiness:

  • 82% of IT leaders say on-premises infrastructure is not fully ready to support AI workloads — consistent with the Cloudera 7% data-fully-ready finding from a different angle
  • 65% run AI workloads via managed service providers rather than own infrastructure — outsourcing the readiness problem without solving it; managed-service delivery reduces data control and customization capacity
  • 80% rank data sovereignty a high-priority infrastructure requirement; 57% require domestically-based infrastructure — a data-residency constraint that must be resolved before any data readiness work can produce AI value at scale
  • The organizational mechanism: 82% cite silos between IT and business units as the primary execution barrier — data readiness cannot be achieved when the teams that own the data (business units) and the teams that need to govern it (IT) cannot align on requirements

Source: research/07-adoption-challenges/nutanix-enterprise-cloud-index-2026.md · Nov 2025 fieldwork · MEDIUM-HIGH (Nutanix vendor; Wakefield independent; TIER 1)

Salesforce State of Data & Analytics, 2nd Edition (n=7,652, Aug 2025)

The largest practitioner survey on data readiness in the 2026 corpus. Key confidence-gap findings:

  • 84% of data/analytics leaders say their data strategies need complete overhauls before AI initiatives can succeed — yet data foundation readiness has improved only 3pp per dimension since 2023
  • 26% of enterprise data is estimated “untrustworthy” by the people managing it; 54% of business leaders lack confidence their required data is even accessible
  • 89% of organizations with AI deployments have experienced inaccurate or misleading outputs — most traceable to upstream data quality failures, not model capability
  • 32% of business leaders have made gut-based decisions because data was unreliable or inaccessible — a direct measure of data readiness failure manifesting as decision failure
  • 70% believe their most valuable insights are trapped in unstructured data — LLMs now unlock this, but governance protocols for unstructured retrieval are absent in most organizations (38% have data lineage tracking)
  • 4x more CIO budget goes to data infrastructure than to AI models, yet readiness improvements remain marginal — confirming the Gartner “4x investment differential” finding from a different angle
  • Annual data volume growth: 30%/yr (up from 23% in 2023) — the readiness problem is not static; it compounds each year without governance investment

Source: research/05-analyst-firms/salesforce-state-of-data-analytics-2nd-edition-2025.md · June–August 2025 fieldwork · MEDIUM (Salesforce vendor; third-party panelists; n=7,652; TIER 2)

Salesforce State of Sales 2026 — CRM Data Quality as the AI Agent Gate (n=4,000+, Feb 2026)

The function-specific confirmation that data readiness failures block AI agent ROI in sales operations:

  • 74% of sales professionals cite data cleansing and integration as a top priority — not as a future goal, but as a present gap blocking AI agent deployment
  • 51% of sales leaders say disconnected systems slow down their AI initiatives — the CRM-to-AI pipeline breaks at the integration seam, not the model capability seam
  • High performers know this: 79% of top-performing sales teams prioritize data hygiene vs. 54% of underperformers — the data quality gap is the leading indicator of the performance gap, not a lagging indicator
  • AI agents in sales run on CRM data; if contact records are incomplete, pipeline stages inconsistently defined, or activity history sparse, agent output degrades proportionally

The 74%/51% pair is the clearest function-specific confirmation of the cross-enterprise data readiness pattern: even in the most AI-forward enterprise function (87% AI adoption), data architecture is the binding constraint on agent ROI.

Source: research/05-analyst-firms/salesforce-state-of-sales-2026.md · Feb 2026 · MEDIUM (Salesforce vendor; n=4,000+; TIER 1)

OneStream / Harris Poll — Bad Data as AI Risk Amplifier (n=352, March 2026)

The clearest evidence that data readiness failures are now a CFO-level liability, not just an IT problem:

  • 47% of C-suite executives (CFOs, CIOs, CTOs, CDOs) made a material business decision on inaccurate, incomplete, or outdated data in the past 12 months — at companies with $50M+ revenue
  • 72% report damages of $500K+; 37% report $1M+ — the financial impact of data quality failure is no longer an abstract risk category
  • Heavy AI users are 4x more likely to have made a bad-data decision — AI amplifies the existing data quality problem at higher velocity; it does not introduce a new one
  • Only 19% pull the majority of AI inputs from a single centralized system; 50% lack a consistent source of truth, data quality rules, or automated reconciliation
  • 79% believe governance supports large-scale AI; 61% second-guess their data at least monthly — the gap between stated confidence and operational behavior is the diagnostic signal
  • Finance-IT ownership paradox: 85% of CIOs say they lead data governance; 78% of CFOs say they do; organizations with complete alignment are 5.5x more likely to trust data completely

Source: research/07-adoption-challenges/onestream-bad-data-ai-risk-enterprise-2026.md · March 2026 · MEDIUM (OneStream vendor; Harris Poll independent; n=352; TIER 1)

GrowthLoop 2026: Marketing SSOT Revenue Premium

GrowthLoop/Ascend2 (n=300+, $100M+ revenue, US/Canada, May 2026) adds a marketing-specific revenue outcome to the data infrastructure gap:

  • Only 46% of marketing organizations have a fully centralized SSOT for customer data
  • SSOT ownership correlates with 44% revenue growth vs. 8% without — a 36-point differential (GrowthLoop is a CDP vendor; treat as directional)
  • Only 23% of marketers can reliably link marketing actions to business outcomes
  • 77% see “winning” A/B tests fail at full-scale deployment — the data latency mechanism
  • Teams using real-time signals are 22pp less likely to see test failures and 2x more likely to see high-impact experiment results

The marketing layer adds a revenue-outcome consequence to data fragmentation that most data readiness research frames in terms of cost or quality. When experiments cannot distinguish current from historical behavior, the entire test-and-learn methodology produces systematically misleading signals.

Source: research/05-analyst-firms/growthloop-ai-marketing-performance-index-2026.md · May 2026 · MEDIUM / TIER 1

Celonis 2026 Process Optimization Report — Process Data as the Missing AI Input (n=1,649, June–July 2025)

The Celonis survey adds an under-researched dimension to data readiness: not data quality in the traditional sense (missing values, schema inconsistency), but process data — the operational context that AI agents need to understand how work actually flows.

  • 90% of leaders say process improvement requires accurate, contextual data about how the business actually runs. 67% have concerns about the data they currently use for this purpose.
  • The departmental visibility gap is severe: 78% of Supply Chain leaders can’t get real-time end-to-end supply chain visibility; 74% of Finance/SS leaders say fragmented systems prevent real-time process views; 66% of IT leaders lack visibility into how technology is actually used across the organization.
  • 72% report that different departments have different views of the same process — meaning there is no shared ground truth that an AI agent can use as its operational model. This is the data readiness problem specific to agentic deployments: not whether data is clean, but whether it is coherent across organizational boundaries.
  • Starting-point mismatch compounds the data problem: 52% begin process improvement with workshops or BI dashboards that don’t produce the machine-readable operational data AI requires. This creates a feedback loop: poor process data → AI failures → continued use of insufficient tools.
  • Corroborates OneStream/Harris Poll (n=352, March 2026): heavy AI users are 4× more likely to have made a bad-data decision; 47% made material decisions on faulty data.

Source: research/07-adoption-challenges/celonis-2026-process-optimization-report.md · Celonis/Insight Avenue, n=1,649, June–July 2025 · MEDIUM-HIGH / TIER 2

Salesforce State of Marketing 2026 — The Personalization Data Bottleneck (n=4,450, Oct–Nov 2025)

The Salesforce 10th annual marketing survey (n=4,450, 26 countries, published Feb 2026) documents data fragmentation as the primary constraint on AI effectiveness in marketing — reinforcing the pattern from the State of Data & Analytics (2025) and State of Sales (2026) entries above.

  • 98% of AI-using marketing teams report at least one data-related personalization barrier. The near-universal prevalence confirms that data fragmentation is structural, not a matter of individual organizational maturity.
  • The average marketing organization manages 7 separate data sources. Only 51% have access to the cross-functional data (Sales, Service, Commerce) they need to build complete customer profiles. The other half is personalizing against an incomplete view of the customer.
  • Only 26% of marketers are satisfied with their organization’s data unification capability — consistent with the 84% “needs overhaul” finding from the State of Data & Analytics (n=7,652, 2025).
  • 51% say campaigns still feel generic despite AI adoption. This is the consumer-visible consequence of the data fragmentation: AI tools cannot produce relevant outputs from incomplete inputs.
  • High performers — the 2x more likely agentic AI users — save 8 hours per week through automation, but the differentiating factor is not tool access; it is data access. Organizations with unified customer data are the ones able to deliver the personalization that drives performance.

The marketing finding reinforces a pattern visible across functions: AI capability is rarely the binding constraint. The constraint is whether the underlying data infrastructure supports the use case. Solving the data problem before adding more AI tools is the higher-leverage investment.

Source: research/05-analyst-firms/salesforce-state-of-marketing-2026.md · Salesforce, n=4,450, 26 countries, Oct–Nov 2025 fieldwork · MEDIUM / TIER 1

Dun & Bradstreet Global AI Momentum Survey — The 5% Data Readiness Floor (n=10,000, Q1–Q2 2026)

The largest enterprise AI survey in the current corpus (n=10,000, 32 countries) produces a single number that anchors every data readiness conversation: only 5% of organizations say their data is fully ready for AI. This is the most credible floor estimate available because of sample scale.

  • 97% have active AI initiatives; 5% have ready data. The 92-point gap is not a failure story — it is a sequencing story. Organizations are deploying AI faster than they are resolving the data conditions those deployments require.
  • Only 10% express high confidence in their ability to identify and mitigate AI-related risks. The other 90% are deploying against a risk surface they cannot yet map.
  • Top data obstacles across 10,000 organizations: limited data access (50%), privacy and compliance risk (44%), data quality and integrity (40%), lack of cross-system integration (38%). Four of the top five barriers are data-layer problems.
  • The agentic escalation: D&B frames the data readiness gap as more consequential for agentic AI than for assistive AI. An agent that acts across systems — updating records, triggering workflows, querying supplier data — requires consistent entity identity across every system it touches. Most enterprise data environments cannot provide this at the agent’s operational speed.
  • ROI distribution: 60% report some measurable ROI; only 10% report strong ROI from multiple projects. The distribution is consistent with BCG’s finding (5% substantial value at scale) — ROI is real but concentrated among organizations that have resolved data prerequisites first.

Source: research/01-ai-native-landscape/dnb-global-ai-momentum-survey-2026.md · D&B, n=10,000, 32 countries, Q1–Q2 2026 · HIGH / TIER 1


Mayfield CXO Survey 2026 (n=266, January 2026) — Five Consecutive Years as the #1 Blocker

  • 58% of CXOs name data readiness/quality as the #1 blocker to agentic AI adoption. This is the fifth consecutive year this finding has topped Mayfield’s annual survey across Fortune 50–Global 2000 enterprises.
  • Sustained pattern across six years (2019–2026): The blocker has not moved despite sustained AI investment. This is not a “we need one more year to fix data” situation — it is a structural gap that changes slowly regardless of AI tool purchases layered on top.
  • Practical implication: As Tsvi Gal (CIO, Memorial Sloan Kettering) frames it: “AI demand is growing faster than compute, data pipelines, or governance can keep up. The only way forward is platformization — shared compute, shared data, shared guardrails.”
  • Consistent with D&B (n=10,000): 97% active AI initiatives vs. 5% ready data. AWS/HBR (n=623): 13% data architecture well-equipped for agentic AI. The Mayfield CXO finding is corroborated across multiple methodologies at different enterprise scales.

Source: research/12-agent-workers/mayfield-agentic-enterprise-cxo-survey-2026.md · Mayfield, n=266 CXOs Fortune 50–Global 2000, January 2026 · MEDIUM-HIGH / TIER 1

Enterprise Abstention + RAG Pipeline — Data Quality as a Hallucination Driver (May 2026)

The abstention and hallucination literature surfaces data quality as a direct upstream cause of unreliable AI output:

  • Hallucination rates above 15% in structured analysis tasks even with best prompting; legal domain 58–88% (arXiv:2603.08274, 172B tokens, March 2026). Poor source data quality is the most common proximate cause.
  • Document curation is flagged as critical by both Microsoft Azure and AWS Prescriptive Guidance: RAG chunks must be self-contained with clear titles — garbage-in produces hallucination-out regardless of model quality.
  • 39% of AI-powered customer service bots were pulled back due to hallucination errors in 2024. The primary remediation in most cases was data pipeline cleanup, not model replacement.
  • Enterprises choose RAG for 30–60% of use cases demanding high accuracy and auditability (AWS, 2025) — making data pipeline quality the primary determinant of whether those use cases succeed or fail.

Source: research/21-benchmarks/enterprise-abstention-whitepaper-2026.md · Synthesis of vendor whitepapers + arXiv:2603.08274 · MEDIUM-HIGH / TIER 1–2

Moody’s 2026: Data Quality as the Dividing Line in Risk and Compliance AI (January 2026)

Survey of 600 global risk and compliance professionals (Moody’s, January 13, 2026) reveals the sharpest documented gap between AI users and non-users along the data quality axis:

  • 59% of active AI users rate their organization’s compliance data as high quality
  • Only 31% of non-users (those resisting or not yet trialing AI) rate their data as high quality
  • 69% of non-users cite inconsistent or unstructured data as their primary barrier

The 28-point data quality gap is not coincidental — it is bidirectional causation. Organizations with structured compliance data deployed AI successfully and became users. Organizations without it could not deploy effectively, remained non-users, and consistently report data quality problems. The practical implication: AI readiness assessments focused on tool selection are asking the wrong question. The correct diagnostic starts with the compliance data audit — what percentage is structured, current, and linked across systems?

Source: research/06-security-frontier/moodys-ai-risk-compliance-survey-2026.md · Moody’s · n=600 · Jan 2026 · MEDIUM / TIER 1


Stanford Enterprise AI Playbook (n=51 cases, Apr 2026) — LLMs as Data Cleaners

The conventional narrative that AI needs clean, centralized data before deployment is obsolete, according to Stanford Digital Economy Lab’s study of 51 production deployments.

  • Only 6% of implementations had data that was fully ready for AI deployment. The rest faced moderate to severe data challenges — and most succeeded anyway.
  • 88% of cases had LLMs unlock data that was previously inaccessible — voice transcripts, scanned documents, legacy code, scattered knowledge bases that earlier approaches (OCR, rules engines, manual tagging) couldn’t process at the required accuracy and scale.
  • 91% of implementations successfully processed unstructured data that would have been unusable two years earlier.
  • 59% had data scattered across multiple systems owned by different teams. Only 16% had fully centralized data. Success did not require centralization. It required access — integration layers (APIs, RAG architectures, multi-agent frameworks) that connected scattered data.
  • 75% of implementations cited proprietary data as a key factor in AI strategy; 47% explicitly described accumulated data as a competitive moat. The organizations generating the most value had been storing data — even imperfect data — long before they knew how they would use it.
  • “We’ve had partners tell us, hey, it would have taken us two months to clean this up, and you guys flagged all the data issues within a day.” — VP of AI, Professional Services Firm (Stanford DEL interview)

The design decision that mattered was not data cleanliness but access: whether the deployment had integration layers to reach data wherever it lived, and retrieval architectures (RAG, MCP) designed to work with fragmented sources.

Source: research/04-consulting-firms/stanford-enterprise-ai-playbook-51-deployments-2026.md · Stanford Digital Economy Lab, April 2026 · MEDIUM-HIGH / TIER 1

Source: research/07-adoption-challenges/pwc-digital-trends-operations-2026.md · PwC, n=767 US operations/supply chain leaders revenue ≥$100M, Jan–Feb 2026 · MEDIUM / TIER 1

  • 87% of US operations and supply chain executives say poor data quality has hampered their progress on digital initiatives — the highest blocking-constraint figure in the corpus for data quality.
  • Only 30% report significant improvement in data quality or reliability despite years of digital investment. Investment in infrastructure does not automatically translate to usable data.
  • The data quality gap is cross-industry: the survey covers industrial products (18%), energy/utilities (16%), tech/telecom (16%), pharma/life sciences (11%), health (11%), consumer markets (13%), and financial services/insurance (10%) — no industry segment escapes the constraint.
  • Only 4% of organizations succeed across all four AI maturity dimensions simultaneously; data quality is the foundational constraint underlying each of the other three.
  • Contrast with Stanford DEL finding (51 cases): LLMs can work with imperfect data through integration layers — but this requires intentional retrieval architecture (RAG, MCP), not the assumption that data problems will resolve before deployment.

OECD 2026 — Health Sector Data Utilization Gap (Country-Level)

Source: research/06-industry-verticals/oecd-scaling-ai-health-2026.md · OECD Health Division, 38-country survey, April 2026 · HIGH / TIER 1

The OECD provides a system-level data readiness measure for healthcare — the sector with the highest stakes for getting data infrastructure right.

  • Health data is 30% of all data generated globally, growing faster than any other sector — and yet less than 5% of that data is used for decision making (World Bank Group, 2023). The gap between data generation and data utilization is larger in healthcare than in any comparably data-rich sector.
  • Only 10% of medical imaging AI applications have reached national scale across OECD member countries, despite AI being used in health administration universally (100% of surveyed countries). The bottleneck is not technology — 75% of health AI solutions evaluated through RCTs demonstrate positive clinical impact. The bottleneck is data infrastructure: fragmented EHR systems, incompatible interoperability standards, and underdeveloped health data governance.
  • 63% of OECD countries have a national interoperability strategy for health data; 24% have none. Without interoperability, AI models trained on institution-level data cannot be deployed at population scale — the same model performs differently across facilities and produces systematically skewed outputs for populations underrepresented in the training institution’s patient cohort.
  • The OECD recommends FAIR data principles (Findable, Accessible, Interoperable, Reusable) as the policy foundation for health data governance — the same principles that underpin enterprise data architecture best practices in other sectors.

The healthcare data utilization gap is the most concrete system-level measure in the corpus of the cost of inadequate data infrastructure. For enterprise AI practitioners outside healthcare: the same fragmentation dynamics that limit health AI to <5% data utilization apply wherever data is siloed across departments, legacy systems, or incompatible schemas — the difference is that healthcare’s stakes are higher and its governance infrastructure more constrained.

ILO–NASK WP140 — Workforce Data Readiness for AI-Exposed Roles (Global, May 2025)

Source: research/07-adoption-challenges/ilo-nask-genai-jobs-global-index-2025.md · ILO Working Paper 140, n=52,558 task data points, 1,640 survey respondents · HIGH / TIER 2

The ILO–NASK exposure index has a specific implication for data readiness: passive AI deployment into high-exposure roles (Gradient 3–4) without role redesign is structurally equivalent to deploying AI without workflow redesign — the tool addresses individual tasks, but the organizational data flows, handoffs, and accountability structures remain calibrated to a pre-AI task bundle.

  • 24% of global workers are in occupations with measurable GenAI exposure. In high-income countries, 34% sit in some exposure gradient. For organizations with 500+ employees, this typically means 120–170 roles where task bundles have been partially disrupted.
  • The data problem is the task-bundle problem. AI tools can address individual tasks within a role; roles are defined by bundles of tasks. Unless the data flows, reporting structures, and output expectations of the role are redesigned around the new task bundle, AI handles the task but not the workflow — leaving the organizational data architecture unchanged and the productivity gain unrealized at the firm level.
  • The ILO finding that Gradient 4 exposure is concentrated in clerical and administrative roles (data entry clerks, accounting clerks, administrative secretaries) means these are precisely the roles generating the structured, high-volume transactional data that enterprise AI systems depend on. Automating the data-entry task without redesigning the data pipeline the task was maintaining introduces downstream data quality risk.

AI Daily Brief — Applied Compute Practitioner Fieldwork: Data Readiness as State of Mind (April 2026)

Source: research/13-multimodal-sources/ai-daily-brief/2026-04-xx-six-big-questions-enterprise-ai-adoption-compounding.md · AI Daily Brief citing Applied Compute (Michael Chan) fieldwork, April 2026 · MEDIUM / TIER 1

Applied Compute’s 6-month embedded fieldwork at large enterprises deploying AI into production provides practitioner-specific data readiness evidence:

  • “Data readiness is just a state of mind.” The gap between “we have data” and “we have data in a format an AI system can learn from” surprises every team — including teams that have already wrangled enterprise data for other purposes.
  • Most enterprise data was never structured with AI consumption in mind. Every agent deployment is fundamentally a data problem, regardless of how it is framed by the team initiating it.
  • The org chart is the actual deployment environment. Who controls data access is rarely a single person or system — it requires navigating the real org chart (not the one on paper) to unlock data flows for AI.
  • Enterprises do not realize all the pre-conditions (data provisioning, compute access, access governance) they must address before agents can operate. Timelines are not just slow — they are slower than the organization’s own estimates.

AI Daily Brief — Maturity Maps: Data as Ceiling Constraint Across All Enterprise Functions (Q2 2026)

Source: research/13-multimodal-sources/ai-daily-brief/2026-04-xx-ai-maturity-maps-q2-enterprise-readiness-benchmarks.md · Super Intelligent / Nathaniel Whittemore, 480+ studies, Q2 2026 · MEDIUM / TIER 1

The Super Intelligent maturity maps place data as structurally different from the other five dimensions — not one pillar among six, but the floor constraint that caps all others.

  • 8 of 10 enterprise functions score 1 or 1.5 on data readiness (on a 5-point scale where 3 = on track). The finding is consistent across all function-specific surveys in the 480+ study aggregate.
  • Without proprietary context feeding AI systems — customer history, deal data, codebase, operational records — organizations cannot progress past basic assisted usage regardless of how capable the underlying models become.
  • The data gap has a specific shape: it’s not “we don’t have data,” it’s “we don’t have data accessible to AI systems in usable form.” MCP servers, vector stores, and clean data pipelines are the infrastructure gap.
  • This finding directly reinforces Applied Compute practitioner fieldwork (six-big-questions episode): “Data readiness is just a state of mind” — the gap between having data and having AI-consumable data surprises every enterprise team.

See also