AI-ready data

Clean Data vs AI-Ready Data

Adding semantic context to raw data triples AI accuracy on business questions.

Columnist · · 10 min read · Updated
AI-Ready Data Layer · September 26, 2026 · 10 min read · 2,331 words

Clean data removes errors. AI-ready data does something harder: it carries meaning, context, and governance an AI system needs to act correctly, and clean data was never built to carry any of that. Clean data was designed for human analysts who bring their own interpretation to a report, while AI agents have no institutional memory to draw on and nothing to fill in the blanks with.

A BI dashboard can tolerate blank fields, inconsistent naming, and vague column labels because the person reading it supplies the missing context from memory. A large language model fed that same dataset has nothing to work from but the schema itself, and a schema carries structure, not meaning. A column labeled "revenue" tells a model nothing about whether the number is gross or net, tax-inclusive, or already adjusted for refunds.

It helps to think of this as three eras of data readiness, stacked on top of each other. AI-ready data has to carry that context inside the data itself, in a form any general-purpose reasoner can pick up and execute without a human standing by to translate. That last shift is the one enterprises are still working through, and it is why the phrase "AI-ready" gets used loosely when it should be used precisely.

Making data AI-ready is an act of translation rather than an act of cleanup. It means transcribing what used to live only in someone's head, the tribal knowledge of what a field really means and which table is the one to trust, into a form a machine can act on directly. That distinction sets up everything that follows: the cost of skipping it, the properties that close it, and the architecture that makes closing it repeatable.

What the accuracy gap shows

The clearest evidence for this distinction comes from a controlled comparison of the same schema, the same model, and the same questions, with only one variable changed: whether the data carried explicit semantic context. Sequeda et al. (2023), as reported in the architecture research referenced here, found that zero-shot GPT-4 answered business questions over a raw enterprise schema correctly roughly one time in six. The only addition was explicit context, and it more than tripled accuracy on the same task.

The benchmark used an insurance industry standard data model, with tables covering claims, payments, and premiums, and questions ranging from straightforward reporting to multi-step metric calculations. What shifted was whether the meaning behind the schema had been written down somewhere the model could read it.

The result deserves a caveat, because the high-accuracy condition did not happen for free. None of that undercuts the core finding. It reframes it: the gain is real and it is substantial, but it has to be built, and the question for any enterprise is whether building it is worth the lift.

A second finding compounds the first. The shortfall is architectural, a limitation in the layer sitting beneath the model.

Diagram: Raw Schema vs. Semantic Context: The Accuracy Gap. Visualizes: Show a stark before/after magnitude contrast between two conditions tested on the same schema, same model (GPT-4), and same benchmark questions (insurance industry data model…

The five properties that separate AI-ready data from clean data

AI-ready data requires five properties that clean data was never asked to have: consistent structure across every source, resolved completeness, business context embedded directly in the data, traceable origin, and governed access. A gap in any one of them propagates into every output the AI produces downstream, because a model has no way to correct for a property that was never supplied upstream.

Business context is the property that traditional data engineering has spent the least time on, and it is the one this piece returns to most, because it is the one the accuracy gap above is actually measuring. Fields need tags and definitions attached so a model or an agent knows that "revenue" in this table means the same thing finance means when it says "revenue": gross or net, tax-inclusive or not, refunds subtracted or left in. A semantic layer, understood as a unified interface layered on top of existing infrastructure, is the mechanism that does the actual translating here: it takes what an analyst used to supply silently, which table is authoritative, what a given metric actually measures, and turns it into explicit metadata and definitions that any general-purpose reasoner can query directly.

Traceable origin means standardized identifiers and full lineage, so that when an output looks wrong, someone can walk it back to its source. It is the precondition for explaining an AI decision to a regulator or an auditor after the fact. Governed access means controls that make compliant, ethical use the default behavior of the data rather than a review step bolted on at the end, and for AI agents that act on data without a human checking every intermediate step, a permission that was never enforced was never a real permission to begin with.

One property on this list cuts against a decade of conventional data engineering wisdom. The practices enterprises refined for years, standardizing formats, deduplicating records, stripping out anomalies, can make data less ready for AI, not more. A fraud-detection model depends on the rare, outlying, anomalous record precisely because that is what fraud looks like. The five properties above give a working framework for AI readiness, grounded in guidance from the OvalEdge framework and the ReVisionz guide, but none of them resolve to a universal checklist. There is no such thing as data that is "AI-ready" in the abstract. The same dataset can be fully ready for a demand-forecasting model and dangerously unfit for a customer-facing agent making decisions in real time.

Why governance is the property most definitions skip

Governance is the property that separates AI-ready data from merely clean data, because lineage, access controls, sensitivity tagging, and business definitions are what make an AI output auditable, explainable, and defensible after the fact. Without that layer, every other quality property above is only half-built: data can be structurally consistent, complete, and context-rich, and still produce an output nobody can account for once a regulator or a customer asks how it was reached.

The real shift governance requires is moving a rule from a written policy into an encoded one. A field approved for fraud prevention and blocked for credit risk use carries that restriction into any system that touches it, and a model attempting to use it for the wrong purpose gets stopped at the data layer rather than relying on a reviewer who may not even be in the room to catch it.

A common assumption holds that governance slows AI projects down. The opposite is usually true: weak governance is what slows projects down, because teams lose time to legal review, security review, exception handling, and remediation when the underlying data estate was never organized for accountable reuse. Sensitivity also has to be evaluated at the point where data gets combined, not only at the level of a single field or source, since two individually harmless fields can become a genuine privacy or compliance risk the moment they are joined. Platforms built around a semantic layer are designed to enforce exactly this at query time, evaluating permissions under the real user's identity rather than a shared service account, so a join that should never happen simply doesn't execute.

Regulation is turning this from a best practice into a requirement. The EU AI Act, DORA, and the expanding set of data sovereignty frameworks all demand transparency, audit trails, and governance built into AI systems by design. The OvalEdge readiness framework lays out the encoded-policy argument directly, and the regulatory dimension sharpens why it matters now rather than later.

Five failure modes at the data layer

Across documented enterprise AI engagements, the model is rarely the point of failure. The data layer is, and it tends to fail in one of five consistent ways. The first is untrusted metric definitions: "revenue" means something different in finance, in sales, and in the data warehouse, and a model has no way to know which definition applies to the question in front of it. The second is fragmented entity resolution, where there is no certified, single source for the entities a model needs to reason about, so the same customer, asset, or product exists under several different identities across systems.

The third failure mode is batch pipelines feeding inference workloads that were never designed for them. A pipeline built for overnight reporting cannot support an agent making a decision in real time, and an agent running on that pipeline ends up acting on a version of reality that is already out of date. The fourth is governance that exists only on paper: policies are written down somewhere, but never encoded into the data or enforced at the moment a query runs. They offer no actual protection when it counts. The fifth is a metric layer that re-derives the same KPI differently in every tool that touches it. When one number is calculated four different ways across four different systems, every AI output built from it is contested before anyone even has a chance to review it.

The scale of this problem is not a marginal concern. A Cloudera survey found that 80% of global IT leaders admit their AI and data initiatives are constrained by limited data access across environments. A survey of senior IT leaders at large enterprises found that data permission or governance problems arose at some point during the AI project lifecycle in the vast majority of cases, and that most enterprises were missing at least one foundational capability required for responsible AI data activation.

None of these five failure modes is a one-off accident or a sign of a team that didn't try hard enough. They are the predictable output of a data architecture built for human analysts, who could absorb an inconsistent metric or a delayed report without much damage, and never redesigned for machine consumers that cannot.

The semantic layer as the primary architectural answer to the context gap

A semantic layer is a structured, governed abstraction that translates raw enterprise data into business meaning both humans and AI systems can trust, and by 2026 it has moved from an analytics nice-to-have into essential infrastructure for any enterprise AI initiative expected to reach production. For an AI system specifically, a semantic layer supplies what the schema alone cannot: table descriptions, metric definitions, and entity relationships an agent can read on its own, without needing a human to explain what a column means.

That structure resolves the failure modes named above almost directly. A single, governed definition for each metric means the same KPI no longer means something different in every tool that touches it, closing the fifth failure mode. Permissions enforced at the moment a query runs, evaluated under the real user's identity rather than assumed from a shared account, closes the fourth failure mode by encoding governance into the query layer itself instead of leaving it written down and ignored. Full lineage running from a query's output back to its source data means every AI output can actually be audited after the fact.

The industry has started standardizing around this idea rather than treating it as proprietary. The Open Semantic Interchange specification, published as an initial v0.1 release in January 2026, defines a vendor-neutral format for sharing business context, definitions, relationships, and access policies, between semantic layers and the AI systems that consume them. Partners in that effort include Snowflake, Salesforce, dbt Labs, BlackRock, RelationalAI, Atlan, Alation, Mistral AI, and ThoughtSpot. That list signals something specific: the semantic layer is becoming a standard interface across the industry, not a feature locked to any single vendor's stack.

A semantic layer by itself does not close the entire context gap. Research across a large set of enterprise queries found that agents given unified, multi-dimensional context outperformed agents using semantic definitions alone by a substantial margin. The fuller architecture for production AI in 2026 runs on four components working together: a knowledge graph for entity and relationship mapping with lineage, a governed semantic layer for metric definitions and business logic, a retrieval layer for unstructured content, and a governance infrastructure that keeps every output traceable. Each piece addresses a different shape of the same underlying problem, and none of them substitutes for the other three.

The pattern is already running in production, not just in specification documents. Distillery implemented AtScale's MCP Server to bring natural language data access into Slack and Google Meet, and large enterprises have gone on to standardize MCP across multiple LLMs, including Claude, GPT, and internal models, so that every AI system in use shares the same semantic foundation rather than reinventing its own. Strategy Mosaic, described in industry reporting as a universal semantic layer, has been running this same pattern in production at Pfizer, Hilton, and GUESS since before agentic AI became its own category. The work required to build an ontology and mappings by hand, the same kind of work that produced the accuracy jump in the Sequeda study, is precisely what a platform built around semantic governance is meant to automate across many sources at once rather than one schema at a time.

What AI-ready data demands from the pipeline

Even a carefully governed semantic layer fails if the data flowing beneath it is a delayed, inconsistent, or siloed version of reality. AI-ready data needs freshness and continuity calibrated to its actual use case, not inherited wholesale from batch schedules built decades ago for overnight reporting.

That freshness requirement is not uniform across every system. A data snapshot that is perfectly fine for a monthly dashboard is already stale the moment an agent tries to use it to make a customer-facing decision in real time. An agent making a decision without that check has no such safeguard, so the pipeline itself becomes part of what "AI-ready" has to mean.

This is the piece of the architecture that schema design alone cannot fix. Closing the context gap takes a semantic layer. Closing the freshness gap takes a pipeline built for the decision speed the business actually needs, rather than the reporting cadence it has always had.

Sources

  1. What Is AI-Ready Data? Complete Guide
  2. Data AI Readiness: What It Means and Why It Matters (2026)
  3. Writing Context Into Your Data — AI-Ready Data, Semantic Layers, Knowledge Graphs, Ontologies
  4. The State of the Semantic Layer: 2025 in Review
  5. Semantic Layer for Enterprise AI in 2026: What Production Use Requires

More in AI-Ready Data Layer