Share with your CDO
The AI stack isn’t a new frontier sitting alongside your data platform. It’s the same platform, wearing different clothes. Bill Schmarzo’s argument in this TechTarget opinion piece lands on a specific structural claim: two companies licensing the same model get different results because one feeds it governed, well-defined data and the other feeds it whatever it can find. An Omdia survey of 400 organizations found that 47% had governed semantic layers in some domains, but only 37% had extended them enterprise-wide. That gap is where AI failures live.
What this means for your business
If your organization has been treating its AI build as a parallel infrastructure project, separate pipelines, separate copies, separate governance, you’re not accelerating. You’re duplicating. The CDOs most exposed here are the ones who approved vector database deployments and embedding pipelines without first asking whether the definitions those systems retrieve are consistent across business units. “Revenue” meaning different things in finance and in the sales CRM is an analytics problem. Delivered confidently by an AI agent at scale, it becomes an operational one.
Schmarzo’s core structural argument holds, and it’s sharper than it first appears. Semantic layers, the governed dictionaries that give “active customer” or “net revenue” a single agreed definition across the enterprise, were built to make analytics trustworthy. They do the same work for AI retrieval, and the investment is already largely sunk for any organization running a mature analytics function. The mistake isn’t building AI capabilities. It’s building them outside the governance perimeter rather than on top of it, which forces a second trust-earning cycle that usually only starts after something visibly breaks.
The compounding dynamic Schmarzo identifies is the right frame for budget conversations with the CFO. Every reusable data product and governed definition lowers the marginal cost of the next AI use case. Organizations that extend their existing data architecture for AI purposes get that compounding. Organizations that build a parallel AI data environment pay setup costs twice and governance costs indefinitely. The falsification condition is straightforward: if AI use cases at your organization require fewer new definitions and shorter data preparation cycles over time, the architecture is compounding. If every new use case feels like starting from scratch, you’ve built a duplicate.
Concept deep-dive: Semantic layer
A semantic layer sits between raw data and the people or systems consuming it, acting as a translation dictionary that enforces shared definitions across the business. Think of it as the rule that says “active customer” means the same thing whether finance, marketing, or an AI agent is asking the question. Without it, every system that touches the data can interpret ambiguous terms differently. For AI specifically, that ambiguity doesn’t produce visible errors, it produces confident wrong answers.
Based on reporting from Why AI should build on, not copy, your data stack, originally published 2026-09-15 17:05:00.
