When AI Models Become a Commodity, Your Data Foundation Is the Last Moat

When AI Models Become a Commodity, Your Data Foundation Is the Last Moat

As AI models converge in capability, enterprise competitive advantage shifts to data quality and process standardization. SAP's Chief Architect explains the sequence that actually works — and why most enterprises are three years behind.

Published

Last Modified

Topic

AI Adoption

Author

Amanda Miller, Content Writer

TLDR: As AI models converge in capability and cost, enterprise competitive advantage is shifting to data quality, process standardization, and architectural foundation. The organizations that win the next decade of AI won't be the ones who chose the best model. They'll be the ones who made their data ready to use it.

Best For: COOs, CIOs, and enterprise operations leaders in traditional industries evaluating how to build durable AI capability when every competitor has access to the same tools.

In June 2026, Guillermo Vazquez, Chief Architect for Business Transformation Services at SAP America, sat down with the Emerj AI in Business Podcast to say something that should unsettle most enterprise AI strategies: the models don't matter as much as you think they do.

Not because AI models aren't powerful. They are. But the gap between GPT-5, Claude Sonnet, Gemini Ultra, and a dozen other frontier systems is narrowing faster than most enterprises can respond. Price per token is collapsing. Capability parity is real. Any mid-market manufacturer or healthcare operator can access the same inference engine as their largest competitor.

The battle for AI advantage has already moved downstream, to the quality, completeness, and structure of the data those models run on top of. What Vazquez described wasn't a technology problem. It was an architectural one. And it's a problem Assembly encounters at nearly every enterprise engagement we run.

The commoditization curve has already happened

When cloud computing arrived, the same dynamic played out. Amazon, Google, and Microsoft offered increasingly identical compute at increasingly low prices. The organizations that won weren't the ones who bet on the right cloud provider. They were the ones who built applications and data pipelines that could run anywhere and improve continuously.

AI is following the same curve, faster.

In 2023, having access to GPT-4 was a meaningful advantage. By mid-2025, the frontier had extended, prices had fallen by 80%, and open-source alternatives had caught up to the 2023 frontier. By 2026, enterprise AI model selection is roughly analogous to enterprise cloud selection circa 2018: important, but not where competitive differentiation lives.

Vazquez's point is precise. When every enterprise can access the same models, what those models have to work with becomes the differentiating factor. Harmonized data. Standardized processes. A clear map of which workflows are differentiating and which are table stakes.

That reframes the entire AI adoption conversation. The question isn't "which AI should we use?" It's "are our data and processes ready for AI to make a difference?" For most traditional industry enterprises — manufacturing, distribution, healthcare operations, financial services — the honest answer is no. And that gap is closing faster from the outside (model capability) than from the inside (data readiness). For a broader look at why this matters, see why AI transformation fails to deliver measurable results.

What SAP's Chief Architect found in the field

Vazquez works with organizations at the exact moment they're trying to bolt AI onto existing operational infrastructure. What he consistently finds is a structural problem that no model upgrade can fix.

Fragmented data is the first one. Most enterprises have accumulated data across dozens of systems: legacy ERP, acquired systems, departmental databases, spreadsheets that became critical operations tools. Each system has its own schema, its own update cadence, its own definition of "customer" or "inventory unit" or "completed order." When AI tries to reason across these systems, it encounters a Tower of Babel. The outputs are unreliable because the inputs are inconsistent.

Then there's process variation. Even when two facilities do the same thing — say, receiving a shipment — they often do it differently. One enters a purchase order number; another enters a vendor reference number. One updates inventory immediately; another batches updates overnight. Process variation that humans navigate intuitively becomes catastrophic noise for AI systems trying to automate or predict outcomes.

Third, most enterprises haven't differentiated their workflows. Not every process is equal. Demand forecasting, customer pricing, product configuration: these are where the business actually competes. Accounts payable, expense reporting, compliance documentation: these are table stakes that every competitor executes at roughly the same cost. Vazquez's observation is that most enterprises haven't done the hard work of mapping which is which. They're trying to apply AI everywhere, which means they're improving nothing in a lasting way.

These three problems compound each other. Fragmented data makes process standardization harder. Process variation makes it impossible to build reliable training data. Without a clear sense of which workflows are differentiating, there's no principled way to prioritize the cleanup.

The result is what Assembly's teams see in the diagnostic phase of nearly every client relationship: a wide gap between the AI ambitions on slide 4 of the leadership deck and the actual state of the data strategy those ambitions would need to run on.

The sequence that actually works

What Vazquez outlined — and what Assembly has validated across dozens of enterprise engagements — is a sequence to AI readiness that can't be shortcut.

Step 1: Establish data harmonization before you touch AI

This doesn't make the press release. But it's the step that everything else depends on. Data harmonization means creating a single, authoritative definition for every business entity (customer, product, supplier, location) and enforcing that definition across systems.

In practice, this usually requires a master data management initiative that predates any AI deployment. It means resolving conflicts between systems, cleaning historical inconsistencies, and establishing governance so that new data entering the organization conforms to the standard.

The organizations that started this work in 2022 or 2023 are the ones seeing real AI ROI in 2026. The ones who skipped it in favor of model selection are the ones running their third pilot that also didn't scale.

Step 2: Standardize processes before automating them

There's a principle in process improvement that predates AI by decades: don't automate a bad process. Automation amplifies whatever the process does, including its defects.

The same applies to AI. A system trained on process variation will encode that variation as normal behavior. A system asked to automate a process that humans complete in seven different ways will produce seven different outputs, none of which leadership trusts.

Vazquez is specific about what standardization means in the context of AI-enabled ERP. Before you deploy AI to a workflow, you need to document the single correct way it should execute, train your people to do it that way, and run it consistently long enough to generate reliable training data. This usually takes three to six months for a single process. Most enterprises have hundreds. Which is why the organizations who started years ago have such a durable advantage: not because they chose a better vendor, but because they did the boring foundational work first.

Step 3: Map differentiating workflows explicitly

Not all processes deserve the same AI investment. The processes where your business actually competes — where you win or lose customers, where margin is created or destroyed — deserve deep AI integration. The table-stakes processes deserve standardization and basic automation, but not the AI investment the differentiating ones do.

Most enterprises do this backwards. They deploy AI broadly, see underwhelming results, and then try to retroactively figure out where it should have been applied. Strategic mapping first, process standardization second, AI deployment third.

Step 4: Build the foundation so adaptation is seamless

This is the payoff Vazquez emphasized. Organizations that invest in data harmonization and process standardization don't just benefit from today's AI capabilities. They're positioned to absorb tomorrow's, too.

When agentic AI becomes the norm, when AI-to-AI workflows become standard enterprise infrastructure, the organizations with harmonized data and standardized processes will integrate those capabilities in weeks. The ones who skipped the foundation will spend months or years catching up, the same way they're doing right now with capabilities that already exist.

Why traditional industries face a steeper version of this problem

Assembly works primarily with enterprises in traditional industries: manufacturing, distribution, healthcare operations, professional services. These organizations face a version of the data foundation problem that's more acute than what digital-native companies deal with.

The reasons are structural.

Legacy systems run deeper. A company that has operated for 40 years has 40 years of accumulated technical debt. ERP systems that were state-of-the-art in 2005 but haven't been updated since. Middleware that connects systems in ways nobody fully understands. Business rules encoded in applications that haven't been documented in a decade. These companies don't just need to clean their data. They need to excavate it. For a practical look at integrating AI with legacy ERP systems, Assembly's framework outlines three layered patterns that work without a full system replacement.

Process variation is higher. Traditional industry operations grew organically — through acquisitions, regional expansion, decade after decade of local adaptation. A manufacturer with 12 facilities almost certainly has 12 different ways of doing the same fundamental operations. Standardizing those processes before AI deployment isn't just a data problem; it's a change management problem, a labor relations problem, sometimes a cultural one.

Data governance is also less mature. Digital-native companies were forced to develop data governance practices early because their entire business ran on data. Traditional industry companies often treated data as a byproduct of operations rather than a strategic asset. The concept of a "data owner" — a person accountable for the quality of a specific data domain — is often foreign. Building that governance structure takes years.

None of this makes AI transformation impossible for traditional industry enterprises. It makes it more important to sequence correctly. Skipping the foundation work doesn't accelerate the timeline. It creates rework loops that are far more expensive than doing it right the first time.

What Assembly sees in the diagnostic phase

When we run an AI readiness assessment for a new client, we're explicitly mapping the Vazquez problem: how wide is the gap between where the organization's data and processes are today and where they need to be for AI to deliver reliable, scalable results?

We assess five dimensions.

Data quality and harmonization. How consistent are business entity definitions across systems? How complete is historical data in key operational domains? How much reconciliation happens manually that AI would have to replicate?

Process standardization. For the top 20 workflows by operational impact, how many process variants exist across facilities, regions, or departments? How close are teams to actually following the documented processes?

Workflow differentiation. Has leadership mapped which workflows are competitive differentiators versus table stakes? Is there a clear prioritization framework for AI investment, or is deployment being driven by vendor pitches?

Technology architecture readiness. What's the state of integration between core systems? Are APIs available for the data sources AI would need? Is there a data warehouse that aggregates key operational data, or is integration still point-to-point?

Governance maturity. Are there designated data owners for key domains? Is there a process for resolving data quality disputes? Is there a change management function that can support the process standardization work that precedes AI deployment?

Most enterprises score well on one or two dimensions and poorly on three or four. That's not a failing grade. It's a starting point. The diagnostic tells us where to sequence the work.

What the Vazquez episode reinforces is that organizations with high scores across all five dimensions didn't get there by accident. They made deliberate decisions, years earlier, to treat their data and processes as strategic infrastructure rather than operational overhead.

The financial case for treating data as infrastructure

The marginal cost of accessing frontier AI capability is approaching zero. OpenAI, Anthropic, Google, and dozens of smaller players are competing to commoditize intelligence as a service. Within 24 to 36 months, accessing a frontier AI model will be roughly analogous to accessing compute on AWS: a utility cost, not a strategic differentiator.

What doesn't approach zero in cost is data infrastructure. Building a master data management capability, standardizing processes across a multi-site operation, establishing data governance that actually holds — these investments take years and require sustained organizational commitment. They can't be purchased overnight. They can't be copied by a competitor who hasn't done the work.

The organizations building data foundations today are building moats. Not the temporary advantage of being the first to adopt a new AI tool, which competitors can replicate in months. A structural moat that takes years to build and years to replicate.

Vazquez's framing from the Emerj episode maps directly to this. When models become commodities, differentiation isn't the model selection. It's the quality of what the models have to work with. And that quality is determined by decisions being made right now — decisions that most traditional industry enterprises haven't yet made.

What to do this quarter

Three concrete priorities:

Audit your data in the two or three domains where you're most likely to deploy AI in the next 18 months. Demand planning, quality control, customer service, finance, wherever the operational stakes are highest. For each domain, map every data source, identify conflicts in business entity definitions, and quantify the gap between your current data quality and what a reliable AI deployment would require.

Count process variants for your highest-impact workflows. For your five or ten operationally most important processes, do structured observation across facilities or departments to count how many process variants exist in practice. The number is almost always higher than leadership expects. That count becomes the input to your process standardization roadmap.

Assign data owners. Data quality improves when someone is accountable for it. Identify a data owner for each of your top operational data domains: not the IT person who manages the system, but the business leader who is responsible for the accuracy and completeness of the data.

None of these require an AI vendor. None require a platform selection. They require leadership attention and organizational commitment. And in 18 to 24 months, they'll determine whether your AI deployments scale or stall.

Source: This article draws on "AI Models as a Commodity and Why Data Foundations Decide Who Wins," an interview with Guillermo B. Vazquez, Chief Architect, Business Transformation Services, SAP America, published June 9, 2026 on the Emerj AI in Business Podcast.

Frequently Asked Questions

What does "AI model commoditization" mean for enterprise operations leaders?

AI model commoditization means frontier AI systems from OpenAI, Anthropic, Google, and others are converging in capability and falling in price, making model selection a commodity decision rather than a strategic one. For enterprise leaders, it means competitive advantage shifts from choosing the right AI model to building the data foundation those models run on.

What is a data foundation in the context of enterprise AI transformation?

A data foundation is the combination of harmonized data, standardized processes, and data governance that makes enterprise operations AI-ready. It includes master data management, consistent business entity definitions across all systems, and designated ownership for data quality — the infrastructure an AI system needs to produce reliable, production-grade results rather than demo-only outputs.

What are the four steps to building an AI-ready data foundation?

The four-step sequence is: establish data harmonization first, then standardize processes before automating them, map which workflows are competitive differentiators versus table stakes, then build the governance infrastructure that makes future AI adaptation seamless. Each step depends on the prior one — skipping ahead produces pilots that succeed in demos and fail in production.

What are the five dimensions Assembly uses to assess AI readiness?

Assembly's AI readiness diagnostic assesses five dimensions: data quality and harmonization, process standardization, workflow differentiation (competitive vs. table stakes), technology architecture readiness (APIs, data warehouses, system integration), and governance maturity (data owners, quality dispute resolution, change management capacity). Most enterprises score well on one or two dimensions and need targeted work on the others.

Why do AI pilots fail when companies skip the data foundation step?

AI pilots fail without data foundation work because they train and run on fragmented, inconsistent data, producing results that look accurate in a controlled demo but break under real operational conditions. Leadership loses confidence in AI broadly, not just in the failed tool. The resulting organizational skepticism often takes years to rebuild — longer than the foundational work would have taken.

What are the three structural data problems that prevent enterprise AI from scaling?

The three structural problems are fragmented data (inconsistent schemas and definitions across systems), unstandardized processes (operational variation that AI cannot reliably automate), and undifferentiated workflows (no clear map of which processes are competitive versus table stakes). According to SAP America's Chief Architect, these three compound each other and must be addressed in sequence.

What competitive moat does a strong data foundation create for traditional industry enterprises?

A strong data foundation creates a structural competitive moat because it cannot be replicated overnight. While a competitor can access the same AI models within weeks, building harmonized data, standardized processes, and mature data governance takes two to four years in a traditional industry enterprise. Organizations that built this foundation in 2022-2023 have a compounding advantage that grows each year.

How does data harmonization translate into measurable AI ROI?

Data harmonization improves AI ROI by eliminating the manual reconciliation, correction, and exception handling that consumes time in fragmented data environments. When business entity definitions are consistent across systems — customers, products, suppliers, locations — AI systems can automate decisions reliably rather than flagging exceptions for human review, converting manual labor costs directly into operational throughput.

What is the first step enterprises should take before deploying any AI tools?

The first step is data harmonization — establishing a single, authoritative definition for every business entity (customer, product, supplier, location) across all systems before any AI deployment. This predates tool selection, vendor evaluation, or pilot design. Organizations that begin with model selection and skip data harmonization reliably produce pilots that fail to scale once real operational data conditions apply.

How do you audit your data before an AI deployment?

An AI readiness data audit maps every data source in the domains you plan to deploy AI, identifies conflicts in business entity definitions across systems, and quantifies the gap between your current data quality and the threshold reliable AI requires. For automation use cases, that threshold is typically 95%+ completeness and accuracy for the fields the AI model relies on.

Who should own data quality in an enterprise AI transformation?

Data owners — business leaders accountable for the accuracy and completeness of a specific data domain — should own data quality in enterprise AI transformation. This is not the IT person who manages the system; it is the operations or finance or commercial leader whose function depends on the data being correct. Without designated data owners, quality improvements don't stick.

How do traditional industries build data governance for AI?

Traditional industry enterprises build data governance for AI by designating domain data owners, establishing a process for resolving data quality disputes across systems, and creating change management functions that support process standardization work. Because these industries grew organically through acquisitions and local adaptation, governance must also address process variants across facilities — not just data definitions.

How long does data harmonization typically take for a manufacturing company?

Data harmonization for a manufacturer typically takes six to twelve months for companies with fewer than five systems and relatively clean data, and two to four years for those with decades of technical debt, multiple acquired systems, and significant process variation across facilities. This is why starting now matters — most traditional industry enterprises are three to four years behind.

When should an enterprise start AI foundation work relative to tool selection?

Enterprises should start data foundation work before selecting any AI tools. Tool selection should be driven by what your data architecture can support, not the other way around. Beginning with tool selection and retrofitting data readiness to match it is the most common driver of AI pilot failure — and why so many organizations are running their third pilot without scale.

What is process standardization and why does it matter for enterprise AI?

Process standardization means ensuring that the same operational task — receiving a shipment, creating a customer order, closing an invoice — is executed the same way across all facilities, regions, and teams before AI is deployed. Without it, AI systems trained on process variation encode inconsistency as normal behavior, producing outputs no operations leader will trust for automated decisions.

Does having a modern ERP system solve the data foundation problem?

A modern ERP system significantly improves the starting point but doesn't close the data foundation gap. ERP systems enforce data standards within their modules, but most enterprises have critical data in legacy systems, spreadsheets, and acquired-company systems outside the ERP. ERP modernization is the beginning of integrating AI with legacy infrastructure, not the end of it.

When should an enterprise use external help for data foundation work before AI deployment?

External help for data foundation and AI readiness work is most valuable when internal teams lack data governance expertise or when the diagnostic needs to stay platform-agnostic rather than anchored to an existing vendor's roadmap. Companies managing an active transformation benefit especially from outside sequencing guidance. Assembly's AI readiness assessment surfaces these gaps in four to six weeks.

Your AI Transformation Partner.

Your AI Transformation Partner.

© 2026 Assembly, Inc.