88% of enterprise AI agents fail to reach production. See the 3 infrastructure gaps Comcast identified for scaling AI from pilot to production.
Published
Last Modified
Topic
AI Adoption
Author
Jill Davis, Content Writer

TLDR: Scaling AI from pilot to production is where enterprise transformation initiatives succeed or collapse. According to a June 2026 episode of the Emerj AI in Business Podcast featuring Shri Nandan, VP of AI Experiences at Comcast, the deciding factors are not technical: they are gaps in data continuity, human-agent coordination, and cross-team governance. This article extracts the concrete lessons from Comcast's deployment experience and applies them to the enterprise context.
Best For: COOs, VP Operations, and transformation leads at mid-to-large enterprises in traditional industries who have approved an agentic AI pilot and need a structured approach to avoid the 88% pilot-to-production failure rate.
Agentic AI is a category of AI systems that take autonomous, multi-step actions on behalf of users or business processes, executing tasks across multiple systems without requiring human approval at each step. Unlike earlier automation, agentic AI adapts to new information and orchestrates decisions across workflows in real time. For enterprises in manufacturing, logistics, financial services, and operations-heavy industries, the pilot-to-production gap is where most AI transformation programmes stall. The failure is rarely the technology. It is the surrounding infrastructure that was never built to support it at scale.
Why Scaling AI from Pilot to Production Is the Hardest Part of Enterprise AI
Scaling AI from pilot to production is harder than most enterprise leaders anticipate because the conditions that make a pilot successful are the exact conditions that production environments do not have. A controlled pilot runs on clean, curated data. Production runs on the messy, fragmented, multi-system reality that defines how large enterprises actually operate.
The numbers are stark. According to IDC research cited by Folio3 AI, for every 33 AI prototypes built, only 4 reach production, an 88% failure rate. A broader survey of enterprise AI deployments found that 95% of AI pilots fail to scale beyond proof of concept. And a February 2026 CrewAI survey found that 79% of enterprises have adopted AI agents in some form, yet only 11% are running them in full production, a 68-percentage-point deployment gap that represents the largest backlog in enterprise technology history.
This is the context in which Shri Nandan, VP of AI Experiences at Comcast, joined the Emerj AI in Business Podcast in June 2026 to explain what separates enterprises that scale agentic AI from those that fragment. Comcast operates one of the largest customer service networks in the US, and customer experience is among the highest-stakes environments for AI deployment: it is real-time, visible to customers, tightly regulated, and spans dozens of internal systems. The lessons from that environment apply broadly.
Why Agentic AI Makes the Gap Worse
Traditional AI deployments (predictive analytics, document classification, anomaly detection) operate in bounded contexts. They ingest defined inputs and produce defined outputs. Agentic AI is different: it routes, decides, hands off, and executes across systems, sometimes without a human in the loop at all. Every gap that existed in a bounded AI deployment becomes magnified.
Analysis from Databricks' State of AI Agents 2026 found that governance and evaluation failures are the most consistent source of agentic AI collapse at scale, not model quality. Accelirate's 2026 agentic AI research attributes 41% of production failures to infrastructure gaps, 38% to governance and security barriers, and just 33% to ROI measurement issues. Put differently: the technology is rarely the problem. The problem is the operating environment the technology was deployed into.
The Cost of Skipping the Foundation
The cost of a production failure in agentic AI is not just the wasted build cost. It is the cost of customer erosion, compliance exposure, and the organizational trust that gets spent on a programme that doesn't deliver. Gartner projects that more than 40% of agentic AI projects will be cancelled by the end of 2027. That is a governance and infrastructure prediction, not a technology one.
For enterprises that do get agentic AI to production, the returns are significant. Research compiled by Digital Applied found an average 171% ROI (192% in the US) for agentic AI agents that reach production, rising to 540% ROI within 18 months for full production deployments. The delta between the 88% that fail and the 12% that don't is not technology or budget. It is foundational infrastructure.
Before committing to a production deployment, most enterprises benefit from an honest AI readiness assessment to understand where their real gaps lie. The three gaps Comcast identified are a useful diagnostic lens.
The Three Data Foundations Comcast Built Before Scaling AI
The first and most underestimated blocker to scaling AI from pilot to production is data. Not data strategy in the abstract sense, but specific data architecture decisions that determine whether an AI agent can maintain context across a customer interaction, a production workflow, or a multi-system transaction.
In the Emerj episode, Nandan identifies three data foundations required for what she calls "context continuity in production." Context continuity is the ability of an AI agent to know, at every step of an interaction, what has already happened and what the user or system state actually is. Without it, the agent starts every sub-task from scratch, produces contradictory outputs, and requires human intervention to resolve conflicts that the data layer should have prevented.
Foundation 1: Unified Interaction History
The first data requirement is a unified record of all prior interactions with a user, process, or entity. In a customer service context, this means every previous call, chat, service request, and account change is accessible in a single schema that the AI agent can query in real time. In an operations context, it means every prior workflow state, exception, and resolution is part of the agent's working memory.
Most enterprises fail this requirement because their interaction history is fragmented across business units, each of which runs its own CRM, ticketing system, or ERP instance. The pilot environment curated data manually, making the problem invisible. Production exposes it immediately.
Foundation 2: Real-Time System State
The second data requirement is accurate, up-to-date system state. An agentic AI making a decision about a customer account, an inventory level, or an approval routing path needs to know the current state of the systems it is acting on, not a cached snapshot from four hours ago.
Research from ClarityArc Consulting found that data quality and latency issues are the primary cause of production AI failures in 60% of enterprise deployments. Batch data pipelines are sufficient for reporting. They are not sufficient for an AI agent that is executing transactions in real time.
Foundation 3: Explicit Context Transfer Between Agent Handoffs
The third data requirement is the most frequently missed. When an agentic workflow passes a task from one AI sub-agent to another, or from an AI agent to a human, the receiving party needs explicit, structured context: not just a summary, but the precise state that the sending agent was operating in.
In Comcast's environment, this became critical when AI agents handed off to human agents for complex or sensitive interactions. Without structured context transfer, human agents received incomplete information, customers repeated themselves, and call resolution times increased. Comcast built an explicit context package that travels with every handoff, and that architecture change was what moved their query resolution rate from 65% to 95% over 12 months, as documented in their earlier CX deployment work.
The journey from 65% to 95% took a year of iteration, a dedicated business intelligence team mining interaction data daily, and consistent investment in the data layer rather than in new AI capabilities. That sequencing matters. The instinct in most enterprises is to invest in more capable AI. The correct investment, at this stage, is cleaner data.
Human-AI Orchestration: The Problem Most Production Plans Miss
The second major gap in scaling AI from pilot to production is human-AI orchestration. This is the architecture that determines when the AI acts autonomously, when it flags for human review, and how the handoff between AI and human is structured.
Pilot environments paper over this problem because pilots run with attentive oversight. Every borderline case gets reviewed. Every edge condition gets handled manually. In production, that overhead is not available. Without a deliberate orchestration architecture, the AI either escalates too much (reducing efficiency gains) or too little (creating compliance and quality risk).
Defining the Decision Boundary
The most important orchestration decision is where the autonomous decision boundary sits. Below the boundary, the AI acts without human review. Above it, the AI flags the decision for human confirmation or takes no action. Defining this boundary requires input from legal, compliance, operations, and customer experience, not just the technology team that built the pilot.
In customer-facing workflows, the decision boundary often depends on transaction type, customer segment, and regulatory context. A telecoms customer service AI might handle routine billing queries fully autonomously, but flag any churn escalation to a human. A logistics AI might reroute shipments autonomously below a certain value threshold but require approval for exceptions above it.
Analysis of enterprise AI deployments from Accelirate found that organizations with clearly defined decision boundaries have a 3x higher production success rate than those where the boundary is implicit or undefined. Only 21% of organizations have a mature governance model for autonomous AI agents, according to agentic AI statistics compiled for 2026.
The Human Side of the Equation
Nandan's experience at Comcast highlights something that most production planning documents miss: the impact on the human agents who work alongside AI. When the AI handles routine volume, human agents have more time per complex interaction. Comcast's contact center data showed agents spending 20 minutes with individual customers on difficult problems. That is a meaningful difference from the pre-AI pattern of rushed calls and high hold times. Employee satisfaction improved alongside customer metrics.
This is not a soft outcome. It is an orchestration design outcome. Building an AI deployment that improves human agent experience requires explicit design choices about what the AI takes on, what it leaves to humans, and how the transition between the two is structured. For enterprises considering how AI impacts their workforce, a structured AI workforce upskilling roadmap is essential preparation before production deployment.
Cross-Team Governance: Why the North Star Matters More Than the Technology
The third gap is governance, and it is the most organizationally difficult to close. Scaling AI from pilot to production requires a governance model that spans CX, IT, and operations, not a governance committee that sits in one of those functions while the others operate independently.
Nandan's framing in the Emerj episode is blunt: cross-team governance is "a single North Star across CX, IT, and operations." That phrase describes something specific. Not a steering committee that meets quarterly. Not a shared Slack channel. A unified decision-making authority with a single set of success metrics that all three functions are accountable to.
What Governance Failure Looks Like
The most common pattern is governance that exists at the pilot stage and then fragments at scale. The pilot is small enough that the original project team makes all the decisions. When production expands to multiple departments, those decisions need to be made by people who weren't part of the original implementation. Without a governance structure that extends to them, each department makes locally optimized choices that create global problems.
52% of organizations cite data quality as the biggest blocker to AI deployment, but data quality problems are almost always governance failures in disguise. If IT owns the data schema and CX owns the interaction design and operations owns the process exceptions, no single team is accountable for the data quality that the AI agent actually requires. The result is a production system that degrades faster than it is maintained.
Building a governance structure before production deployment is a known best practice that most enterprises skip because it is organizationally difficult. Enterprises that have built AI Centers of Excellence consistently report that cross-functional governance is the highest-leverage investment they made, not the AI capability itself.
What Good Governance Architecture Looks Like
A production-ready governance model for agentic AI has four components:
First, a unified success metric that all three functions (CX, IT, operations) agree on before launch. In Comcast's case, the North Star was customer resolution rate, a metric that required CX to own the interaction design, IT to own the data infrastructure, and operations to own the process exceptions. No single function could optimize for it alone.
Second, an escalation protocol that defines who makes which decisions when the AI encounters edge cases. This cannot be improvised in production; it must be defined before the first production deployment.
Third, a feedback loop that routes production signal back to the team that can act on it. Most governance models have monitoring. Far fewer have closed feedback loops where production signal triggers specific remediation actions by named owners.
Fourth, a refresh schedule for the AI's underlying data and decision logic. GEO 2026 research on AI agent performance decay documents that agents operating in dynamic environments require systematic refresh cycles to maintain performance. Without governance ownership of that refresh, AI performance degrades silently.
The 3 Infrastructure Gaps That Cause Agentic AI to Stall at Scale
Based on Comcast's deployment experience and the broader pattern of enterprise agentic AI failures, the three gaps that most consistently kill production deployments are:
1. Data Fragmentation
Context that worked in a curated pilot environment breaks down at production volume when data is distributed across siloed systems. The fix is not a new data platform; it is explicit architecture decisions about how context is packaged and passed between systems and agents. The data work required is boring and time-consuming, which is exactly why most enterprises skip it until production has already failed.
2. Governance Vacuum
The informal approval processes that work for pilots collapse at enterprise scale. Multi-department production deployments require explicit ownership, defined escalation paths, and shared success metrics before launch. Organizations that have built this rigorously report 3x higher production success rates than those that rely on implicit governance.
3. Human-Agent Coordination Gaps
Handoff protocols designed for pilots, where attentive humans are always present, do not survive production conditions. Explicit orchestration architecture, defining when AI acts autonomously and how it transfers context to humans, must be designed before production launch. The enterprises that get this right see measurably better outcomes for both customers and employees, as Comcast's contact center data demonstrates.
For organizations managing risk alongside these deployments, AI risk management in regulated industries provides a practical framework for building compliance into the production architecture rather than bolting it on after incidents occur.
What Skeptics Get Wrong About Scaling AI from Pilot to Production
"We have the technology; we're ready to scale." The Comcast experience and the broader statistics both contradict this. Technology readiness and production readiness are different assessments. Production readiness requires data architecture, governance design, and orchestration definition, none of which are technology problems. An enterprise can have best-in-class AI technology and still fail the 88% failure rate test because the surrounding infrastructure was not built.
"Our IT team can handle governance." Governance for production agentic AI requires joint ownership across IT, business operations, and customer-facing functions. When governance is owned by IT alone, it optimizes for infrastructure stability rather than business outcomes. When it is owned by a business unit alone, it optimizes locally at the expense of data quality and systemic risk. The Comcast model, a shared North Star across all three functions, reflects what governance at scale actually requires.
"We can figure out orchestration as we go." This is the most expensive misconception. Orchestration decisions made during a production incident, under pressure and without documented protocols, become technical debt that compounds over time. Comcast built the orchestration architecture before the first production user landed on it. That sequencing is what allowed the query resolution rate to improve systematically rather than fluctuate with each incident.
Building the Foundation Before You Scale
The pattern from Comcast's experience is not novel, but it is consistently ignored. Enterprises focus their AI investment on AI capabilities and underfund the infrastructure, governance, and human coordination systems that determine whether those capabilities reach production. The result is the 88% failure rate the industry documents year after year.
The enterprises that reach production and sustain performance do the same three things: they build data infrastructure that supports context continuity, they design human-agent orchestration before launch rather than during incidents, and they stand up cross-team governance with a single North Star rather than siloed accountability. None of these are technology investments. All of them are operational design decisions.
For most enterprises, the right starting point is understanding where they currently stand against these foundations, not another AI pilot. A structured AI transformation roadmap built around these three infrastructure requirements is a more reliable path to production than any new capability investment.
Frequently Asked Questions
What does "scaling AI from pilot to production" mean in an enterprise context?
Scaling AI from pilot to production means moving from a controlled proof-of-concept deployment, run with curated data and manual oversight, to a live system that operates across production workflows, real customers, or full operational volume. According to Folio3 AI, only 4 in 33 AI prototypes successfully complete this transition.
Why does agentic AI fail at the pilot-to-production transition so often?
Agentic AI fails at scale primarily due to infrastructure gaps rather than technology problems. Accelirate's 2026 research attributes 41% of production failures to infrastructure, 38% to governance gaps, and 33% to ROI measurement failures. The conditions that make a pilot work (clean data, attentive oversight, bounded scope) do not exist in production environments.
What are the three data foundations required for context continuity in production agentic AI?
The three data foundations are: a unified interaction history accessible to all agents in a workflow, real-time system state that reflects current rather than cached information, and structured context transfer packages that travel with every agent-to-agent or agent-to-human handoff. Comcast's experience shows that all three must be in place before production launch.
What is the pilot-to-production failure rate for agentic AI in 2026?
88% of AI agents fail to reach production, according to IDC data cited by Folio3 AI. A CrewAI enterprise survey from February 2026 found that 79% of enterprises have adopted AI agents but only 11% run them in full production, representing a 68-percentage-point deployment gap.
What is human-AI orchestration and why does it matter for production deployments?
Human-AI orchestration is the architecture that defines when an AI agent acts autonomously, when it escalates for human review, and how context is transferred at each handoff. Without a defined orchestration model, production systems either over-escalate, reducing efficiency, or under-escalate, creating compliance and quality risk. This architecture must be designed before production launch, not during incidents.
What does "cross-team governance" mean for an enterprise AI deployment?
Cross-team governance in agentic AI means a unified decision-making structure with shared success metrics across CX, IT, and operations. In Comcast's model, all three functions aligned to a single North Star metric before launch, with explicit escalation protocols and feedback loops. Without this, departments optimize locally and create data quality and accountability gaps that degrade production performance.
How long does it realistically take to scale AI from pilot to production?
Timeline depends on infrastructure readiness. Comcast's journey from 65% to 95% query resolution in their AI CX deployment took 12 months of iterative improvement with a dedicated business intelligence team. Enterprises with fragmented data and no governance structure in place typically require 12 to 24 months from pilot approval to stable production performance.
What is the ROI of agentic AI deployments that reach production?
Digital Applied's compiled research shows an average 171% ROI for agentic AI agents that reach production (192% in the US), rising to 540% ROI within 18 months for full-scale production deployments. This is the return available to organizations that successfully cross the pilot-to-production gap.
What percentage of organizations have mature governance for autonomous AI agents?
Only 21% of organizations have a mature governance model for autonomous AI agents, according to 2026 agentic AI statistics. The remaining 79% operate with informal governance structures that work at pilot scale but fail when deployments expand to multiple departments or full production volume.
What was Comcast's experience scaling AI in customer experience?
Comcast, led by Shri Nandan as VP of AI Experiences, built agentic AI into its customer service workflows and reached a 95% query resolution rate after 12 months of iteration, starting from 65%. Key investments were data architecture, a dedicated business intelligence team, and structured agent-to-human handoff protocols. The deployment was discussed in a June 2026 episode of the Emerj AI in Business Podcast.
What role does data quality play in AI pilot failures?
52% of organizations cite data quality as the biggest blocker to AI deployment, according to zbrain.ai's enterprise deployment analysis. Data quality failures are almost always governance failures in disguise: when data ownership is split across departments with no unified accountability, the quality required for production AI is never achieved and maintained.
What is the difference between agentic AI and traditional AI automation?
Agentic AI takes autonomous, multi-step actions across systems without requiring human approval at each step. Traditional automation executes predefined rule sets in bounded contexts. Agentic AI adapts to new information, makes decisions across ambiguous situations, and orchestrates multiple systems simultaneously, which is why the production infrastructure requirements are significantly more demanding than those for earlier automation deployments.
How should enterprises decide when an AI pilot is ready to scale to production?
A pilot is ready to scale when three conditions are met: the data layer supports context continuity at production volume, the orchestration architecture is documented and tested, and cross-team governance with shared success metrics is in place. An AI production readiness checklist helps teams audit these conditions before committing to a production launch.
What should enterprises do first if their AI pilot has stalled before production?
The first step is diagnosing which of the three infrastructure gaps is the primary blocker. In most cases, the audit reveals data architecture problems fragmented interaction history or batch data pipelines that cannot support real-time AI decisions. Starting with an AI readiness assessment that specifically covers data, governance, and orchestration readiness gives transformation teams a clear prioritization.
What does a "single North Star" governance model mean in practice?
A single North Star governance model means one shared success metric that all functions involved in an AI deployment are accountable to, along with unified escalation protocols and closed feedback loops. In Comcast's case, the North Star was customer resolution rate. Every data, orchestration, and governance decision was evaluated against that single metric rather than against each function's local priorities.
What is Gartner's prediction for agentic AI projects through 2027?
Gartner projects that more than 40% of agentic AI projects will be cancelled by the end of 2027. The projection reflects governance and infrastructure failure, not technology failure. This is a direct consequence of the pattern this article describes: enterprises launching agentic AI without the data foundations, orchestration architecture, and governance model required for sustainable production performance.
Legal
