Top-down AI mandates reproduce the same failure modes regardless of budget. See the use case discipline framework that enterprise leaders use to reach scale.
Published
Last Modified
Topic
AI Adoption
Author
Jill Davis, Content Writer

TLDR: Enterprise AI programmes stall not because of weak technology or bad data, but because of poor decision sequencing before and after a pilot succeeds. Top-down AI mandates routinely reproduce the same failure modes regardless of budget because they bypass use case discipline, skip meaningful kill-switch authority, and treat production readiness as a technology problem rather than an organizational one.
Best For: COOs, Chief Transformation Officers, and VP Operations at mid-to-large enterprises who have received executive pressure to "move on AI" and are watching initiatives stall at the production threshold, despite apparent pilot success.
Enterprise AI implementation challenges are not primarily technical. According to RAND Corporation research, 80.3% of enterprise AI projects fail to deliver promised business value, and McKinsey's 2025 Global AI Survey found that while 88% of organizations now use AI in at least one function, only 39% can link any EBIT impact to it at the enterprise level. The gap is not a technology problem. It is a sequencing problem, a governance problem, and above all, a use case discipline problem.
What a top-down AI mandate actually does to an organisation
Top-down AI mandates fail because they invert the sequence in which good AI decisions get made. When a board or CEO issues a mandate to "have an AI strategy" or to "deploy AI across operations," the organizational response is predictable: an RFP goes out, a vendor or consulting firm gets selected, a pilot gets designed to satisfy the mandate rather than to test a specific operational hypothesis, and the pilot team spends the following months building something impressive in a controlled environment rather than something functional inside a real workflow.
Ronny Fehling, Chief AI Transformation Officer at HTEC, described this pattern with precision in a May 2026 episode of the Emerj AI in Business Podcast. The reason enterprise AI programmes stall is not the technology, Fehling argued. It is the sequence in which decisions are made before and after the pilot succeeds. Most pilots, he observed, are built "next to reality, not inside reality." Once production begins, the integration issues, compliance requirements, and user adoption obstacles that were carefully sidestepped during the pilot phase surface all at once, and the initiative runs out of momentum.
This is what distinguishes a mandate-driven AI programme from a discipline-driven one. The mandate creates pressure to show a working demo within a fixed timeframe. The discipline creates pressure to prove that an AI system works reliably inside the operational environment where it will actually run.
The mandate pressure loop
When executive pressure arrives before operational clarity exists, three things happen in sequence. First, use case selection gets driven by what can be demoed quickly rather than what will deliver measurable business value at scale. Second, the pilot team optimizes for the demo environment, making architectural and data choices that would never survive production conditions. Third, when the pilot succeeds and the board approves funding for scale, the team discovers that the conditions enabling pilot success do not exist in the real operating environment.
Gartner predicted in June 2025 that over 40% of agentic AI projects will be cancelled by the end of 2027 due to escalating costs, unclear business value, or inadequate risk controls. The root cause cited was not model capability or data infrastructure. It was that most AI projects are driven by hype and misapplied, executed by teams optimizing for exploration rather than delivery.
The mandate pressure loop is the organizational mechanism that produces exactly this outcome. It creates incentives for visible activity over operational discipline, and it consistently reproduces the same failure modes regardless of how large the AI budget is.
Why increasing the budget does not help
The most common executive response when an AI initiative stalls is to increase the budget. More engineers, more data infrastructure, more vendor support. This response is almost always wrong, and Fehling's analysis explains why. If the problem is not the technology, more technology spending does not fix it. If the problem is that the pilot was built next to reality rather than inside it, more engineers working on the same architecture will produce a better-engineered version of the same wrong thing.
A March 2026 survey found that 78% of enterprises have AI pilots underway, but only 14% have reached production scale. The enterprises sitting in that 64% gap between pilot and production are not there because they lack resources. Large enterprises with over 10,000 employees abandoned an average of 2.3 AI initiatives in 2025 alone, according to research by Folio3 AI, each carrying an average sunk cost of $7.2 million. The sunk cost is not the result of underinvestment. It is the result of investment directed at the wrong problem.
What use case discipline actually requires
Use case discipline is selecting, scoping, and sequencing AI initiatives based on operational reality rather than executive aspiration. It sounds obvious. In practice, it is rare, because it requires three conditions that top-down mandates actively undermine.
Condition 1: A hypothesis-driven selection process
Every AI initiative should begin with a falsifiable operational hypothesis, not a technology aspiration. The difference is significant. "We will use AI to improve our demand forecasting" is a technology aspiration. "If we deploy AI on our 90-day rolling demand signal for our top 200 SKUs in the Northeast distribution region, we expect to reduce inventory carrying costs by 12 to 18% within two quarters, and we will kill the project if we see less than 8% improvement after one full demand cycle" is a hypothesis.
The hypothesis makes the use case bounded, measurable, and killable. The aspiration makes it open-ended, hard to evaluate, and politically difficult to stop. Top-down mandates generate aspirations. Discipline generates hypotheses.
For enterprises looking to build this kind of discipline into their AI transformation roadmap, the hypothesis-first approach means delaying executive demos in favour of operational design work. That is a harder internal sell than a well-produced pilot, but it is the only sequence that produces reliable scale.
Condition 2: Production slices as the unit of work
A production slice is a narrow, bounded deployment of an AI system inside an actual operational workflow, not a parallel system running alongside the workflow but fully integrated into it. The distinction matters because the failure modes of "next to reality" systems only become visible when the system has to interact with real data quality, real exception handling, real user behaviour, and real integration constraints.
Fehling's framing of production slices as the unit of work changes how AI initiatives are scoped. Instead of building a pilot that demonstrates what AI can do under ideal conditions and then attempting to retrofit it into production, the production slice begins by identifying the narrowest possible real-world deployment that can produce a measurable operational signal. A production slice for a claims processing AI in an insurance operation might be a single claims type, in a single geography, for claims below a certain complexity threshold, running for a defined period with full audit logging and a clear rollback path.
The production slice is not a smaller pilot. It is a different kind of object. It is designed to surface integration failure, not to demonstrate capability under controlled conditions. This is why enterprises that build in production slice methodology as a standard approach scale AI more reliably, even when their initial deployments are smaller and less impressive than mandate-driven pilots.
Before running any production slice, most enterprises benefit from running an honest assessment of whether their AI pilot is actually ready to scale into real operational conditions.
Condition 3: Use case portfolio management
Most enterprises run AI initiatives as independent projects rather than as a managed portfolio. The consequence is that no one is making the cross-functional decision about which initiatives to accelerate, which to pause, and which to kill. Without portfolio oversight, the initiatives that survive longest are the ones with the most internal advocacy, not the ones producing the most operational value.
A McKinsey analysis found that only 6% of organizations qualify as AI high performers, defined as those seeing 5% or greater EBIT impact from AI. One of the distinguishing features of high performers is that they actively prune their AI portfolios, killing initiatives faster than average organizations and reallocating resources to initiatives showing early production traction. That pruning behaviour requires a governance structure with genuine authority to make those calls.
Decision gates and the kill-switch problem
One of the most common structural failures in enterprise AI programmes is the absence of what Fehling calls decision gates with genuine kill-switch authority. A decision gate is a pre-defined evaluation point at which an AI initiative must demonstrate progress against specified criteria before receiving further resources. Kill-switch authority is the organizational power to enforce a negative outcome at that gate without political interference.
Both elements need to be present for the gate to function. Decision gates without kill-switch authority are essentially retrospectives. The initiative continues regardless of what the gate reveals, because no one in the room has the mandate or the incentive to stop it. Kill-switch authority without pre-defined decision criteria is arbitrary. Stopping a project because it "doesn't feel like it's working" does not generate organizational learning and creates risk aversion that slows future initiatives.
What genuine kill-switch authority looks like
Effective kill-switch authority has three characteristics. First, it is held by a person or body that is not directly invested in the initiative's continuation. A project sponsor who built their quarterly objectives around the initiative succeeding cannot hold kill-switch authority over that initiative. The evaluating body needs to be separate from the delivery team.
Second, it is exercised at pre-defined intervals aligned to production milestones, not to calendar quarters. AI initiatives should be evaluated when they reach meaningful operational checkpoints, not because a fiscal quarter ended. Fehling's observation that enterprises need decision gates oriented toward learning rather than destination suggests that the question at each gate is not "are we on track to hit the annual target?" but "what did this production slice reveal, and does that justify the next slice?"
Third, kill-switch decisions need to be documented and shared. One of the highest-leverage learning mechanisms available to an enterprise is a well-written post-mortem on a killed AI initiative. Why was it stopped? What did the production slice reveal that the pilot had hidden? What would need to change before a similar initiative could succeed? Organisations that treat killed initiatives as failures to be quietly buried repeat the same structural mistakes. Organisations that treat them as expensive diagnostics accumulate the kind of operational AI knowledge that eventually becomes a competitive differentiator.
The enterprise AI stall problem is directly connected to the absence of this structure. When there is no mechanism for honest evaluation and clean termination, initiatives do not get killed. They get deprioritised, starved of resources, handed to smaller teams, and eventually abandoned without documentation. The organisation never learns what went wrong.
The 2026 governance shift
In 2026, the AI governance conversation in enterprise operations has shifted from "do we need guardrails?" to "what governance architecture enables speed?" This shift matters because it repositions governance as an accelerant rather than a brake. Enterprises that built meaningful governance structures earlier, before pressure to scale was acute, have the decision infrastructure to move faster now because they can evaluate and scale production slices with confidence rather than anxiety.
Gartner's January 2025 research found that 42% of organizations had made conservative investments in agentic AI, with most still in early-stage experimentation. The organizations that will benefit from agentic AI at scale are the ones building governance architecture now that can handle autonomous action with bounded authority, human-in-the-loop escalation paths, and audit trails that satisfy both operational and compliance requirements.
What sceptics get wrong
"We can't be disciplined about use case selection when the board wants to see results now."
This is the most common objection from operations leaders facing mandate pressure, and it is understandable. But it is not a choice between discipline and speed. It is a choice between apparent speed and real speed. Moving quickly to a polished pilot demo and then spending 18 months trying to make it work in production is slower than spending 8 weeks on use case scoping and building a production slice from the start. The discipline-first approach feels slower at the beginning. It is faster in aggregate.
"Our situation is different because we've already invested heavily in the pilot."
Sunk cost reasoning is the most reliable predictor of an AI initiative that will never reach production. The question is never "how much have we already spent?" It is "what is the probability that this initiative produces meaningful operational value, and what would we do with the resources if we freed them?" If the honest answer to the first question is below some reasonable threshold, the invested cost is not a reason to continue. It is information about how the initiative was structured.
"Our AI partner said the technology is production-ready."
Vendor claims about production readiness are claims about the technology, not about the operational environment the technology is being deployed into. The most capable AI system in the world will fail in production if it encounters data quality conditions it was not trained on, integration patterns it was not designed for, or user workflows it was not aligned with. Production readiness in the context of enterprise AI is a property of the deployment, not the model. No vendor can confirm it before the production slice runs.
The transition: from mandate to discipline
Making the shift from mandate-driven to discipline-driven AI requires three structural changes and one cultural one.
Structural changes: establish a use case governance function with portfolio authority; embed production slice methodology into the standard AI project template; create decision gates with pre-defined criteria and genuine kill-switch authority before any initiative receives full production funding.
The cultural shift is harder. It is the willingness to treat a killed initiative as useful information rather than a political liability. Organisations that can do this learn faster, recover faster, and eventually scale more reliably than those that maintain the fiction that every initiative is proceeding on track.
For enterprises ready to build this into a formal structure, connecting it to a broader AI governance framework ensures the decision gates and kill-switch mechanisms have institutional support rather than depending on individual leaders to enforce them.
The Ronny Fehling framework from the Emerj AI in Business Podcast episode (May 2026) is not a technology prescription. It is an organizational one. Enterprises that close the pilot-to-production gap in 2026 will not do so by selecting better AI models or hiring more engineers. They will do so by sequencing decisions correctly, scoping deployments inside reality rather than next to it, and building the governance infrastructure to kill initiatives cleanly when the production evidence calls for it.
The use case discipline framework: a summary table
Element | Mandate-Driven Approach | Discipline-Driven Approach |
|---|---|---|
Use Case Selection | What can we demo quickly? | What operational hypothesis can we falsify? |
Pilot Design | Built for controlled demo environment | Built as a production slice inside real workflow |
Success Criteria | Pilot impressiveness | Operational signal quality in production |
Governance | Progress reviews with steering committee | Decision gates with kill-switch authority |
Failed Initiative Response | Deprioritise and quietly abandon | Document, kill cleanly, share learning |
Resource Reallocation | Based on internal advocacy | Based on production traction signals |
Timeline Orientation | Calendar-driven quarterly updates | Milestone-driven production checkpoints |
Frequently Asked Questions
What is use case discipline in enterprise AI?
Use case discipline is the practice of selecting, scoping, and sequencing AI initiatives based on operational hypotheses rather than executive aspiration. It requires that every initiative begins with a falsifiable success criterion, is bounded to a specific operational environment, and has a pre-defined termination condition if production traction does not materialize within a defined timeframe.
Why do top-down AI mandates fail?
Top-down AI mandates fail because they invert the sequence of good AI decision-making. Mandate pressure drives teams to build demos that succeed under controlled conditions rather than production slices that succeed inside real operational workflows. According to RAND Corporation research, 80.3% of enterprise AI projects fail to deliver promised value, with poor sequencing and governance identified as the primary root cause rather than technology or data problems.
What is a production slice in enterprise AI?
A production slice is a narrow, bounded deployment of an AI system fully integrated into an actual operational workflow rather than running parallel to it. Unlike a pilot optimised for demo conditions, a production slice surfaces real integration failures, data quality issues, and user adoption obstacles in a controlled scope. It is designed to produce honest operational signal, not to demonstrate capability under ideal conditions.
What are decision gates in AI governance?
Decision gates are pre-defined evaluation checkpoints at which an AI initiative must demonstrate progress against specified criteria before receiving additional resources. Effective decision gates require genuine kill-switch authority held by a body separate from the initiative's delivery team. Without both elements present, gates function as retrospectives rather than real decision points, and failing initiatives continue regardless of what the evaluation reveals.
How does kill-switch authority work in practice?
Kill-switch authority is the organizational power to terminate an AI initiative at a decision gate based on pre-defined criteria, without political interference from initiative sponsors. It is held by a governance body independent of the delivery team, exercised at production milestones rather than calendar quarters, and its decisions are documented and shared as organizational learning. Ronny Fehling of HTEC argues this is one of the organizational conditions that most distinguishes enterprises that reach production scale from those that stall.
Why does increasing AI budget not fix the pilot-to-production gap?
Increasing budget does not fix the pilot-to-production gap because the gap is not a resource problem. It is a sequencing and governance problem. Enterprises abandoned an average of 2.3 AI initiatives in 2025 with sunk costs averaging $7.2 million each, according to Folio3 AI research. Adding budget to initiatives built next to reality rather than inside it produces better-engineered versions of the same wrong architecture.
What is the difference between a pilot and a production slice?
A pilot demonstrates what AI can do under controlled, often ideal conditions and is optimised for persuasiveness and demo quality. A production slice is a narrow deployment inside a real operational workflow, designed to surface failure modes rather than demonstrate capability. A pilot answers "can this AI work?" A production slice answers "does this AI work inside our operational reality and at what accuracy level?"
What percentage of enterprise AI pilots reach production scale?
Only 14% of enterprise AI pilots reach production scale, according to a March 2026 survey, despite 78% of enterprises running active pilots. The 64-point gap between pilot prevalence and production scale is the defining enterprise AI challenge of 2026, and it is driven by governance failures and missequenced decision-making rather than by technology limitations.
Why do agentic AI projects have higher cancellation rates?
Agentic AI projects carry higher cancellation rates because they inherit the same sequencing problems as standard AI projects but at greater organizational complexity and cost. Gartner predicts over 40% of agentic AI projects will be cancelled by end of 2027 due to unclear business value, escalating costs, and inadequate risk controls, with most current projects driven by hype rather than operational discipline.
How do AI high performers select use cases differently?
AI high performers use hypothesis-driven selection processes with falsifiable success criteria and defined termination conditions, rather than aspirational or mandate-driven selection. McKinsey research found that only 6% of organizations see 5% or greater EBIT impact from AI. High performers distinguish themselves by actively pruning their AI portfolios and reallocating resources to initiatives showing early production traction rather than maintaining failing initiatives for political reasons.
What organizational conditions determine whether an AI initiative reaches production?
The organizational conditions most predictive of production success are: use case selection grounded in operational hypotheses, pilot or production slice design inside real workflows rather than parallel to them, decision gates with genuine kill-switch authority held by a body independent of the delivery team, and a cultural willingness to document and share learnings from killed initiatives. According to Ronny Fehling of HTEC, these conditions matter more than technology capability or budget size.
How do you build a kill-switch into an AI programme?
Building a kill-switch into an AI programme requires three elements: a pre-defined set of criteria that would trigger termination (specific accuracy thresholds, user adoption rates, or operational metric targets), a governance body with the authority to enforce termination independent of initiative sponsors, and a documentation protocol that captures what the production evidence revealed. The ServiceNow AI Control Tower and similar enterprise governance platforms can automate kill-switch triggers for technical criteria, but the organizational authority component requires deliberate governance design.
What is the right role for an external AI partner in this framework?
An external AI partner in a discipline-driven programme serves a different function than in a mandate-driven one. Rather than delivering a complete solution to a broadly defined problem, the partner works inside the production slice methodology: scoping the narrowest viable first deployment, surfacing integration and data quality issues before they become production emergencies, and supporting the governance structure that enables honest decision-making at each gate. Partners who resist kill-switch authority or who scope work to avoid honest production evaluation are a red flag regardless of their technical capability.
What does the transition from mandate-driven to discipline-driven AI require?
Transitioning from mandate-driven to discipline-driven AI requires three structural changes and one cultural shift. Structurally: establish a use case governance function with portfolio authority, embed production slice methodology into the standard AI project template, and create decision gates with pre-defined criteria and genuine kill-switch authority before any initiative receives full production funding. The cultural shift is the willingness to treat killed initiatives as valuable organizational learning rather than political liabilities.
How does an enterprise AI transformation roadmap support use case discipline?
An AI transformation roadmap supports use case discipline by making the sequencing of decisions explicit rather than implicit. A well-built AI transformation roadmap defines which use cases are sequenced in which order, on what criteria, with what governance gates between phases. It prevents individual mandate pressure from overriding the portfolio logic by making the portfolio logic visible and agreed-upon at the executive level before individual initiative pressure arrives.
When should an enterprise bring in an external AI transformation partner?
Bringing in an external AI transformation partner is most valuable at two points: before use case selection, when an independent assessment of operational readiness and use case viability can prevent mandate-driven selection errors; and at the point where a production slice has revealed structural issues that internal teams do not have the architecture or governance experience to resolve. Partners engaged only after a pilot has already been built to an inflexible design have limited ability to change the sequencing failures that will prevent it from reaching production.
Legal
