Why portfolio company AI pilots fail: the operating model was never tested, not the tech. Score your stalled pilots on 5 tests and see which earns scale money.
Published
Last Modified
Topic
AI Adoption
Author
Amanda Miller, Content Writer

TLDR: The honest answer to why portfolio company AI pilots fail is that the operating model around the pilot was never tested, not that the technology failed. Most stalled pilots worked as demonstrations. They died because nobody in the portco owned the number, the workflow ran beside the agent instead of through it, and the released hours never left the budget. The 5-Test Operating Model Autopsy scores a finished pilot on ownership, decision rights, workflow integration, capacity release, and repeatability, and tells an operating partner which pilots deserve scale money before the next partner meeting.
Best For: AI, digital, and value-creation operating partners at mid-market private equity funds whose portfolio companies are enterprise-scale (1,000 to 15,000 employees), who have two or three portcos with AI pilots that "worked" and nothing yet in EBITDA to show for them.
A portfolio company AI pilot is a time-boxed deployment of AI on one workflow inside a portco, funded to prove a value hypothesis that the fund can underwrite. Most of them stall, and the stall gets blamed on the wrong thing. When an operating partner asks why portfolio company AI pilots fail, the answer that comes back from the portco is usually about the technology: the model was not accurate enough, the data was messy, the vendor overpromised. Sometimes that is true. More often the pilot did what it was built to do, and the organization around it was never set up to absorb the result. The operating model was assumed, not tested. This post gives the operating partner's version of the diagnosis: how to tell a technology failure from an operating model failure, a five-test autopsy that produces a score, and a rule for which stalled pilots get a second round of money.
Why portfolio company AI pilots fail: the operating model was never tested
Portfolio company AI pilots fail because the fund and the portco test the technology and assume the operating model. The pilot proves an agent can classify invoices or draft responses. Nobody proves that a named manager will own the throughput number, that the workflow will route through the agent rather than around it, or that the hours released will come out of a budget line. When the pilot ends, the technology has a result and the organization has a slide.
The numbers back this up from several directions. FTI Consulting surveyed 200 fund and operating leaders for its 2026 Private Equity AI Radar and found that 95 percent of funds report AI initiatives meeting or exceeding their original business case, yet only 7 percent of portfolio companies have reached enterprise-scale deployment. Both facts are true at once. The business case was written for the pilot, and the pilot met it. Scale was never part of the case.
BCG reached a similar split in January 2026: roughly 60 percent of companies have yet to realize measurable value from AI, and only 5 percent qualify as what it calls future-built. The Bain and StepStone 2026 GP Outlook, a survey of more than 100 investment professionals in early 2026, found nearly 40 percent of GPs do not expect any material financial impact from AI in their portfolios this year. That is not a technology forecast. It is a confession that the pilots already running are not connected to anything.
The demo trap
A demo answers one question: can the AI do the task? A pilot that scales has to answer four more: who owns the output, what changes in the workflow, where do the released hours go, and can the next portco run it without the first one's champion. Funds skip those four because the first one is exciting and the other four are boring. EY calls the result "use case syndrome" and reports that 62 percent of PE respondents in its Q4 2025 AI pulse cannot link specific productivity gains to AI adoption. Gains that cannot be linked cannot be underwritten, and gains that cannot be underwritten do not get scale money.
What "operating model" means for one pilot
An operating model, at the scale of one pilot, is five things: a person who owns the number, a written rule for who decides what the agent may do, a workflow that routes through the agent, a budget line that shrinks when hours are released, and a runbook someone else could follow. If any one is missing, the pilot can succeed technically and still produce nothing the CFO can find. The EBITDA map for portfolio companies shows where AI value sits in a portco; this post is about why the value stays theoretical when the operating model is skipped.
Technology failure vs. operating model failure: telling them apart
A technology failure and an operating model failure look the same from the partner meeting: a pilot that ran, a vendor invoice, and no EBITDA. They are different problems with different fixes, and the first job of an operating partner is to sort each stalled pilot into the right column. Technology failures are rarer than portco CEOs report, and cheaper to fix, because they show up in measurable accuracy and cycle-time data.
Symptom | Technology failure | Operating model failure |
|---|---|---|
Accuracy or error rate | Below the threshold set at pilot start, on clean input | At or above threshold; errors cluster on inputs the process never standardized |
Throughput | Agent is slower than the manual step or times out | Agent is fast; the queue in front of it or behind it did not move |
Who noticed the stall | The pilot team, from system logs | Nobody, until the CFO asked what changed |
Released hours | None, because the agent could not do the work | Hours released, but headcount, overtime, and contractor spend unchanged |
Second portco | Cannot be attempted; the tool does not work | Never attempted; the first portco's champion left or got promoted |
Fix | Different model, cleaner input, tighter scope | Ownership, routing, budget, and runbook, in that order |
Read the table against the stalled pilots in your portfolio and most will land in the right column. Deloitte surveyed 100 private company leaders in March 2026 and found 48 percent citing difficulty scaling beyond the pilot stage and 72 percent citing data quality, which sounds like technology until you ask why the data was never standardized. That is a process decision nobody made, and process decisions are operating model.
Why the right column is more common
The right column is more common because the pilot was designed by people rewarded for the left column. A vendor is paid for accuracy. An internal data lead is judged on whether the model works. Nobody in the pilot is judged on whether the AP manager's team got smaller or whether the second portco could repeat it. PwC surveyed 4,454 CEOs in late 2025 and found only 12 percent reporting both cost and revenue benefits from AI, with 56 percent reporting no significant financial benefit at all; the 12 percent were two to three times more likely to have embedded AI across operations rather than run it beside them. Embedding is an operating model choice.
The 5-Test Operating Model Autopsy
The 5-Test Operating Model Autopsy is a scored review an operating partner runs on a finished pilot to establish why portfolio company AI pilots fail in that portco and whether the pilot deserves scale money. Each test scores 0, 1, or 2 on evidence the portco can produce, not on what the pilot team says. Seven or more out of 10 is a scale candidate. Below 5 is a demo, however well the technology performed.
Test | Question | Input the portco must produce | 0 | 1 | 2 |
|---|---|---|---|---|---|
1. Ownership | Who loses something if this stops working? | Name and title of the owner; the KPI on their scorecard | No named owner, or the owner is the pilot lead | Named owner outside the pilot team, no KPI | Named P&L owner with the metric on their scorecard |
2. Decision rights | What may the agent decide alone, and who signs off on the rest? | Written approval rules and thresholds | Nothing written; humans re-check everything | Rules exist but were set by the vendor | Rules set by the function, reviewed by finance or compliance |
3. Workflow integration | Does the work route through the agent or beside it? | System logs showing share of volume handled by the agent | Under 25 percent of volume; agent is optional | 25 to 75 percent; parallel manual path still open | Over 75 percent; manual path closed except exceptions |
4. Capacity release | Where did the hours go? | Before and after hours, plus the budget line that changed | Hours reported, no budget line moved | Hours reallocated informally | Headcount, overtime, or contractor line reduced or explicitly redeployed |
5. Repeatability | Could a second portco run this without the first one's champion? | Runbook, data requirements, and the name of the second portco | No runbook; knowledge sits with one person | Runbook exists, not yet tested | Second portco has run a dry pass from the runbook |
Running the autopsy in one week
Run the autopsy in five working days, one test per day, with the portco CFO and the function head in the room and the vendor out of it. Day one is a single question to the CEO: who owns this number? If the answer takes more than a sentence, score it 0 and move on. Day three usually produces the biggest surprise, because system logs almost always show a parallel manual path the pilot team did not mention. Day four is where the 3-column spend register earns its keep: if the register shows AI spend rising and no budget line falling, test four scores 0 regardless of what the hours report says.
Reading the score
A high score on tests one and two with a low score on three and four means the portco has the will and not the plumbing; fund the workflow fix, not a bigger model. A high score on three and four with a low score on one means the pilot is running on a champion's enthusiasm and will stop when they leave. A low score on five with high scores elsewhere is the most common pattern in a fund's best portco, and it is the reason the second portco never happens. IBM found in its 2025 survey of 2,000 CEOs that only 16 percent of AI initiatives had scaled enterprise-wide, against 25 percent that had met their ROI target; the gap between those two numbers is test five.
What the fund does that the portco cannot
The fund's job in a stalled pilot is to supply the three things a single portco cannot generate for itself: a comparison across portcos, a scale gate with money attached, and permission to close the manual path. A portco CEO will not shut down the parallel process on their own, because the pilot might fail and the old way is safe. The operating partner can make closing it a condition of the next tranche.
The sequencing rule is simple to state and hard to hold: no second portco gets funded on a use case until the first portco scores 7 or more on the autopsy. That rule stops the pattern where three portcos each run their own pilot on the same workflow, none reaches scale, and the fund pays for the same lesson three times. McKinsey analyzed 471 PE-backed companies across 30 countries in June 2026 and found that the ones embedding AI across the business trade at median revenue multiples more than twice those of companies running productivity pilots alone. The multiple is paid for the operating model, not the pilot.
For a fund with three or more portcos already spending on AI and nothing in the value creation plan to show for it, with a partner meeting inside a quarter, a generic approach is not enough; the answer is a dedicated portfolio program that first autopsies every live pilot against the five tests and publishes the scores to the deal teams, then picks the single highest-scoring portco and workflow and funds only the operating model gaps the autopsy found, and only then writes the runbook and the scale gate that the second and third portcos must pass before they receive money. A lighter version, a survey of the portcos and a maturity heat map, produces a slide and changes no score.
Where the operating partner sits in the pilot
The operating partner should be the person who signs the exit criterion before the pilot starts and runs the autopsy after it ends, and should stay out of the middle. Funds that put the operating partner inside the pilot get a pilot the portco does not own, which fails test one by construction. The first 100 days of AI in a PE-backed company set the pattern for who owns what; the autopsy checks whether the pattern held.
The two objections every operating partner raises
The two objections that come up whenever this method is presented are that the scoring table could be generated by any AI assistant in seconds, and that the whole exercise could be done with rules-based automation. Both are fair, and both point at the same distinction: the table is trivial, the inputs are not.
"I could ask an AI assistant to make that matrix in two seconds"
Yes. The matrix is not the work. The work is getting the portco CFO to produce the budget line that moved, pulling the system log that shows what share of volume the agent actually handles, and getting a CEO to name an owner in one sentence. Those inputs take a week of access and a fund's authority to request. An assistant can draft the questions; it cannot compel the answers, and it cannot tell the difference between a runbook that exists and one that has been used. Gartner expects more than 40 percent of agentic AI projects to be canceled by the end of 2027 for unclear business value, and unclear business value is exactly what an autopsy with real inputs makes clear.
"This could be done with RPA"
Sometimes it could, and if the workflow is fixed-rule and stable, it should, because rules-based automation is cheaper to run and easier to audit. The autopsy applies to an RPA pilot just as well as an AI pilot, and RPA pilots stall on the same five tests. The technology question, whether the workflow has enough variation to need AI, is the RPA-in-disguise check that belongs in a proposal review, before money is spent. The operating model question comes after, and it is the one that decides whether anything reaches EBITDA. The 2025 MIT NANDA research, reported by Forbes, put the share of enterprise AI pilots with no measurable P&L impact at 95 percent, and its explanation was not model quality but the avoidance of the friction that comes with changing how work is done.
Why portfolio company AI pilots fail differently from corporate pilots
Portfolio company AI pilots fail differently from pilots at an independent enterprise because the hold period compresses the timeline, the fund adds a second layer of ownership, and the value has to be legible in an exit model rather than in an internal dashboard. A corporate pilot can drift for two years. A portco pilot has to move a line in a value creation plan inside a hold period that is often five years or less.
The two-owner problem
At an independent company, a pilot has one owner: the function head who asked for it. In a portco there are two candidates, the portco function head and the fund's operating partner, and the pilot often ends up with neither. The function head assumes the fund is driving it because the fund suggested it; the fund assumes the portco owns it because it is their P&L. Test one exists because of this. KPMG found in January 2026 that 65 percent of large-company leaders cite agentic system complexity as their top barrier to scaling agents; in a portco, half of that complexity is figuring out which of two organizations is accountable.
How the industry's thinking on this has moved
Three years ago, fund AI programs were mostly a tools rollout: give every portco access to an assistant and count logins. By 2025 that had shifted to use cases, with each portco running its own pilot and the fund keeping a list. The current phase is operating model, and the evidence for the shift is in what high performers now report. McKinsey found in its 2026 State of AI survey that nearly three-quarters of AI high performers had fundamentally redesigned workflows because of AI, up from 55 percent the year before, while only 37 percent of all respondents could attribute any EBIT impact to AI. A September 2026 Harvard Business Review analysis by two operations professors made the same point from the other side: the pilots that fail are the ones that automated an old process instead of designing a new one.
What to bring to the partner meeting
The operating partner should bring three artifacts to the partner meeting: the autopsy score for every live pilot, the one portco and workflow that clears the 7-of-10 gate, and the budget line that will move when it scales. A fund that has those three can ask for scale money and defend it. A fund that has a heat map of AI maturity across the portfolio has a conversation.
Present the scores as a table with the five tests as columns and the pilots as rows, and let the partners see which column is empty. In most portfolios the empty column is capacity release, and the discussion turns from "does the AI work" to "why has no budget line moved," which is the right discussion. Pair the table with the CFO-ready view of how AI shows up in EBITDA so the partners can see what the scaled version would report. Where a pilot scores under 5, say so and close it; a closed pilot with a clear reason is easier to defend than a third one that is "still learning."
One trade-off should be stated before anyone asks. The autopsy produces uncomfortable scores for pilots that portco CEOs are proud of, and it will name owners who did not know they were owners. The alternative is to keep funding pilots that meet their own business case and never reach the value creation plan, which is the pattern the FTI, BCG, and Bain numbers above describe. Uncomfortable now is cheaper than unexplained at exit.
This analysis was developed using methodologies and operating experience from Assembly.
Frequently Asked Questions
Why do portfolio company AI pilots fail to scale?
Portfolio company AI pilots fail to scale because the operating model around the pilot was never tested. The technology usually did its job. What was missing was a named owner of the number, decision rights written by the function, a workflow routed through the agent, a budget line that moved, and a runbook a second portco could follow.
How do you tell a technology failure from an operating model failure?
A technology failure shows up in the data: accuracy below the pilot threshold on clean input, or throughput slower than the manual step. An operating model failure shows a fast, accurate agent while the queue around it never moved, hours reported but no budget line changed, and nobody noticing the stall until the CFO asked.
What is the 5-Test Operating Model Autopsy?
The 5-Test Operating Model Autopsy is a scored review of a finished AI pilot covering ownership, decision rights, workflow integration, capacity release, and repeatability. Each test scores 0 to 2 on evidence the portco produces, such as system logs and budget lines. Seven or more out of 10 makes the pilot a scale candidate; below 5 it is a demo.
Who should own an AI pilot inside a portfolio company?
A named P&L owner in the portco outside the pilot team should own the pilot, with the metric on their scorecard. A pilot owned by the pilot lead, or informally by the fund's operating partner, fails the ownership test. The operating partner signs the exit criterion before the pilot and runs the autopsy after, and stays out of the middle.
What does workflow integration mean for an AI pilot?
Workflow integration means the work routes through the agent rather than beside it. The measurable input is the share of volume the agent handles, taken from system logs. Under 25 percent means the agent is optional; over 75 percent, with the manual path closed except for exceptions, means the workflow changed. Most stalled pilots sit in between.
What is capacity release and why does it matter?
Capacity release is the test of whether hours freed by AI came out of a budget line. Hours reported on a slide with headcount, overtime, and contractor spend unchanged score zero. Hours that reduced a line, or were explicitly redeployed to named work, score two. Capacity release is the column most often empty across a fund's pilots.
What is a scale gate for portfolio AI?
A scale gate is a rule that no second portco gets funded on a use case until the first portco scores 7 or more on the operating model autopsy. The gate stops three portcos from piloting the same workflow with none reaching scale. The fund ties the next tranche to the gate and makes closing the manual path a condition.
How long does the operating model autopsy take?
The autopsy takes five working days, one test per day, with the portco CFO and the function head present and the vendor absent. Day one asks the CEO who owns the number. Day three pulls the system logs showing the agent's real share of volume. Day four checks the spend register for a budget line that fell.
How is a portfolio company AI pilot different from a corporate AI pilot?
A portfolio company AI pilot has a compressed timeline, two candidate owners, and a value test set by the exit model. A corporate pilot can drift for two years with one owner and an internal dashboard. A portco pilot must move a value creation plan line inside a hold period, and often ends up owned by nobody.
Could most portfolio AI pilots be done with RPA instead?
Some could, and if the workflow is fixed-rule and stable, rules-based automation is cheaper to run and easier to audit. The technology choice belongs in the proposal review before money is spent. The operating model autopsy applies to RPA and AI pilots alike, because both stall on the same five tests when nobody owns the number.
What data does the operating partner need to run the autopsy?
The operating partner needs five inputs the portco must produce: the name and scorecard of the owner, written approval rules and thresholds, system logs showing the agent's share of volume, before and after hours with the budget line that changed, and the runbook plus the name of a second portco. Statements from the pilot team do not count as evidence.
What do surveys say about why AI pilots stall in portfolio companies?
Surveys show pilots meeting their own business case while almost none reach scale. FTI Consulting's 2026 survey of 200 fund leaders found 95 percent of AI initiatives met their business case yet only 7 percent of portfolio companies reached enterprise scale. BCG found roughly 60 percent of companies had yet to realize measurable value from AI in January 2026.
What should an operating partner bring to the partner meeting about stalled AI pilots?
An operating partner should bring three artifacts: the autopsy score for every live pilot, the one portco and workflow that clears the 7 of 10 gate, and the budget line that will move at scale. Present scores as a table with tests as columns so partners see which column is empty. Close any pilot under 5 with a stated reason.
What is the demo trap in portfolio AI?
The demo trap is funding a pilot that answers whether AI can do the task and skipping the four questions that decide scale: who owns the output, what changes in the workflow, where released hours go, and whether another portco can repeat it. EY calls this use case syndrome; 62 percent of PE respondents cannot link gains to AI.
How should a fund sequence AI pilots across portfolio companies?
A fund should sequence pilots so one portco reaches a passing autopsy score before any other portco starts the same use case. Run the autopsy on every live pilot, fund only the operating model gaps in the highest-scoring portco, write the runbook, and set the scale gate. Later portcos start from the runbook, not a fresh vendor pilot.
When should a fund bring in outside help for stalled portfolio AI pilots?
Outside help is worth considering when three or more portcos have live pilots, none appears in the value creation plan, and the partner meeting is within a quarter. A useful partner delivers scored autopsies built on the portco's own evidence, a runbook, and a scale gate the fund keeps, rather than a maturity heat map.
Legal
