How to evaluate an AI proposal from a portco comes down to four questions: baseline, defined value, RPA in disguise, exit criterion. See how to score yours.
Published
Last Modified
Topic
AI Diligence
Author
Jill Davis, Content Writer

TLDR: How to evaluate an AI proposal from a portco comes down to four questions: is there a measured baseline, is the value defined before the build, is it RPA in disguise, and what is the exit criterion. Each question has named inputs and a 0 to 2 score, and the total maps to one of three decisions: fund the build, fund measurement only, or decline. The grid takes two seconds to draw and two weeks to fill, and the filling is the work.
Best For: AI, digital, and value-creation operating partners and value-creation directors at mid-market PE funds whose portfolio companies are enterprises of 1,000 to 15,000 employees, with an AI proposal from a portco on the desk and a partner meeting ahead.
An AI proposal from a portco is a request for fund-level approval or capital to build or buy an AI system inside a portfolio company, usually arriving as a vendor quote, a projected saving, and a slide that says "agent." Evaluating one is different from evaluating a capex request, because the saving is unmeasured, and different from evaluating an RPA proposal, because the work being automated may involve judgment the proposal has not accounted for. The operating partner's job is to turn the proposal into something the investment committee can underwrite or to say why it cannot be. This post gives the four questions that do that, the inputs each question needs, and the rule that turns the scores into a decision.
Why most operating partners have no method for how to evaluate an AI proposal from a portco
Most operating partners have no method for how to evaluate an AI proposal from a portco because the proposals started arriving faster than the fund built a standard for them. Each portco sends its own format, each vendor supplies its own ROI, and the operating partner ends up judging enthusiasm rather than evidence. The result is a portfolio where AI activity is high and proof is scarce.
The surveys now describe that portfolio in detail. Grant Thornton's 2026 PE AI Impact survey found 46% of PE leaders scaling AI across functions while only 24% report revenue growth from it, and just 9% are very confident they could pass an AI governance audit within 90 days, against 22% in other industries. The Bain and StepStone 2026 GP Outlook found nearly 40% of GPs expect no material financial impact from AI in their portfolio companies during 2026, with whatever benefit exists skewed toward cost savings. Those two numbers together are the problem. The portcos are spending, the GPs do not expect it to show up, and nobody has a method for deciding which proposals deserved the money.
The proposal arrives with the wrong numbers
A typical portco AI proposal contains a vendor's per-transaction saving multiplied by the portco's annual volume. It rarely contains the current cost of the workflow, the share of volume the system will actually handle, or what happens to the items it cannot handle. FTI Consulting's 2026 Value Creation Index, drawn from 555 PE leaders, found that 66% reported AI benefits within 12 months, which sounds encouraging until you read the next line: only 31% described their AI implementation as efficient, and only 5% of non-high-performer firms exceeded their AI business case. The business case was the vendor's number. The 5% is what happened to it.
Why the hold period makes this urgent
The arithmetic of the hold period has changed, and that is what makes the evaluation standard matter now rather than next year. Bain's 2026 midyear PE report calculates that a deal which needed 5% annual EBITDA growth a decade ago now needs roughly 12% to return the same 2.5x over five years. An AI proposal that burns a year on a pilot with no baseline is a year of that 12% gone. The operating partner cannot afford to approve on enthusiasm, and cannot afford to decline everything either.
A short history of the question
In 2023 the operating partner's AI question was "should we have a policy." In 2024 it was "which portcos should pilot something." By 2025 the pilots existed and MIT's NANDA report was reporting that 95% of enterprise AI pilots showed no measurable P&L impact. The 2026 question is narrower and harder: this specific proposal, from this specific portco, yes or no, and on what evidence. That is what the grid below answers.
How to evaluate an AI proposal from a portco: the 4-Question Decision Grid
How to evaluate an AI proposal from a portco is a four-question test applied to the proposal as written: is there a measured baseline, is the value defined before the build, is it RPA in disguise, and what is the exit criterion. Each question is scored 0, 1, or 2 against named inputs. The total maps to fund the build, fund measurement only, or decline.
The grid is deliberately short. Operating partners sit on several boards and see a proposal a month; a 40-point scorecard will not be used. Four questions will. Here is the grid with the inputs each question requires, what a 0, 1, or 2 looks like, and the red flag that should end the conversation.
Question | Inputs to request from the portco | Score 0 | Score 1 | Score 2 | Red flag |
|---|---|---|---|---|---|
1. Is there a measured baseline? | Monthly volume, touch time per item, handoffs per item, exception rate, measurement window, data source | No baseline, or a survey-based estimate | Baseline reconstructed from system timestamps for 4 to 8 weeks | Baseline measured from system data over a full quarter, signed by the process owner | Baseline supplied by the vendor |
2. Is the value defined upfront? | The metric that will move, the expected delta as a range, the date it will be read, who reads it | Value stated as a projected annual saving only | One operating metric named with a target but no read date | Metric, delta range, read date, and reader all named, with the cost per run alongside | Value defined as "productivity" with no metric |
3. Is it RPA in disguise? | A written description of the decision the system makes, the inputs it reads, the share of items needing judgment | Deterministic rules on structured inputs, no judgment step, proposed as AI | Mixed: structured intake with a judgment step the proposal does not describe | Unstructured inputs and a judgment step, with the human review path described | Proposal cannot say what decision the system makes |
4. What is the exit criterion? | The measured result at which funding stops, the date of that reading, who makes the call | No stop rule | A review date with no threshold | A numeric threshold on the named metric, a date, and a named decision owner | Exit defined as "we will reassess" |
Question 1: Is there a measured baseline?
A measured baseline is the current cost of the workflow in time and errors, taken from system data, over a stated window. It is the single input that separates an underwritable proposal from a pitch. Ask for ticket, queue, or ERP timestamps rather than a survey of the team. If the portco cannot produce one, the proposal is not ready, and the honest response is to fund four weeks of measurement rather than the build. The baseline method is the same one used when assessing AI readiness during diligence, and a portco that was scored at acquisition often already has half of it.
Question 2: Is the value defined upfront?
Value defined upfront means one operating metric, an expected delta stated as a range, a date on which it will be read, and the name of the person who will read it. KPMG's Q2 2026 AI pulse of 2,145 senior leaders found that only 7% of organizations had reached established ROI on AI, and that those with full cost visibility reached it at 15% against 3% for those without. Cost per run is the visibility. A proposal that states the metric but not the cost per run (vendor fee, compute, human review time per item) has defined half the value. Score it 1.
Question 3: Is it RPA in disguise?
RPA in disguise is a proposal that applies deterministic rules to structured inputs with no judgment step, priced and positioned as AI. It is the second objection every operating partner raises, and it is a fair one. The test is to ask the portco to write down, in one sentence, the decision the system makes. If the answer is "it matches the PO number to the invoice and posts it," that is a rule, and a rules engine will do it for a fraction of the cost per run with a cleaner audit trail. If the answer is "it reads an unstructured denial letter, identifies the payer's stated reason, and drafts the appeal for review," there is judgment in the loop, and AI is the right tool. Gartner's June 2025 forecast that over 40% of agentic AI projects will be canceled by 2027 cites unclear business value as one of three causes, and Gartner's own estimate is that only about 130 of the thousands of vendors selling "agents" are real. Many of the rest are selling RPA with a new word on it.
Question 4: What is the exit criterion?
An exit criterion is the measured result at which the fund stops paying, with a date and a named decision owner. Proposals almost never include one, because nobody writes a stop rule into a pitch. Insist on it anyway. The investment committee will ask how the fund gets out of a failed AI program, and "we will reassess" is not an answer. IBM's 2025 CEO study found only 25% of AI initiatives delivered expected ROI and just 16% scaled enterprise-wide. The other 84% kept running for a while because nobody had written down when to stop them.
The decision rule
The four scores sum to a number between 0 and 8, and the decision rule is fixed in advance so the operating partner is not negotiating it with the portco CEO in the room. A total of 7 or 8 with no red flag means fund the build. A total of 4 to 6, or any proposal that scores 0 on question 1, means fund measurement only: four weeks to build the baseline and define the value, then re-score. A total below 4, or any red flag, means decline, with the grid returned to the portco so they know what a fundable proposal looks like. In the proposals Assembly has reviewed with operating partners across enterprise-scale portfolio companies in traditional industries, the most common first score is between 3 and 5, and the most common missing input is the exit criterion, not the baseline. The portcos usually know roughly what the workflow costs. Almost none of them have thought about when to stop.
AI proposal vs. capex request vs. RPA proposal
An AI proposal from a portco differs from a capex request in that its return is unmeasured until a baseline exists, and from an RPA proposal in that the work may involve judgment the proposal has not described. Operating partners default to treating all three the same way, through a payback period, and that is where most AI approvals go wrong.
Capex request | RPA proposal | AI proposal | |
|---|---|---|---|
What is being bought | A known asset with a known output | A rules engine for a defined process | A system that makes a judgment on variable inputs |
Return basis | Depreciation and capacity math | Transactions times cost per transaction | Baseline cost minus post-deployment cost, including human review and error handling |
Biggest unknown | Utilization | Process exceptions | Share of items the system handles correctly without review |
Right evaluation tool | Payback and IRR | Process map and exception rate | The 4-Question Decision Grid |
How it fails | Under-utilization | Exceptions pile up outside the bot | No baseline, so nobody can tell whether it worked |
The distinction matters more in 2026 than it did two years ago because the spend is moving. BCG's September 2026 Applied AI Index reports that AI spending nearly doubled in under a year, from 1.7% of revenue to 3.3%, and that more than 80% of it now sits outside the IT budget. In a portco, that means the request is coming from the CFO or the head of operations, through whatever approval path they use for other operating spend, and the capex discipline the fund applies to a plant expansion is not being applied to it.
When the grid is not enough
The 4-Question Decision Grid is enough for a single proposal in a single portco with a workflow that sits inside one function. It is not enough when the proposal touches customer-facing or regulated work, when the vendor is asking for a platform fee across several portcos, or when the fund is being asked to approve the same system in three companies at once. Those situations need a program, and the grid becomes its intake step rather than its whole method.
For a fund being asked to approve platform spend across several portfolio companies where the proposed system touches customer-facing or regulated workflows, a generic approach is not enough; the answer is a dedicated assessment program that does three things: it runs the same baseline method in every portco before any vendor is engaged so the proposals are comparable, it computes cost per run per workflow with human review time included so the platform fee can be judged against the actual unit economics, and it writes one exit criterion per portco into the approval so the fund can stop one company's deployment without stopping the others. The spend register that holds those numbers across the portfolio is described in how to govern AI spend across a portfolio.
Fund measurement, not belief
The most useful decision the grid produces is the middle one. A portco that scores 4 or 5 is usually not wrong about the opportunity; it is early. Four weeks of funded measurement, with the baseline method agreed and the cost per run defined, turns most of them into a 7. McKinsey's 2025 State of AI survey found that only 39% of companies report any EBIT impact from AI, and that the ones who do are far more likely to have redesigned the workflow before deploying. Funded measurement is where the redesign conversation starts, because the baseline usually shows that the workflow has more handoffs than anyone in the proposal thought.
What the investment committee will ask
The investment committee asks three things about portco AI spend: what it is measured against, when the fund finds out whether it worked, and how the fund stops it. The grid answers all three if the scores are attached to the approval memo. PwC's 2026 Global CEO Survey found 56% of CEOs saw neither revenue gains nor cost reductions from AI, and Forrester expects enterprises to defer a quarter of planned AI spend into 2027 because fewer than one-third of decision-makers can tie it to financial growth. An IC that has read those numbers will not approve a proposal on a vendor's slide. It will approve one with a baseline, a read date, and a stop rule.
The objections operating partners raise, and the honest answers
The objections operating partners raise to a formal AI proposal grid are usually two: that the grid is trivial to produce, and that the thing being proposed could be done with RPA. Both are partly right, and both are answered by what the grid asks for rather than what it looks like.
"I could ask an AI assistant to make that matrix in two seconds." Yes. The grid is four rows and anyone can draw it. The two weeks go into filling it: pulling timestamps from the portco's systems to build the baseline, getting the CFO to name the metric and the date, getting the vendor to put cost per run in writing, and getting the CEO to agree a number at which the fund stops paying. The matrix is free. The inputs are the work, and the inputs are what the IC underwrites. If a portco can fill the grid in two seconds, approve it in two seconds.
"What are you really doing here?" Turning a pitch into an underwritable request, or establishing that it cannot be. Nothing more. The operating partner is not evaluating the technology, which changes every quarter, and is not evaluating the vendor, which is the portco's procurement decision. The operating partner is evaluating whether the fund can tell, on a known date, whether the money worked. L.E.K.'s 2026 PE Pulse found 68% of funds rely on external advisors for value creation support and 46% expanded internal operations teams in the past year. Whoever does the work, the grid is what they are filling in.
"This could be done with RPA." Sometimes it could, and question 3 exists to find out. If the portco cannot describe a judgment step, send the proposal back as an RPA request with a lower cost per run and a cleaner control story. Nobody loses face. The portco gets its automation and the fund avoids paying AI prices for a rules engine. The pattern of pilots that survived approval and then failed for exactly this reason is laid out in why portfolio company AI pilots fail, and the map of where AI does produce EBITDA inside a portco, which is where question 3 should send a good proposal, is in how operating partners drive EBITDA with AI.
This analysis was developed using methodologies and operating experience from Assembly.
Frequently Asked Questions
How do you evaluate an AI proposal from a portco?
You evaluate an AI proposal from a portco with four questions: is there a measured baseline, is the value defined before the build, is it RPA in disguise, and what is the exit criterion. Each scores 0 to 2 against named inputs. A total of 7 or 8 funds the build, 4 to 6 funds measurement only, below 4 declines.
What is the 4-Question Decision Grid?
The 4-Question Decision Grid is a scoring method for a single portco AI proposal. It asks for a measured baseline, a value defined upfront with a read date, a written description of the judgment the system makes, and a numeric exit criterion. Each question scores 0, 1, or 2, and the total maps to fund, fund measurement only, or decline.
What should an AI proposal from a portfolio company contain?
An AI proposal from a portfolio company should contain a system-data baseline of the workflow's volume, touch time, handoffs, and exception rate, one operating metric with an expected delta range and a read date, the cost per run including human review, a one-sentence description of the decision the system makes, and a numeric stop rule with a named decision owner.
What is a measured baseline in an AI proposal?
A measured baseline is the current cost of a workflow in time and errors, taken from system timestamps over a stated window of at least four weeks and signed by the process owner. A survey of the team is not a baseline, and a baseline supplied by the vendor is a red flag. Without one, the proposal cannot be underwritten.
How do you tell if an AI proposal is RPA in disguise?
An AI proposal is RPA in disguise when the system applies deterministic rules to structured inputs with no judgment step. Ask the portco to write the decision the system makes in one sentence. Matching a PO to an invoice is a rule. Reading an unstructured letter and drafting a response for review involves judgment. Only the second needs AI.
What is an exit criterion for a portco AI investment?
An exit criterion is the measured result at which the fund stops paying, with the date of the reading and the person who makes the call. Gartner expects over 40% of agentic AI projects to be canceled by 2027; a written stop rule means the fund runs that decision rather than the vendor.
When should an operating partner fund measurement instead of the build?
An operating partner should fund measurement instead of the build when the proposal scores 4 to 6 on the grid, or scores 0 on the baseline question. Four weeks of funded measurement builds the baseline, names the metric, and defines cost per run. Most proposals that start at 4 or 5 re-score at 7 afterward.
How does evaluating an AI proposal differ from evaluating a capex request?
Evaluating an AI proposal differs from a capex request because the return is unmeasured until a baseline exists. A capex request has a known asset and known output, so payback math works. An AI proposal's biggest unknown is the share of items the system handles correctly without review, which only a baseline and a read date can resolve.
What does the investment committee ask about portco AI spend?
The investment committee asks what the AI spend is measured against, when the fund finds out whether it worked, and how the fund stops it. The grid answers all three when its scores are attached to the approval memo. PwC's 2026 CEO Survey found 56% of CEOs saw no financial benefit from AI, so committees now expect evidence.
How do you compare AI proposals from different portfolio companies?
You compare AI proposals across portfolio companies by scoring each on the same four questions and comparing cost per run on the named metric. The grid makes proposals from a distributor and an insurer comparable because both must show a baseline, a read date, and a stop rule. Vendor projections are excluded from the comparison entirely.
Why do most portco AI proposals score low on the first pass?
Most portco AI proposals score low on the first pass because they contain a vendor projection instead of a baseline and no exit criterion. In the proposals Assembly has reviewed with operating partners, the typical first score is 3 to 5, and the missing input is most often the stop rule, not the baseline.
Is the decision grid too simple for a serious AI investment?
The decision grid is simple to draw and hard to fill, which is the point. The four rows take seconds; the inputs take about two weeks: system timestamps for the baseline, a named metric from the CFO, cost per run from the vendor in writing, and a stop rule from the CEO. The committee underwrites the inputs.
What is cost per run and why does the grid ask for it?
Cost per run is the full cost of the AI system handling one item: vendor fee, compute, and human review time per item, set against the baseline cost of that item today. KPMG's Q2 2026 pulse found organizations with full cost visibility reached established ROI at 15% against 3% without, so the grid treats it as half of question 2.
When does a fund need an assessment program instead of a grid?
A fund needs a dedicated assessment program when a proposal touches customer-facing or regulated workflows, or when platform spend is requested across several portcos. The program runs the same baseline method in every company, computes cost per run with review time included, and writes one exit criterion per portco so one deployment can stop without stopping the others.
How long does it take to evaluate an AI proposal from a portco?
Evaluating an AI proposal from a portco takes about two weeks when the portco has system data, and four to six weeks when the baseline must be built first. FTI Consulting's 2026 index found only 31% of PE firms describe their AI implementation as efficient; two weeks on inputs prevents a year on an unmeasured pilot.
Should the operating partner evaluate the vendor as part of the proposal?
The operating partner should not evaluate the vendor; the grid evaluates whether the fund can tell on a known date whether the money worked. Vendor selection is the portco's procurement decision. The operating partner's questions are about baseline, defined value, the judgment step, and the stop rule, all of which hold regardless of which vendor the portco picks.
Legal
