How to Measure AI ROI per Workflow Across a Portfolio: The 4-Input Formula

How to Measure AI ROI per Workflow Across a Portfolio: The 4-Input Formula

How to measure AI ROI per workflow: four inputs in minutes, one capacity ratio, one table a committee can underwrite. See how your portcos compare on one scale.

Published

Last Modified

Topic

AI Diligence

Author

Amanda Miller, Content Writer

TLDR: Knowing how to measure AI ROI per workflow is what separates a portfolio AI program a partner can underwrite from one that only reports usage. The method has four inputs per workflow: baseline minutes per transaction, AI minutes per run, human in the loop minutes per run, and error minutes per run. Multiply the net by monthly volume and you get hours returned per workflow, which can be compared across portfolio companies that share nothing else.

Best For: AI operating partners, value creation directors, and digital operating partners at mid-market private equity funds whose portfolio companies are enterprise scale (1,000 to 15,000 employees), who need one number per workflow that holds up in an investment committee and works across companies with different systems.

Updated: October 2026

Measuring AI ROI per workflow is a unit economics method that treats each automated workflow as a production line with a before cost, a per run cost, and a defect cost, all expressed in time, not in tool spend. The method exists because portfolio AI spending is now reported at the company level, usage is reported at the tool level, and neither can be compared across a portfolio or defended in front of an investment committee. A workflow is the only unit that is small enough to baseline, large enough to matter, and common enough across portcos to benchmark. For an operating partner asked where to concentrate effort across six companies, a per workflow number is the first thing that makes the question answerable.

Why Does an Operating Partner Need to Know How to Measure AI ROI per Workflow?

An operating partner needs to measure AI ROI per workflow because every other unit of measurement fails the investment committee test. Company level spend cannot be tied to an outcome. Tool level usage cannot be tied to a workflow. Workflow level unit economics can be tied to both, and they can be compared across companies that run different systems, serve different customers, and report different margins.

The scale of the problem is now documented. BCG's September 2026 Applied AI Index found that AI spending has roughly doubled to 3.3% of revenue in a year, with more than 80% of it sitting outside the IT budget. Most of it is therefore invisible to the controls a fund already has. The same report found only 7.5% of companies have reached the point where AI value compounds, and that group is showing 2.8 times the EBITDA growth of laggards. The gap between the two groups is not spend. Both groups are spending.

What the committee actually asks

The investment committee asks three questions in sequence: what was the workflow costing before, what does it cost now, and how do we know the difference is real. PwC's 2026 Global CEO Survey found 56% of CEOs have seen no significant financial benefit from AI, and the companies that have are two to three times more likely to have embedded AI in specific processes instead of spreading it across the board. A per workflow measure is how an operating partner answers all three questions with the same table.

Where portfolio reporting breaks today

Portfolio AI reporting breaks at the handoff between the portco and the fund. The portco reports licenses, active users, and a vendor supplied productivity estimate. The fund receives three numbers that cannot be added together or compared with the company next door. Bain's 2025 Global Private Equity Report found that while most portfolio companies were testing AI, fewer than 20% had operationalized a use case with concrete, measured results. The 80% have mostly not failed at AI. They have skipped the measurement that would show whether it worked. The 3-column register for governing AI spend across a portfolio catches the spend side; this post supplies the outcome side.

The 4-Input Formula: How to Measure AI ROI per Workflow

The formula for how to measure AI ROI per workflow uses four inputs, all in minutes: baseline minutes per transaction before AI, AI minutes per run including routing, human in the loop minutes per run for review and approval, and error minutes per run (rework minutes per error times the error rate). Net minutes returned is the baseline minus the other three. Multiply by monthly volume for hours returned per workflow per month.

Written out:

Hours returned per month = (Baseline minutes − AI minutes − Review minutes − Error minutes) × Monthly volume ÷ 60

Every term is observable from system timestamps or a two week sample. None requires a vendor's estimate, a survey, or a dollar conversion. That last property is deliberate: the moment the formula is expressed in money, each portco's finance team reconverts it with its own loaded rates and the numbers stop being comparable.

Input 1: Baseline minutes per transaction

Baseline minutes per transaction is the median elapsed working time from trigger to close for the workflow before AI touched it, measured from timestamps in the system of record. Invoice received to invoice approved. Claim opened to claim adjudicated. Contract request to signature. If the portco cannot produce timestamps, the baseline is built from a two week observed sample, and the formula is flagged as estimated until timestamps exist. McKinsey's 2025 State of AI found that only 39% of organizations reported any EBIT impact from AI, and the high performers were distinguished by having redesigned and measured specific workflows instead of deploying tools broadly. No baseline, no workflow in the program.

Input 2: AI minutes per run

AI minutes per run is the elapsed time the automated step takes, plus the time spent routing the transaction into and out of the agent. For most back office workflows the model's own processing time is trivial; the routing time is not. Stanford's 2025 AI Index documented a roughly 280 fold decline in the cost of running AI at a fixed capability level between late 2022 and late 2024. This input is the smallest of the four in almost every workflow, and operating partners who focus on it are measuring the wrong thing.

Input 3: Human in the loop minutes per run

Human in the loop minutes per run is the review, approval, and exception handling time a person spends on each transaction after the agent has acted. This is the input most portcos forget, and it is routinely the largest. A workflow where the agent drafts and a person rewrites has not returned any hours. The 5-domain scored checklist for assessing AI readiness during diligence scores process maturity partly on this question: how many steps still need a human decision. Gartner's November 2025 finance survey found 91% of finance organizations reporting only low to moderate impact from AI so far; a review step that absorbs the time the agent saved is the most common reason.

Input 4: Error minutes per run

Error minutes per run is the rework time caused by agent errors, computed as rework minutes per error multiplied by the error rate per transaction. A 4% error rate with 45 minutes of rework per error is 1.8 minutes per run. That sounds small until the workflow handles 20,000 transactions a month. Gartner's June 2025 forecast that over 40% of agentic AI projects will be canceled by the end of 2027 cites inadequate risk controls alongside unclear value; the error term is where those two causes meet in a single number.

Workflow Unit Economics vs. Tool Level ROI: What Each Can and Cannot Tell You

Workflow unit economics and tool level ROI answer different questions, and portfolios that confuse them end up with usage dashboards that cannot be underwritten. Tool level ROI measures how much a license is used and what the vendor claims it saves. Workflow unit economics measures how many hours a specific process returned after review and rework are subtracted. Only the second can be compared across portfolio companies or presented to a committee as a baseline and a delta.

Dimension

Tool level ROI

Workflow unit economics

Unit of measure

Licenses, active users, vendor estimate

Minutes per transaction, hours per month

Baseline

Usually none

Required, from timestamps

Captures review time

No

Yes, as a named input

Captures rework from errors

No

Yes, as a named input

Comparable across portcos

No, tools differ

Yes, workflows recur

Survives committee questioning

Rarely

If the baseline is signed

What it is good for

Vendor negotiation, adoption tracking

Capital allocation across a portfolio

BCG's October 2025 analysis of the AI value gap found that 60% of companies reported no material value from AI despite investment and only 5% were generating value at scale; the 5% were distinguished by focusing on a small number of workflows with measured outcomes, not broad tool rollouts. Tool level ROI cannot see that difference. Workflow unit economics is built to.

Why minutes beat money for cross portfolio comparison

Minutes beat money because every portco loads labor differently, and a formula in money gets reconverted by each finance team until the numbers no longer agree. Hours returned per workflow per month is the same unit in a distributor in Ohio and an insurer in Texas. Conversion to EBITDA happens once, at the fund, using each portco's own rates, after the operating comparison has already been made. The CFO ready method for tracking AI's EBITDA impact in a portfolio company covers that final conversion; this formula stops one step before it.

How to Compare Portfolio Companies Using Cost per Run

Comparing portfolio companies on AI ROI per workflow requires normalizing hours returned by the size of the workflow, so a large distributor's accounts payable and a small services firm's can sit in the same table. The normalization is the capacity ratio: hours returned per month divided by baseline hours per month for that workflow. A ratio of 0.30 means the workflow returned 30% of the time it used to consume, after review and rework.

The portfolio comparison table

The comparison table has one row per workflow per portco and five columns: baseline hours per month, hours returned per month, capacity ratio, review share (review minutes divided by baseline minutes), and error share (error minutes divided by baseline minutes). Review share and error share are diagnostic. A workflow with a low capacity ratio and a high review share has an agent that works and a control design that eats the benefit. A workflow with a high error share has an agent that does not yet work. The two problems have different owners and different fixes, and the table tells them apart before anyone visits the portco.

IBM's 2025 CEO study found that only 25% of AI initiatives had delivered their expected return and only 16% had scaled across the enterprise. In a portfolio comparison table those two numbers become visible per workflow: the 25% have a capacity ratio above zero, the 16% have one that held for two consecutive quarters.

The concentration rule

The concentration rule follows from the table: concentrate operating partner time on the workflows with the highest baseline hours and the lowest review share, because those return the most capacity for the least redesign. Deloitte's 2025 State of Generative AI in the Enterprise found two thirds of organizations expected 30% or fewer of their experiments to scale within six months; a concentration rule based on measured baseline hours is how a fund gets its own scale rate above that line without funding every pilot.

For a portfolio where two or three companies are already spending on AI with nothing measured, a generic ROI template is not enough. The answer is a dedicated measurement program that does three things: baselines the top five workflows per portco from system timestamps before any new tool is approved, instruments review and error minutes as named inputs rather than folding them into a vendor estimate, and publishes one comparison table to the partner meeting each quarter with the capacity ratio per workflow. The properties that matter are a timestamp baseline per workflow, review and error measured separately, and a fund level table that uses the same units for every company.

What Skeptical Operating Partners Say About Per Workflow Measurement

Skeptical operating partners raise two objections to per workflow measurement, and both deserve a direct answer because both are partly right. The first is that the formula is trivial and a chatbot could produce the matrix in seconds. The second is that this is robotic process automation with a new name. Neither objection survives contact with the baseline step, which is where the actual work is.

"I could get a chatbot to make that matrix in two seconds"

The matrix, yes. The inputs, no. The formula is four terms because four terms is what a committee can follow; the value is in the baseline minutes per transaction, measured from timestamps the portco has to be persuaded to export, and in the review minutes, which no vendor reports and no portco volunteers. MIT's 2025 State of AI in Business report found 95% of enterprise pilots produced no measurable profit and loss impact, and the common factor in the 5% was a defined workflow with a measured before state. The matrix is the easy part. The before state is the work.

"This is RPA in disguise"

The measurement would work for RPA too, and that is the point. Unit economics per workflow does not care whether the automation is a rules engine or an agent; it cares whether the workflow returned hours after review and rework. What changes with AI is the error term and the review term: agents handle exceptions that rules could not, so baseline coverage goes up, and they make different mistakes, so error minutes have to be measured, not assumed. Capgemini's 2025 research on agentic AI found only 2% of organizations had deployed agents at scale; a fund that measures per workflow can tell which of its portcos is in that 2% and which is running a rules engine with a new label.

"Our portcos will not produce timestamps"

Some will not, and that tells you something. A portco that cannot export trigger to close timestamps for its top five workflows is ready for a systems project, not an AI program. EY's December 2025 research found most companies reinvesting AI productivity gains instead of cutting roles, so hours returned is the number the portco's own CFO will eventually need; timestamps are how they get it. The operating model autopsy for failed portfolio company AI pilots treats missing timestamps as the first test, for the same reason.

This analysis was developed using methodologies and operating experience from Assembly.

Frequently Asked Questions

What does it mean to measure AI ROI per workflow?

Measuring AI ROI per workflow means treating each automated process as a production line with a before cost, a per run cost, and a defect cost, all expressed in minutes. The output is hours returned per workflow per month, which can be compared across portfolio companies that run different systems and report different margins.

What are the four inputs of the formula?

The four inputs are baseline minutes per transaction, AI minutes per run, human in the loop minutes per run, and error minutes per run. Net minutes returned is the baseline minus the other three. Multiplied by monthly volume and divided by 60, it gives hours returned per workflow per month for the comparison table.

Why is the formula expressed in minutes instead of money?

Minutes are the same unit in every portfolio company, while money gets reconverted by each finance team using its own loaded rates until the numbers stop agreeing. Hours returned per workflow is compared first at the fund level; conversion to EBITDA happens once, afterward, using each company's own rates.

How do you get the baseline minutes per transaction?

Baseline minutes per transaction come from system of record timestamps, measured as median elapsed working time from trigger to close before AI touched the workflow. If timestamps do not exist, a two week observed sample stands in, and the workflow is flagged as estimated until the system can produce the real figure.

Which input do portfolio companies most often forget?

Human in the loop minutes per run is the input most often forgotten and most often the largest. A workflow where the agent drafts and a person rewrites has returned no hours. Gartner's 2025 finance survey found 91% of finance teams reporting low to moderate AI impact, and absorbed review time is the usual cause.

How is error minutes per run calculated?

Error minutes per run equals rework minutes per error multiplied by the error rate per transaction. A 4% error rate with 45 minutes of rework is 1.8 minutes per run, which looks small until the workflow processes 20,000 transactions a month. The term is where control failures and value failures show up as one number.

What is the difference between workflow unit economics and tool level ROI?

Tool level ROI measures license usage and vendor claimed savings; workflow unit economics measures hours a process returned after review and rework. Only the second has a baseline, captures review and error time, is comparable across portfolio companies, and survives questioning in an investment committee. The first is useful for vendor negotiation and adoption tracking.

What is the capacity ratio?

The capacity ratio is hours returned per month divided by baseline hours per month for the same workflow. A ratio of 0.30 means the workflow returned 30% of the time it previously consumed, net of review and rework. The ratio lets a large distributor and a small services firm sit in the same comparison table.

What columns does the portfolio comparison table need?

The table needs five columns per workflow per company: baseline hours per month, hours returned per month, capacity ratio, review share, and error share. Review share and error share are diagnostic; a low ratio with high review share points to a control design problem, while a high error share points to an agent that does not work yet.

What is the concentration rule for operating partner time?

Concentrate operating partner time on workflows with the highest baseline hours and the lowest review share, because those return the most capacity for the least redesign. Deloitte's 2025 research found most organizations expected 30% or fewer of their AI experiments to scale, and measured concentration is how a fund beats that rate.

Is per workflow measurement just RPA with a new name?

Per workflow measurement works for RPA and for AI agents alike, which is the point. The method cares whether a workflow returned hours after review and rework, not which technology did it. What changes with agents is the error term and the review term, because agents handle exceptions rules could not and make different mistakes.

What if a portfolio company cannot produce timestamps?

A company that cannot export trigger to close timestamps for its top workflows is not ready for an AI program. It is ready for a systems project first. Missing timestamps are information for the readiness score, and they also mean the company's own CFO cannot yet show hours returned to anyone.

How often should the comparison table go to the partner meeting?

The comparison table should go to the partner meeting quarterly, with the capacity ratio per workflow per company and the change since the previous quarter. A workflow counts as scaled when its capacity ratio has held for two consecutive quarters, which separates measured results from a single good month.

How does this differ from measuring AI spend across the portfolio?

Spend governance tracks what each company is paying for and whether it overlaps; per workflow measurement tracks what each workflow returned. BCG's September 2026 Applied AI Index found AI spend near 3.3% of revenue with over 80% outside IT budgets, so a fund needs both registers to see cost and outcome side by side.

Why does the investment committee reject tool level ROI?

Tool level ROI has no baseline, so the committee cannot see a delta. IBM's 2025 CEO study found only 25% of AI initiatives delivered their expected return; a committee that has seen that pattern asks what the workflow cost before, and usage numbers cannot answer the question.

What is the first step for an operating partner starting this measurement?

The first step is to baseline the top five workflows at each portfolio company from system timestamps before approving any new tool. That single rule produces the before state the formula needs, surfaces which companies lack usable systems, and gives the next partner meeting a table instead of a list of pilots.

Your AI Transformation Partner.

Your AI Transformation Partner.

© 2026 Assembly, Inc.