An AI results board presentation operations leaders can defend has four parts: baseline, delta, capacity effect, and what to fund next. See what your CFO asks.
Published
Last Modified
Topic
AI Use Cases
Author
Amanda Miller, Content Writer

TLDR: An AI results board presentation operations leaders can defend has four parts, read in this order: the baseline the agent was measured against, the delta it produced, the capacity effect on the team, and what to fund next. Most AI board updates fail because they lead with usage and anecdotes, which a CFO cannot underwrite. The evidence pack replaces them with operating data from one live use case.
Best For: Heads of Transformation, Heads of AI, CAIOs, and COO or finance leaders carrying the AI mandate at enterprises of 1,000 to 15,000 employees in manufacturing, distribution, healthcare services, hospitality, or insurance, with one AI agent live and a board pre-read due.
An AI results board presentation operations leaders can defend is a short, evidence-first pre-read that shows what one live AI use case changed in a workflow, measured against a baseline the finance team accepts. It is different from an AI strategy deck, which describes intent, and from a vendor dashboard, which describes activity. The pack exists because the board question has changed. Two years ago the question was "what are we doing with AI." In 2026 it is "what did the AI budget return," and the person who has to answer it is usually the transformation lead who ran the first pilot in their spare time. This post lays out the four parts that question needs, the fields inside each, and the order a board reads them in.
Why the AI results board presentation operations leaders usually bring gets filed
An AI results board presentation operations leaders bring to the pre-read usually fails for one reason: it reports activity instead of change. Logins, prompts per week, and documents processed tell the board the agent is running. They do not tell the board what the agent changed in cycle time or error rate, and a CFO will not sign off on a second budget line based on activity.
The numbers behind that failure are now public and consistent. In the PwC 2026 Global CEO Survey of 4,454 chief executives, 56% said AI had produced neither a revenue increase nor a cost decrease, and only 12% reported both. IBM's 2025 CEO study of 2,000 CEOs found that only 25% of AI initiatives had delivered the return they expected, and just 16% had scaled across the enterprise. Those are the people your board talks to at dinner. They walk in skeptical.
Usage is not value, and the board knows it
The mistake is easy to make. Usage data is what the vendor's admin console gives you, so that is what ends up on the slide, usually the night before. But Gartner's July 2026 CFO survey found that 45% of CFOs describe their AI investments as leaning toward productivity, and only 20% toward decision quality. A productivity claim without a time baseline is an assertion. The CFO has learned to treat it that way. In Gartner's May 2026 release, 63% of finance organizations said their AI implementation ran slower than expected in 2025, which is the other thing the CFO remembers when your slide comes up.
The vendor slide problem
The second failure is borrowing the vendor's ROI. A per-transaction saving multiplied by projected volume looks like a result. It is a forecast, built on the vendor's assumptions about your process, and the board has seen enough of these that the format itself now reads as a warning. Forrester's 2026 predictions put it plainly: fewer than one-third of decision-makers can tie AI value to their organization's financial growth, and Forrester expects enterprises to defer a quarter of planned AI spend into 2027 as a result. Deferral is what a board does when the evidence is somebody else's.
A short history of the board's AI question
The question has moved three times. In 2023 boards asked for an AI policy and a risk view. In 2024 and 2025 they asked for a pipeline of use cases and a pilot. In 2026 they ask what the budget returned, and most boards are not equipped to evaluate the answer. A Deloitte survey published in July 2026 found that 51% of boards still lack rules and guidance on AI use in their own work. NACD's 2025 private company survey found that only 15% of boards receive management reporting metrics on AI and only 17% have approved an AI project budget. You are often the first person to put a measured AI result in front of them. That is an advantage if the format is right, because the format you bring becomes the standard they expect.
What an AI results board presentation operations teams can defend contains: the 4-Part Evidence Pack
An AI results board presentation operations teams can defend contains four parts in a fixed order: Baseline, Delta, Capacity effect, and What to fund next. Each part is one page. Each has named fields, a named data source, and a named owner who will stand behind the number if the CFO calls them. The pack is built from one live use case, never from a portfolio of pilots.
The order matters more than people expect. Boards read baseline before delta because a delta without a baseline is a claim. They read capacity effect before the funding ask because the ask is only credible if the capacity freed is already visible somewhere. Here are the four parts with their required fields.
Part | What it answers | Required fields | Data source | Owner |
|---|---|---|---|---|
1. Baseline | What did this workflow cost in time and errors before the agent | Volume per month, touch time per item, handoffs per item, exception rate, rework rate, measurement window | System logs and ticket timestamps, not surveys | Process owner |
2. Delta | What changed, by how much, over what period | Same fields as baseline after go-live, side by side, with the number of weeks observed and the share of volume the agent handled | Same sources as the baseline | Transformation lead |
3. Capacity effect | Where did the freed time go | Hours freed per month, what the team did with them, backlog reduction, overtime change, headcount decisions taken or deferred | Time allocation record and manager confirmation | Functional leader |
4. What to fund next | What the board is being asked to approve and on what evidence | Next use case, its baseline (already measured), expected delta range, the exit criterion that stops funding, the control design | Use case pipeline and the operating model | Transformation lead with CFO sign-off |
1. Baseline: the page the CFO reads first
The baseline is the page that decides whether the rest of the pack is read or filed. It states what the workflow consumed before the agent went live, in units the finance team already uses: items per month, minutes per item, handoffs, exceptions. The baseline must come from system data. A survey of the team is not a baseline, and the board will ask how it was collected. If the pilot went live without one, the method in how to baseline team time before deploying AI can reconstruct it from timestamps for the eight weeks before go-live. State the measurement window on the page. A baseline measured over two weeks in a slow month is a weak baseline, and it is better to say so than to have the CFO find out.
2. Delta: the same fields, after go-live
The delta page repeats every baseline field with the post go-live value next to it. Nothing else. No percentages without the underlying counts, and no annualized projections. The one addition is coverage: what share of total volume the agent actually handled, because an agent that processed 30% of invoices and cut touch time on those by half has changed the workflow less than the headline suggests. McKinsey's 2025 State of AI survey found that only 39% of companies report any EBIT impact from AI at the enterprise level, and the ones that do are the ones that redesigned the workflow rather than adding a tool to it. The delta page is where the board sees whether you did that.
3. Capacity effect: the page most packs leave out
Freed hours are the number boards want and the number most presentations skip, because it requires saying what happened to the people. The capacity page states hours freed per month, where they went (backlog, close acceleration, deferred hiring, overtime reduction), and which decisions were taken as a result. This page is also where the human-in-the-loop design shows up: how many items went to a reviewer, how long review took, and what the reviewer changed. A board in a SOX environment will ask. In the pre-reads Assembly has reviewed from operations leaders at enterprises in traditional industries, the capacity page is the one most often missing, and its absence is the most common reason a funding ask is deferred to the next quarter rather than declined. The CFO does not doubt the agent works. The CFO doubts the hours went anywhere.
4. What to fund next: an ask with an exit criterion
The final page is the funding ask, and it is credible only because of the three pages before it. It names the next use case, shows that its baseline has already been measured, gives an expected delta as a range, and states the exit criterion: the measured result at which the board should stop funding. Boards approve more readily when they are told how to stop. Gartner predicts that over 40% of agentic AI projects will be canceled by the end of 2027 for escalating cost, unclear value, or inadequate controls. A pack with a written exit criterion tells the board you are running the cancellation decision yourself rather than waiting for them to run it for you.
Evidence pack vs. AI dashboard vs. business case
An AI evidence pack is a backward-looking record of measured change from one live use case; an AI dashboard is a live view of activity; a business case is a forward-looking forecast. Boards need the first to trust the third, and the second is what most operations leaders bring by mistake. The three are often confused because they share metrics, but they answer different questions and carry different weight with a CFO.
Evidence pack | AI dashboard | Business case | |
|---|---|---|---|
Time direction | Backward: what happened | Present: what is running | Forward: what will happen |
Core unit | Measured delta against a baseline | Activity counts | Projected savings |
Source of numbers | Your system logs | Vendor console | Vendor model plus your assumptions |
What the CFO can do with it | Underwrite the next ask | Confirm the tool is used | File it until there is evidence |
Failure mode | Weak baseline | Mistaken for results | Never reconciled to actuals |
When to bring it | After one live use case | Monthly ops review, not the board | After the pack, as the funding ask |
KPMG's Q2 2026 AI pulse of 2,145 senior leaders shows why the distinction is now material. 49% of organizations had scaled back, narrowed, delayed, or paused AI agent deployments because cost ran ahead of value, and only 7% had reached established ROI. The organizations with full cost visibility reached established ROI at 15%, against 3% for those without. Visibility is what an evidence pack provides and a dashboard does not.
How to build the pack in the three weeks before the pre-read
Building an AI results evidence pack in three weeks follows a fixed sequence: reconstruct the baseline from system data in week one, pull the post go-live values for the same fields and confirm the capacity story with the functional leader in week two, and write the funding page with the CFO in week three. The sequence is strict because each page depends on the one before it, and because the CFO must see the funding page before the board does.
Week one: baseline from timestamps, not memory
Pull ticket, queue, or ERP timestamps for the eight weeks before go-live. Compute volume, touch time, handoffs, and exception rate per item. If the pilot ran without a baseline, this is reconstruction, and the pack should say so. A reconstructed baseline from system data is still stronger than a survey. The MIT NANDA report from August 2025 found that 95% of enterprise AI pilots produced no measurable P&L impact, and the report's own explanation was that most were never measured against the workflow they were meant to change. Week one is where you leave that group.
Week two: the delta and the capacity conversation
Run the same fields for the weeks since go-live. Then sit with the functional leader whose team the agent serves and record where the freed hours went. This conversation is the most important hour in the three weeks, and it should produce a statement the functional leader will repeat in the boardroom if asked. If the leader cannot say where the hours went, the capacity page says "not yet redeployed," and the funding page shrinks accordingly. Boards remember the honesty longer than they remember the number.
For a multi-entity company where the live agent posts into the ERP under SOX controls and the board is being asked to fund the next three use cases, a generic approach is not enough; the answer is a dedicated measurement program that does three things: it fixes one measurement standard across entities so every use case reports the same baseline and delta fields, it records reviewer interventions as part of the delta so the control design is evidence rather than a diagram, and it assigns one owner for the pack who is not the vendor. Without that program, each pre-read is rebuilt from scratch and the numbers drift between quarters. The lean version of that program is described in what an AI operating model PMO setup looks like.
Week three: write the funding page with the CFO, not for the CFO
Draft the funding page, then review it with the CFO before it goes into the pre-read. The CFO will push on the baseline window, on coverage, and on the exit criterion. Every push you absorb in that meeting is one you do not absorb in the boardroom. BCG's September 2026 Applied AI Index reports that AI spending nearly doubled in a year, from 1.7% of revenue in late 2025 to 3.3% in 2026, and that over 80% of it now sits outside the IT budget. Your CFO is being asked to find that money across functional lines. A funding page the CFO has already signed is the only kind that survives.
The hardest questions the CFO will ask, and what to say
The hardest questions a CFO asks about an AI results board presentation are the ones about method: how the baseline was measured, how much of the volume the agent really handled, and where the freed hours went. The answers are in the pack if it was built in the sequence above. Here are the three that come up most, and the honest response to each.
"Your baseline was measured after the fact. How do I know it is real?" You know because it came from timestamps, the extraction method is stated on the page, and the process owner will walk you through the query. A reconstructed system baseline is weaker than a planned one and the pack says so. It is still the strongest number in the room.
"The savings are hours. Where is the money?" The money is on the capacity page, and it is whatever the functional leader did with the hours: a close that finished two days sooner, a hiring requisition withdrawn, overtime that stopped. If the hours were not redeployed, the page says so, and that is the reason the next use case is sequenced where it is. Accenture's March 2026 research found that 86% of companies plan to increase AI investment while only 18% are generating value from it. The gap is unredeployed capacity. Naming it is how you show you are not in that group.
"Why should the board fund a second use case before the first has paid back?" Because the first already has a measured delta and a stated exit criterion, the second has a measured baseline, and the control design from the first carries over. The board is not being asked to believe a forecast. It is being asked to extend a method that has already produced one number it can check. If that argument does not hold for your next use case, the pack should say the next use case is not ready, and the funding page should ask for measurement time instead. The measurement gaps that usually sit behind this question are laid out in how to prove AI ROI to the board, and the metric set that makes the delta page comparable across quarters is in how to measure AI ROI in enterprise operations.
This analysis was developed using methodologies and operating experience from Assembly.
Frequently Asked Questions
What is an AI results board presentation for operations?
An AI results board presentation for operations is a short pre-read that shows what one live AI use case changed in a workflow, measured against a baseline the finance team accepts. It has four parts: baseline, delta, capacity effect, and what to fund next. It differs from a strategy deck, which states intent, and from a vendor dashboard.
What should an AI results board presentation contain?
An AI results board presentation should contain four one-page parts in a fixed order: the baseline the workflow was measured against, the delta after go-live using the same fields, the capacity effect showing where freed hours went, and a funding ask with a written exit criterion. Each page names its data source and the owner who will defend the number.
Why do most AI board updates get filed without a decision?
Most AI board updates get filed because they report activity rather than change. Logins, prompts, and documents processed show the agent is running, not what it altered. The PwC 2026 CEO Survey found 56% of CEOs saw neither revenue gains nor cost reductions from AI, so boards now expect measured deltas, not usage counts.
What is the 4-Part Evidence Pack?
The 4-Part Evidence Pack is the format for an operator's AI board pre-read: Baseline, Delta, Capacity effect, and What to fund next, each one page with named fields, a data source, and an owner. It is built from one live use case, never a pilot portfolio, and read in that order because each page depends on the one before it.
How do you present AI results to the board without a baseline?
You reconstruct the baseline from system timestamps for the eight weeks before go-live, compute volume, touch time, handoffs, and exception rate per item, and state on the page that it was reconstructed. A reconstructed system baseline is weaker than a planned one but far stronger than a team survey, and the board should be told which it is.
What metrics do boards want to see for AI?
Boards want measured change in operating metrics: items per month, touch time per item, handoffs, exception and rework rates, before and after, plus the share of volume the agent handled and the hours freed. NACD's 2025 survey found only 15% of boards receive AI management reporting metrics, so your format often becomes the standard.
What is the difference between an AI evidence pack and an AI dashboard?
An AI evidence pack is a backward-looking record of measured change against a baseline from one live use case, while an AI dashboard is a live view of activity from the vendor console. The dashboard belongs in a monthly operations review. The pack is what the board needs to underwrite a funding ask, because it reconciles to your system data.
What is the capacity effect page and why does it matter?
The capacity effect page states hours freed per month and where they went: backlog reduction, close acceleration, deferred hiring, or overtime reduction. It is the page most packs omit because it requires saying what happened to the people. In the pre-reads Assembly has reviewed, its absence is the most common reason a funding ask is deferred rather than declined.
How should the funding ask be structured in an AI board presentation?
The funding ask should name the next use case, show its baseline is already measured, give an expected delta as a range, and state an exit criterion. The exit criterion is the measured result at which the board stops funding. Gartner expects over 40% of agentic AI projects to be canceled by 2027, so a written stop rule builds trust.
How long should the AI section of a board pre-read be?
The AI section of a board pre-read should be four pages, one per part of the evidence pack, plus an optional appendix with the baseline extraction method. Boards read baseline before delta and capacity before the funding ask. Anything longer gets skimmed, and anything shorter forces the board to take the delta on faith because the baseline is missing.
How long does it take to build an AI evidence pack?
Building an AI evidence pack takes about three weeks when system data exists: week one reconstructs the baseline from timestamps, week two pulls the post go-live values and confirms the capacity story with the functional leader, and week three writes the funding page with the CFO. The sequence is strict because each page depends on the one before it.
Who should own the AI results board presentation?
The transformation lead should own the pack, with named owners for each page: the process owner signs the baseline, the functional leader signs the capacity effect, and the CFO reviews the funding page before it reaches the board. IBM's 2025 CEO study found only 25% of AI initiatives met expected ROI, and unowned numbers are a large part of why.
How do you handle human-in-the-loop and SOX questions in the board pack?
Human-in-the-loop and SOX questions are answered on the capacity page by reporting how many items went to a reviewer, how long review took, and what the reviewer changed. Recording interventions as part of the delta turns the control design into evidence rather than a diagram, which is what an audit committee needs before it approves ERP write-back.
What do you do when the first AI use case has not paid back yet?
You present the measured delta and exit criterion for the first use case and ask for method extension, not belief in a forecast. If the next use case lacks a measured baseline, the funding page asks for measurement time instead of build budget. KPMG's Q2 2026 pulse found 49% of organizations paused or narrowed agent rollouts when cost outran value.
Should vendor ROI figures appear in an AI board presentation?
Vendor ROI figures should not appear in the evidence pack. A per-transaction saving multiplied by projected volume is a forecast built on the vendor's assumptions, and boards now read the format as a warning. Forrester's 2026 predictions found fewer than one-third of decision-makers can tie AI value to financial growth; your own system data is what closes that gap.
When does an enterprise need a dedicated AI measurement program instead of a one-off pack?
An enterprise needs a dedicated AI measurement program when agents post into the ERP under SOX controls across multiple entities and the board is funding several use cases. The program fixes one measurement standard, records reviewer interventions as part of the delta, and assigns a pack owner who is not the vendor, so each quarter's pre-read reconciles to the last.
Legal
