Where to Start With AI Agents in a Traditional Industry? A 3-Layer Map for Operations Leaders

Where to Start With AI Agents in a Traditional Industry? A 3-Layer Map for Operations Leaders

Where to start with AI agents in a traditional industry: intake work first. Your first workflow needs a queue, a system of record, and a baseline. See the map.

Published

Last Modified

Topic

AI Use Cases

Author

Amanda Miller, Content Writer

TLDR: The answer to where to start with AI agents in a traditional industry is the high-volume back-office intake work: invoices, purchase requests, and contract intake. These workflows have a countable queue, a system of record, and a baseline you can measure in weeks. Customer-facing and judgment-heavy work come later, once the operating model behind the first agent has been tested.

Best For: Heads of Transformation, Heads of AI, Chief AI Officers, and COOs or finance leaders carrying the AI mandate at enterprises in traditional industries (1,000 to 15,000 employees) who have been asked to name the first workflow.

An AI agent in operations is a piece of software that takes a work item from a queue, reads the documents attached to it, decides what should happen next under rules a person set, and hands the result to a system of record or to a reviewer. That definition matters because it tells you where an agent can and cannot start. It needs a queue, it needs documents, it needs rules, and it needs somewhere to write the result. In a manufacturing, distribution, insurance, or healthcare services enterprise, the places that have all four today are the back office: accounts payable, procurement intake, and contract intake. The customer-facing and judgment-heavy work most executives picture when they say "AI" usually has none of the four in a usable state. This post lays out the starting map, the tests for choosing the first workflow, and the reasons the exciting work comes second.

Why the First Workflow Decides Whether the Program Survives

The first workflow decides the program's fate because it produces the only evidence the CFO will accept: a measured before-and-after on one process. A first agent placed in a workflow with no queue and no baseline produces a demo. A first agent placed in high-volume intake produces a number, and that number is what funds the second workflow.

The failure rate is high enough that the choice is not academic. MIT's NANDA initiative reported in 2025 that 95% of enterprise AI pilots delivered no measurable return, and that budgets flowed to sales and marketing even though the returns it could document came from operations and finance. IBM's 2025 CEO study found that only 25% of AI initiatives had delivered the expected return, only 16% had scaled across the enterprise, and 64% of CEOs admitted they were investing in some technologies before they understood the value. Three surveys, one mistake: starting where the story is good rather than where the measurement is possible.

The evidence gap is a sequencing problem

McKinsey's most recent State of AI survey found that nearly nine in ten organizations use AI somewhere, but only 37% can attribute any earnings impact to it, and only 6% qualify as high performers. Nearly three quarters of those high performers had redesigned a workflow before automating it, against a quarter of everyone else. The software is rarely the gap. The gap is whether anyone fixed the process the software was dropped into, and that is a great deal easier in a workflow with a defined front door than in one where work arrives by phone call and hallway conversation.

Why traditional industries feel this harder

A 3,000-person distributor or a regional hospital group runs on an ERP that predates the current CFO, a document management system nobody fully owns, and a dozen spreadsheets that hold the real process. IBM's survey found that 50% of CEOs say recent investment has left them with disconnected, piecemeal technology. An agent has to live inside that reality. The back office is the one place where the piecemeal stack still has a clear front door. Every invoice, requisition, and contract has to enter somewhere, or the company cannot pay, buy, or sign.

Where to Start With AI Agents in a Traditional Industry: The 3-Layer Map

Where to start with AI agents in a traditional industry is Layer 1: high-volume back-office intake, meaning accounts payable, procurement requests, and contract intake. Layer 2 is internal exception handling with a reviewer in the loop. Layer 3 is customer-facing and judgment-heavy work. Each layer depends on the operating habits built in the one before it.

The map below is the honest version, which is to say the unglamorous one. It is ordered by how much of the work an agent can do reliably today and how quickly you can prove it.

Layer

Workflows

What the agent does

What has to exist first

When to start

1. Back-office intake

Invoice intake and matching, purchase request intake, contract intake and clause extraction, vendor onboarding

Reads the document, checks it against the PO, contract, or policy, routes the exception, drafts the entry for a person to post

A queue, a system of record, an owner, a time baseline

Now. First 90 days

2. Internal exception handling

Invoice exceptions, credit holds, claims triage, order changes, shared services tickets

Proposes a resolution with the evidence attached; a reviewer approves or edits

Layer 1 running, an exception taxonomy, approval thresholds written down

After the first Layer 1 number is in

3. Customer-facing and judgment-heavy

Customer service resolution, pricing decisions, clinical or underwriting judgment, supplier negotiation

Acts with the customer or makes a decision with real financial or legal weight

Clean master data, a tested control design, a workforce that trusts the first two layers

After a year of operating discipline, and only for defined, bounded cases

Layer 1: high-volume intake

Accounts payable is the canonical starting point because the queue is visible and the waste is measured. Ardent Partners' 2025 AP metrics put the average invoice at 9.2 days to process against 3.1 days for the best performers, an average exception rate of 14% against 9% for the best, and only 32.6% of invoices processed with no human touch. AP staff spend 21.8% of their time answering supplier inquiries about where an invoice is. Each of those numbers is a baseline your own team can pull from the ERP in a week. Try doing that for "customer experience."

Procurement intake is the second seat. The Hackett Group's 2025 key issues study found procurement workloads rising 10% against budgets rising 1%, with purchase order processing and spend analytics as the top AI use cases; 49% of procurement teams had piloted AI but only 4% had deployed it at scale. Contract intake is the third. World Commerce & Contracting has found for over a decade that poor contract management costs companies about 9% of their bottom line, most of it from obligations nobody tracked after signature. Reading a contract at intake and pulling out the dates, thresholds, and obligations is the kind of bounded, document-heavy task an agent handles well.

Layer 2: exceptions with a reviewer

Once intake is running, the exceptions it routes become the next queue. The agent already knows why an invoice failed the match; the next step is letting it propose the fix and attach the evidence, while a person keeps the approval. PwC's 2025 AI agent survey found 79% of companies adopting agents, but fewer than half (42%) had redesigned any process around them, and at most companies half or fewer of employees touched an agent in daily work. Layer 2 is where that redesign has to happen, whether anyone planned it or not, because the agent cannot use an exception taxonomy or an approval threshold that lives in someone's head.

Layer 3: customer-facing and judgment work

This is the work the board pictures. It is also the work that fails first when attempted first. Gartner predicts more than 40% of agentic AI projects will be canceled by the end of 2027 for unclear value or inadequate risk controls, and a second Gartner analysis predicts half of the organizations that planned to cut customer service headcount will abandon those plans by 2027, with 95% of service leaders keeping human agents. Customer-facing work needs clean customer master data, a tested control design, and a workforce that has watched the first two layers behave. None of that exists on day one at a traditional enterprise.

How to Decide Where to Start With AI Agents in a Traditional Industry: The 4-Test Sequencing Rule

Deciding where to start with AI agents in a traditional industry comes down to four tests applied to each candidate workflow: a countable queue, a system of record with a place to write, a named owner who can change the process, and a time baseline you can capture before go-live. A workflow that fails any one test is not first, whatever its strategic appeal.

The rule is narrow on purpose. BCG's research on AI value found that the 4% of companies generating consistent value pursued about half as many opportunities as everyone else and scaled more than twice as many, and that 70% of implementation difficulty sat in people and process rather than the technology. Fewer candidates, tested harder. That is the whole pattern.

Test 1: Is there a countable queue?

Count the items that entered the workflow last month. Invoices received, requisitions raised, contracts submitted for review. If the number takes more than a day to find, or if the answer is "it depends who you ask," the workflow has no intake point and the agent has nothing to pull from. A practical floor for a first deployment is a few thousand items a month; below that, the measured saving will be too small to defend, and above it the exceptions alone will justify the work.

Test 2: Is there a system of record with a place to write?

The agent's output has to land somewhere a person will act on it. For AP that is a draft entry in the ERP awaiting posting; for procurement it is a requisition with the policy check attached; for contracts it is a record in the contract system with the extracted terms. If the result would land in an email, the workflow is not ready. Enterprises with a heavy legacy stack often find this is the test that fails, and the fix is usually a small integration rather than a platform replacement; the legacy-heavy enterprise strategy covers the patterns that work.

Test 3: Is there an owner who can change the process?

The agent will surface process defects in the first week: duplicate vendor records, POs raised after the invoice, approval chains that loop. Someone has to be able to fix those without a steering committee. If the workflow crosses three functions and no one owns it end to end, pick a different workflow or narrow this one to the slice one person controls.

Test 4: Can you capture a time baseline before go-live?

The saving is only provable if the "before" was measured. Cycle time per item, touches per item, exception rate, and hours spent on inquiries are the four figures that matter, and they need to come from operating data rather than a survey. A baseline captured before deployment is what turns the first agent into evidence. Without it, the CFO's question in month six has no answer.

For a multi-entity enterprise whose first agent will touch the general ledger under SOX controls across several legal entities, a generic approach is not enough; the answer is a dedicated build that does three things: it maps and fixes the intake step before placing any agent, it keeps a reviewer at every posting threshold with a complete audit trail and segregation of duties preserved, and it measures against a time baseline captured before go-live so the result can be defended to the audit committee as well as the CFO.

The Intake Checklist for the First Workflow

The intake checklist is a one-page record you complete for the chosen workflow before any agent is configured. It fixes the queue, the routing rules, the exceptions, and the write-back point in writing, so that the agent inherits a defined process rather than the informal one that existed before. Completing it usually takes two weeks, and in most AP teams it removes a meaningful share of the exceptions before any agent runs, simply because the informal process finally gets written down.

The checklist has eight lines:

  1. Entry point: the single place work items arrive (mailbox, portal, EDI feed), and what happens to items that arrive anywhere else.

  2. Required fields: what a complete item contains, and what the agent should do when a field is missing.

  3. Reference data: the PO, contract, vendor master, or policy the item is checked against, and who maintains it.

  4. Routing rules: the conditions that send an item to a person, and which person.

  5. Exception taxonomy: the eight to twelve reasons an item fails, named consistently.

  6. Approval thresholds: the amount or risk level above which a reviewer must act, written as numbers.

  7. Write-back point: the exact record and status the agent creates in the system of record.

  8. Baseline: cycle time, touches, exception rate, and inquiry hours for the last full quarter.

Most of the checklist can be filled from a workflow audit of the current process, and the lines that cannot be filled tell you which workflow is not ready. Deloitte's 2026 State of AI in the Enterprise found that only one in five companies has mature governance for autonomous agents; the checklist is the minimum governance for one workflow, and it is enough to start.

Why Customer-Facing and Judgment Work Come Later

Customer-facing and judgment-heavy work come later because they lack the four conditions the sequencing rule tests for, and because their failures are visible to customers, regulators, and auditors rather than to an AP clerk. An agent that mis-routes an invoice costs a day. An agent that misquotes a price or misjudges a claim costs a customer and possibly a compliance finding.

The data is not ready

Customer, product, and pricing master data at a traditional enterprise is typically maintained by several teams with different definitions. KPMG's Q2 2026 AI pulse survey found 53% of large organizations deploying agents, but data readiness was the top barrier cited in banking at 63%, and only 26% of leaders had real-time visibility into what their AI was costing to run. Intake workflows can tolerate imperfect master data because a person still posts the result. Customer-facing work cannot.

The controls are not designed

Judgment work sits under controls that were written for people: SOX segregation of duties, underwriting authority limits, clinical sign-off. Those controls have to be redesigned for an agent before the agent acts, and that design is best learned on Layer 2 exceptions, where the reviewer is already in the loop and the stakes are internal. Bain's enterprise AI survey found data security and privacy (44%), lack of in-house expertise (42%), and accuracy concerns (39%) as the top barriers; every one of those is sharper in customer-facing work.

The workforce has not seen it work

Functional leaders resist agents they have not watched. Layer 1 gives the AP manager, the procurement lead, and the contracts team a year of seeing the agent route, draft, and defer to them. That is the trust the customer service or underwriting leader will need before an agent touches their queue. The multi-site admin workflow analysis shows how the same pattern plays out across entities.

What Skeptics Get Wrong About Starting in the Back Office

The most common objection is that back-office intake is too small to matter, and it is wrong because the intake workflows are where the enterprise's cash, commitments, and obligations enter, and because they are the only place a first agent can be measured against a baseline within a quarter. Three objections come up in nearly every leadership discussion.

"This is just RPA with a new name"

Partly true, and the difference is where the work breaks. Scripted automation handles the 32.6% of invoices that already match and falls over on the rest. An agent reads the unstructured part, the PDF invoice with a different line structure, the contract clause worded three ways, and it handles the exception path that scripted tools route to a person. If your candidate workflow has no unstructured documents and no exceptions, use the cheaper tool.

"The CEO wants something visible"

The CEO wants a number at the next board meeting more than a demo. A visible customer-facing pilot with no baseline produces an anecdote; an AP intake agent produces a cycle time that fell from nine days to three, with the ERP data to prove it. The second is what gets the second workflow funded. If visibility is the requirement, publish the Layer 1 number internally and let it travel.

"We don't have the data for any of this"

You have more than you think for Layer 1 and less than you think for Layer 3. The ERP already holds every invoice, PO, and receipt; the contract system holds every signed agreement. What is missing is usually the exception taxonomy and the baseline, and both can be built in two weeks from records that already exist. The customer and pricing master data problem is real. It argues for starting in intake and cleaning the master data while the first agent runs.

What Comes After the First Number

After the first workflow produces a measured result, the next step is a scored prioritization of the remaining candidates using the four tests as gates and the operating data as inputs, followed by an operating model that gives the program an owner and an intake of its own. That is a different problem from where to start, and it gets its own treatment elsewhere. The first decision is narrower than people expect: pick the intake workflow that passes all four tests, fill in the eight-line checklist, capture the baseline, and place the agent. Everything else in the program is built on that number.

This analysis was developed using methodologies and operating experience from Assembly.

Frequently Asked Questions

Where should a traditional-industry enterprise start with AI agents?

Start with high-volume back-office intake: invoice intake and matching, purchase request intake, and contract intake. These workflows have a countable queue, a system of record, a named owner, and a baseline you can capture from the ERP within a week. Customer-facing and judgment-heavy work lack those conditions and should come after the first measured result.

What is an AI agent in enterprise operations?

An AI agent in operations is software that takes a work item from a queue, reads the attached documents, decides the next step under rules a person defined, and writes the result to a system of record or hands it to a reviewer. It differs from scripted automation by handling unstructured documents and exception paths rather than fixed, structured inputs.

Why is accounts payable the most common first AI agent use case?

Accounts payable has the clearest queue and the best-measured waste. Ardent Partners reports an average of 9.2 days per invoice, a 14% exception rate, and only 32.6% of invoices processed without human touch. Every one of those figures can be pulled from the ERP as a baseline before an agent is deployed.

What are the four tests for choosing the first AI agent workflow?

The four tests are a countable queue, a system of record with a write-back point, a named owner who can change the process, and a time baseline captured before go-live. A candidate workflow that fails any single test should not be first, regardless of how strategic it appears. Fewer candidates tested harder is the pattern that scales.

Why do customer-facing AI agents fail when deployed first?

Customer-facing agents fail first because their data, controls, and workforce trust are not ready. Customer and pricing master data is usually inconsistent, controls were written for people, and functional leaders have not yet watched an agent behave. Gartner predicts half of planned service headcount cuts will be abandoned by 2027.

How many items per month does a workflow need to justify a first agent?

A practical floor is a few thousand work items per month. Below that volume the measured saving is too small to defend to a CFO, and the exceptions are too few to teach the team anything. Above it, the exception handling alone justifies the deployment. Count last month's items first to find out where you stand.

What is an intake checklist for an AI agent deployment?

An intake checklist is a one-page record completed before any agent is configured. It fixes the entry point, required fields, reference data, routing rules, exception taxonomy, approval thresholds, write-back point, and baseline in writing. Completing it typically takes two weeks and removes a share of exceptions on its own, because the informal process gets defined for the first time.

What baseline should be captured before deploying an AI agent?

Capture four figures from operating data for the last full quarter: cycle time per item, touches per item, exception rate, and hours spent on status inquiries. Surveys and estimates do not count. The baseline is what converts the first deployment into evidence a CFO accepts, and it cannot be reconstructed after the agent goes live.

Is an AI agent just RPA with a new name?

No, though the overlap is real. Scripted automation handles items that already match a fixed structure and stops on everything else. An agent reads unstructured documents such as a PDF invoice with unusual line items or a contract clause worded three ways, and works the exception path. With no unstructured documents and no exceptions, scripted automation is the simpler tool.

How long does the first AI agent deployment take in a traditional enterprise?

A first intake agent typically takes about 90 days from workflow selection to a measured result, with roughly two weeks on the intake checklist and baseline, several weeks on configuration and integration to the system of record, and the remainder running in production with a reviewer while the before-and-after figures accumulate for the first review.

What does human-in-the-loop mean for a first AI agent?

Human-in-the-loop means the agent drafts and routes while a person keeps every posting and approval decision. For an invoice agent, the person posts the entry the agent prepared; for contracts, a reviewer confirms the extracted terms. Approval thresholds are written as numbers, and every agent action is logged so segregation of duties and audit requirements are preserved.

How do SOX controls affect where an AI agent can start?

SOX controls favor starting in intake, where the agent prepares work and a person posts it. Segregation of duties stays intact because the agent never approves and posts the same item. Judgment work under authority limits requires the controls to be redesigned for an agent first, which is best learned on internal exceptions before anything touches customers or the ledger.

Why do most enterprise AI pilots fail to show a return?

Most pilots fail to show a return because they start where the story is good rather than where measurement is possible. MIT's NANDA initiative found 95% of enterprise pilots delivered no measurable return in 2025, with budgets flowing to sales and marketing while the documented returns came from operations and finance workflows.

What happens after the first AI agent produces a measured result?

The next step is a scored prioritization of the remaining candidate workflows, using the four tests as gates and operating data as inputs, followed by an operating model with an owner and an intake process for new use cases. BCG found value leaders pursue half as many opportunities and scale twice as many.

Should a traditional enterprise wait for cleaner data before deploying AI agents?

No for intake workflows, yes for customer-facing ones. The ERP already holds every invoice, purchase order, and receipt, and the contract system holds every agreement, which is enough for a Layer 1 agent because a person still posts the result. Customer and pricing master data problems are a reason to start in intake, not a reason to delay.

When is an enterprise ready to move AI agents into judgment-heavy work?

An enterprise is ready when three conditions hold: the intake and exception layers have run for about a year with measured results, the controls governing that judgment have been redesigned for an agent and tested, and the functional leader who owns the decision has watched agents defer correctly. Even then, start with defined, bounded cases rather than open-ended authority.

Your AI Transformation Partner.

Your AI Transformation Partner.

© 2026 Assembly, Inc.