Human in the loop AI agents SOX controls need four layers: a reviewer map, thresholds, an audit trail, and a duty split. See which steps keep your reviewer.
Published
Last Modified
Topic
AI Governance
Author
Jill Davis, Content Writer

TLDR: Human in the loop AI agents SOX controls are the set of reviewer checkpoints, approval thresholds, audit trail fields, and duty separations that let an AI agent act inside a financial workflow without breaking internal control over financial reporting. The agent prepares, a named person approves anything above a threshold, every action is logged in a form an auditor can test, and the agent never holds two duties a person could not hold together. The pattern has four layers, and each one has to be designed before the agent touches a ledger-relevant step.
Best For: AI Transformation Leads, Heads of AI, and finance leaders carrying the AI mandate at enterprises of 1,000 to 15,000 employees in traditional industries, whose first agent worked and whose next use case touches a SOX-relevant step.
Human in the loop AI agents SOX controls are a control design that keeps a named person accountable for every financial decision an AI agent influences, while letting the agent do the preparation work that used to take the team's time. The design answers three questions the controller and the auditor will ask before an agent expands from reading and recommending to posting and releasing: which steps keep a reviewer, what the reviewer sees when they approve, and what evidence exists afterwards. Plenty of enterprises get this far with a working agent in accounts payable or procurement and then stall. The vendor's answer to "how does a human stay in the loop" is usually an approval toggle. An approval toggle is not a control. This post lays out the four layers a real control design needs, in the order you build them.
Why Do Human in the Loop AI Agents SOX Controls Matter More Now?
Human in the loop AI agents SOX controls matter now because agents are moving from suggestion to action faster than control designs are being written. In a September 2026 EY survey of senior AI decision makers at large US companies, 85% said their autonomous systems execute actions without real-time human involvement, and 49% said their governance framework had not been updated to cover agents at all.
The confidence gap is the part that should worry a transformation lead. Deloitte's Q2 2026 CFO Signals survey found 96% of CFOs are confident in their AI governance framework, yet 43% of the same group admit they lack visibility into which AI tools are in use. A framework nobody can see into is a policy document, not a control. Grant Thornton's 2026 AI Impact Survey puts a harder number on it: 78% of 950 senior leaders lack strong confidence they could pass an independent AI governance audit within 90 days, while 70% are already giving agents access to data and processes.
The auditor is arriving before the control design does
Internal audit is already treating agents as a fraud and control surface. The IIA and AuditBoard Q4 2025 survey of 373 North American audit leaders found only 40% believe their function is prepared to detect AI-enabled fraud, and 65% name fabricated invoices and financial documents as a primary threat. An AI agent that codes and releases invoices sits directly on that threat. If you have not handed internal audit a control design, they will write one for you, and it will be more restrictive than the one you would have written.
SOX effort was already growing before agents arrived
The SOX workload the agent lands in is not slack. Protiviti's SOX compliance survey of more than 560 audit and finance leaders found 58% reported an increase in SOX hours in the prior year, and 74% were pursuing further automation to cope. A control design that adds three manual review steps to every agent action will be rejected by the same team it is meant to protect. The pattern below is built to add reviewers only where they change the risk.
What Is a Human in the Loop AI Agent Under SOX, and What Is It Not?
A human in the loop AI agent under SOX is an agent that performs preparation, matching, drafting, or routing steps inside a financial process, while a named person with the right authority performs the approval, release, or posting step, and the whole sequence leaves testable evidence. The agent is a preparer. The person is a reviewer and approver. The control is the combination, not either part alone.
Human in the loop vs. human on the loop vs. human out of the loop
The three terms get used interchangeably in vendor material and mean different control designs. Human in the loop means a person acts before the agent's output takes effect: the agent proposes an invoice coding, a person approves it, then it posts. Human on the loop means the agent acts and a person monitors afterwards, with the ability to stop or reverse: the agent posts low-value recurring invoices and a person reviews an exception report daily. Human out of the loop means the agent acts with no person in the path, which is not a SOX-compatible design for anything that reaches the ledger. The FEI framework for AI in internal control over financial reporting, released June 2026, names human in the loop oversight as one of four control approaches, alongside performance testing, multi-model validation, and data analytics monitoring. The point of naming them separately is that each one is a different control, evidenced differently.
Design | Who acts first | Where the control sits | SOX suitability |
|---|---|---|---|
Human in the loop | Agent prepares, person approves before effect | Pre-action approval with evidence | Suitable for posting, release, master data change |
Human on the loop | Agent acts, person monitors and can reverse | Post-action review and exception handling | Suitable for low-value, high-volume, reversible steps within a threshold |
Human out of the loop | Agent acts, nobody in the path | None at the transaction level | Not suitable for ledger-relevant steps |
How thinking on this has moved
Two years ago the question was whether AI belonged in finance at all; Gartner predicted in September 2024 that 90% of finance functions would deploy at least one AI tool by 2026. The question has since moved to what the agent is allowed to do. In February 2026 COSO published guidance on internal control over generative AI that splits AI use into eight capability types, including ingestion, transformation, posting, judgment, and monitoring, each with its own control expectations. That taxonomy is useful for one reason: it separates the agent that reads from the agent that posts, and the control design should do the same.
How Do You Design Human in the Loop AI Agents SOX Controls? The 4-Layer Pattern
Human in the loop AI agents SOX controls are designed in four layers, in this order: the reviewer map (which steps keep a person), the threshold structure (when the person must act before the effect), the audit trail (what gets recorded for every agent action), and the duty split (which combinations the agent may never hold). Assembly calls this the 4-Layer Control Pattern, and the order matters because each layer constrains the next.
Layer 1: The reviewer map
Start by listing every step in the workflow and marking each one with one of three labels: agent only, agent prepares and person approves, person only. The test for "person only" is whether the step involves clinical, legal, or commercial judgment that the company would not want to defend as automated, or whether it is a key control the external auditor already tests. In an accounts payable flow, invoice capture, three-way match, and coding proposal are typically agent only. Coding approval, payment batch release, and vendor master changes are agent prepares and person approves. Approving a new vendor's bank details is person only, because it is the single step most exposed to the fabricated-document fraud the IIA data points at. If the map ends up with a person on every step, you have built a slower version of the current process and should revisit the 4-gate test for agents in legacy stacks before going further.
Layer 2: The threshold structure
Thresholds convert the reviewer map into rules the system can enforce. Each "agent prepares and person approves" step gets three bands. Below the first threshold the agent acts and the person reviews on an exception report (human on the loop). Between the first and second threshold the person approves before the effect (human in the loop). Above the second threshold a second approver is required, matching whatever dual-approval rule already exists for manual transactions. Set the first threshold from your own operating data, not the vendor's default: if you have already baselined where team time goes, you know the value band where 80% of volume sits and the review effort is highest. The Maximor benchmark reported by the Journal of Accountancy in February 2026 found 86% of CFOs had encountered inaccurate or hallucinated data from AI in finance tasks, which is the argument for keeping the low band narrow at first and widening it only when the exception rate has been measured.
Layer 3: The audit trail
An auditor tests a control by pulling a sample and asking for evidence. For an agent action, the evidence has to exist at the moment of the action, not be reconstructed later. Every agent action in a SOX-relevant step should record eight fields: the input document or record the agent used, the rule or model version that produced the output, the proposed output, the confidence or match score, the reviewer identity, the reviewer's decision and any edit, the timestamp of each, and the system of record transaction ID the action created or changed. The FEI framework describes the traditional model plainly: a person reviews, a person approves, and the trail is legible. The eight fields are what makes the trail legible when a person did not do the preparation.
Layer 4: The duty split
Segregation of duties for agents is simpler than it sounds: treat the agent as one user with one role. The agent may prepare or it may approve, never both for the same transaction. The agent may change master data or it may process transactions against that master data, never both. The agent's credentials are its own, not a shared service account borrowed from the ERP integration user, because a shared account collapses the trail from Layer 3. If the vendor's architecture requires the agent to run under a broad system account, that is a finding before it is a deployment.
For a multi-entity company with several ERP instances, intercompany transactions, and an external auditor who tests entity-level controls separately, a generic approach is not enough; the answer is a dedicated control-design build that first maps every SOX-relevant step per entity to one of the three reviewer labels, second sets thresholds from each entity's own transaction distribution rather than a group-wide default, and third produces the audit trail specification and the duty matrix as documents the controller signs before the agent gets write access. Done once, that build takes a few weeks. Skipped, it tends to come back as an audit finding.
What Do Auditors Actually Test in Human in the Loop AI Agent Controls?
Auditors test whether the control operated as designed for the period, whether the evidence supports that, and whether the people in the loop had the authority and independence the design assumed. For agent controls, this means they will sample agent actions, check that the threshold rules were in force, confirm the reviewer was not also the requester, and ask what happened when the agent was wrong.
The three questions to prepare for
The first question is change management: how do you know the agent's logic on the date of the sample was the logic you documented? Model and rule versioning in the audit trail answers it. The second is completeness: how do you know every agent action was logged? The answer is a reconciliation between the agent's action log and the ERP transaction log, run on a schedule. The third is the override: when a reviewer overrode the agent, what happened next? An override rate that nobody reviews is a control that nobody monitors. Gartner's June 2025 forecast that more than 40% of agentic AI projects will be cancelled by the end of 2027 names inadequate risk controls as one of three causes; the override question is where that inadequacy usually shows up first.
What a good exception rate looks like
There is no published benchmark, so set one. A useful starting rule from Assembly's work with finance teams running agents in AP: if reviewers change more than one in ten agent proposals in the middle band, the agent is not ready for a wider low band, and if they change fewer than one in fifty, the middle band is probably too wide and reviewers are rubber-stamping. Both numbers should be reported monthly to the controller. Grant Thornton's 2026 data found only 20% of organizations have a tested AI incident response plan; an exception-rate report with defined triggers is the cheapest version of one.
What Skeptics Get Wrong About Human in the Loop Agent Controls
The most common objections come from the controller, the CFO, and the operations team, and each one has a specific answer rather than a reassurance.
"If a person approves everything anyway, what did the agent save?" The agent saves the preparation time, which is where most of the hours were. The reviewer map only puts a person on the approval step, and the threshold structure takes the person off the lowest band once the exception rate earns it. McKinsey's 2026 State of AI survey found nearly three quarters of high-performing companies redesigned workflows because of AI, against 25% of everyone else; the redesign is what moves the hours, and the controls are what let the redesign reach the ledger.
"Our people will not trust it, so they will re-check everything." They will at first, and the audit trail lets you show them the exception rate. Protiviti's 2025 study of AI agent adoption found only 17% of mid-level and operational staff expect full autonomy, against 31% of C-suite executives. The gap closes with data, not with a mandate.
"We are not even sure this counts as a SOX control." If the agent's output feeds a balance in the financial statements, it is in scope. L.E.K.'s 2025 Office of the CFO survey found only 11% of CFOs currently use AI inside the finance function, with 35% still piloting; most companies are designing this control for the first time, and the auditor knows it. Bringing them the design early is faster than defending an undocumented one later.
How Should a Transformation Lead Sequence the Control Design?
A transformation lead should sequence the control design so that the reviewer map is agreed with the controller before any threshold is set, the audit trail specification is agreed with internal audit before the vendor builds anything, and the duty matrix is signed before the agent receives write credentials. The sequence takes four to six weeks alongside the existing pilot and is the difference between an agent that expands and one that stays a demo.
Run it as a short project with three sign-offs rather than a governance committee. The decision-centric approach to governing AI agents works well here because every layer of the pattern is a decision with an owner. Where the agent needs to write back into an ERP, the layered architecture for integrating AI with legacy ERP systems covers the technical side; the control pattern covers the accountability side, and both need to be in place before the write-back is switched on. Gartner's finance research also found fewer than 10% of finance functions expect AI-driven headcount reduction, which is worth saying out loud to the team that will do the reviewing. The pattern moves their hours into review work, and the data says the jobs mostly stay.
This analysis was developed using methodologies and operating experience from Assembly.
Frequently Asked Questions
What are human in the loop AI agents SOX controls?
Human in the loop AI agents SOX controls are the reviewer checkpoints, approval thresholds, audit trail fields, and duty separations that let an AI agent prepare work inside a financial process while a named person approves anything that reaches the ledger. The agent prepares, the person approves, and every action leaves evidence an auditor can sample and test.
Which steps must keep a human reviewer when an AI agent runs a financial workflow?
Approval, release, and master data changes must keep a human reviewer. In accounts payable that means coding approval, payment batch release, and vendor master edits, with new vendor bank details kept as person-only steps. Capture, matching, and coding proposals can be agent-only steps because they prepare a decision rather than make one.
What is the difference between human in the loop and human on the loop?
Human in the loop means a person acts before the agent's output takes effect; human on the loop means the agent acts and a person monitors afterwards with the ability to reverse. The FEI June 2026 framework treats human in the loop oversight as one distinct control approach among four, each evidenced differently.
What is the 4-Layer Control Pattern for AI agents?
The 4-Layer Control Pattern is a build sequence for agent controls: a reviewer map that labels each step agent-only, agent-prepares-person-approves, or person-only; a threshold structure with three value bands; an eight-field audit trail recorded at the moment of action; and a duty split that treats the agent as one user with one role.
How do you set approval thresholds for AI agent actions?
Set three bands per approval step using your own transaction data. Below the first threshold the agent acts and a person reviews exceptions; between thresholds a person approves before the effect; above the second threshold a second approver is required. Start with a narrow low band and widen it only after the measured exception rate supports it.
What should an audit trail for an AI agent contain?
An agent audit trail should record eight fields at the moment of each action: the input record used, the rule or model version, the proposed output, the match or confidence score, the reviewer identity, the reviewer decision and any edit, timestamps for each step, and the system of record transaction ID. Reconstructed logs do not satisfy an auditor's sample.
How does segregation of duties apply to an AI agent?
Treat the agent as one user with one role. The agent may prepare or approve a transaction, never both, and may change master data or process transactions against it, never both. The agent needs its own credentials; running it under a shared ERP integration account collapses the audit trail and is a finding before it is a deployment.
Why do auditors care about AI agents in finance processes?
Auditors care because agent output feeds financial statement balances. The IIA and AuditBoard Q4 2025 survey found only 40% of internal audit leaders feel prepared to detect AI-enabled fraud, and 65% name fabricated invoices as a primary threat, which puts an invoice-processing agent squarely in scope.
What do auditors test in a human in the loop agent control?
Auditors test that the control operated as designed for the whole period. They sample agent actions, confirm the threshold rules were in force on the sample date, check the reviewer was independent of the requester, verify the agent's logic version matches documentation, and ask what happened after each reviewer override. Version history and a log reconciliation answer most of it.
What is an acceptable exception rate for an AI agent in accounts payable?
A working rule from practice is between one in fifty and one in ten. If reviewers change more than 10% of agent proposals in the middle band, the agent is not ready for a wider low band. If they change fewer than 2%, the band is probably too wide and reviewers are rubber-stamping. Report both to the controller monthly.
How is SOX compliance effort affected by AI agents?
AI agents can reduce SOX effort only if reviewers are placed where they change risk. Protiviti's SOX survey found 58% of companies reported rising SOX hours and 74% were pursuing automation to cope. A control design that adds manual review to every agent step will be rejected by the team it protects.
Why do agent governance frameworks fail in practice?
Agent governance frameworks fail when they are policies without visibility or enforcement. EY's September 2026 survey found 98% of large companies claim formal AI governance policies while 47% admit skipping their own process for urgent deployments and 36% suffered a materially negative AI incident. Thresholds enforced in the system outlast policies.
How long does it take to design SOX controls for an AI agent?
Designing the control pattern takes four to six weeks alongside an existing pilot. The reviewer map is agreed with the controller first, the audit trail specification with internal audit second, and the duty matrix is signed before the agent receives write credentials. The work runs as a short project with three sign-offs, not a standing committee.
What is the first practical step before expanding an agent to a ledger-relevant step?
The first step is the reviewer map. List every step in the target workflow and label each one agent-only, agent-prepares-person-approves, or person-only, then walk the controller through it. If the map puts a person on every step, the process needs redesign before the agent is expanded, not more automation.
What does COSO's 2026 guidance say about AI in internal control?
COSO's February 2026 guidance splits generative AI use into eight capability types, including ingestion, transformation, posting, judgment, and monitoring, each with tailored control expectations, according to the Journal of Accountancy. The practical use is separating the agent that reads from the agent that posts and controlling them differently.
When does an external partner add value to AI agent control design?
An external partner adds value when the company has several entities, several ERP instances, or an auditor who tests entity-level controls separately. In that situation each entity needs its own reviewer map and thresholds set from its own transaction distribution, and the partner's job is to produce the audit trail specification and duty matrix the controller can sign.
Legal
