How to baseline team time before deploying AI: use system timestamps, work sampling, and volume counts, never a survey. Get the record your CFO will accept.
Published
Last Modified
Topic
AI Diagnostic
Author
Amanda Miller, Content Writer

TLDR: Knowing how to baseline team time before deploying AI is the difference between a CFO accepting your number and a CFO politely ignoring it. The method that holds up does not ask people how they spend their day and does not put tracking software on their laptops. It combines three sources you already have: timestamps in the systems where the work happens, a short run of work sampling sessions, and volume counts per work type, recorded in an eight-field baseline table before the agent goes live.
Best For: Heads of Transformation, Heads of AI, CAIOs, and COO or finance leaders carrying the AI mandate at enterprises in traditional industries (1,000 to 15,000 employees) who have one live use case and need a before-number for the next one.
A team time baseline is a validated record of where a team's working hours go, by workflow step, before an AI agent touches the process. Most enterprises skip it, then discover the gap the day finance asks what the first agent returned. McKinsey's 2026 State of AI survey found that 88% of organizations now use AI in at least one function while only 37% can attribute any EBIT impact to it, and that the small group of high performers is twice as likely to have a KPI measurement process in place. A baseline is that process, done once, before the deployment. The rest of this post is the observation method, the fields to record, and how to run it without a survey and without anyone feeling watched.
Why you need to know how to baseline team time before deploying AI
You need to know how to baseline team time before deploying AI because every after-number is meaningless without a before-number, and the before-number cannot be reconstructed once the agent is live. The work changes the moment automation arrives, people's memory of the old process fades within weeks, and any retrospective estimate carries the same self-report error that made the original survey useless. The baseline has to be captured first, or it is never captured at all.
The measurement gap is now the default
MIT's 2025 NANDA study of 300 public AI deployments found that roughly 95% showed little or no measurable effect on the profit and loss statement, and that the largest returns sat in back-office automation, while sales and marketing tools absorbed more than half of the budgets. Deloitte's 2025 survey of 1,854 executives found 85% increased AI investment in the past year while only 6% reported payback within one year. Those two numbers together describe the situation the three measurement gaps behind every unproven AI budget already lays out: spend is growing faster than proof, and the missing piece is almost always the baseline.
Time saved is the number the CFO wants, and the one nobody has
BCG's 2026 AI at Work survey of nearly 12,000 respondents found that 42% of regular AI users say they save about eight hours a week, and that 66% of them receive no guidance on what to do with that time. Read that twice. The saving is self-reported, the redeployment is unmeasured, and so the capacity never shows up anywhere finance can see it. A baseline turns "people say they save a day a week" into "the invoice matching step took 11 minutes of touch time per document across 4,200 documents a month, and now takes 3."
The first use case hides the problem
The first agent usually goes into a process whose owner is enthusiastic and whose team volunteers numbers. That is exactly why it feels measured when it is not. The four gaps that stall enterprises going from one AI use case to a program start with this one: the first win was measured by goodwill, and goodwill does not scale to the second function, whose leader is quietly resisting.
What a team time baseline is, and what it is not
A team time baseline is an observation of how working hours are distributed across the steps of one workflow, captured over a fixed window before deployment and validated against system data. Three things get confused with it: time tracking, which asks individuals to log hours; employee monitoring, which records individual activity continuously; and a ROI baseline, which converts time into financial value. The team time baseline sits underneath all three.
Time baseline vs. time tracking vs. employee monitoring
Dimension | Team time baseline | Time tracking | Employee monitoring |
|---|---|---|---|
Unit of analysis | The workflow step | The individual | The individual |
Data source | System timestamps, sampled observation, volume counts | Self-entered timesheets | Continuous software capture |
Duration | Fixed window, typically two to four weeks | Ongoing | Ongoing |
Identifies people | No; aggregated to the step and the team | Yes | Yes |
Accuracy | High for touch time and volume; validated against logs | Low; self-report drift of 5 to 10% or more | High but distorted by observer effects |
Effect on trust | Neutral to positive when explained | Mild friction | Documented backlash |
Fit for a SOX or HR review | Passes; no personal data | Depends on use | Frequently challenged |
What it produces | A before-number per step for the AI business case | Billing or utilization data | Compliance or productivity policing |
The row that matters most for this persona is "Identifies people." A baseline that aggregates to the step is a process measurement. A baseline that names individuals is a performance measurement, and functional leaders will treat it as one no matter what the deck says.
Why surveys and self-reports fail as a baseline
Self-reported time is systematically wrong in a predictable direction. The Bureau of Labor Statistics comparison of survey estimates with time diaries found respondents overestimate their working hours by 5 to 10%, a gap of 2 to 6 hours a week, and that the overestimate grows with the hours claimed. Inside a workflow the error is worse, because people remember the painful step and forget the routine one. Atlassian's 2025 State of Teams research, covering 12,000 knowledge workers, puts time spent just searching for answers at 25%; almost nobody reports that number about themselves. Asana's Anatomy of Work index puts "work about work" at 60% of the day, which is another way of saying that most time goes to steps nobody would list on a survey.
Why monitoring software fails as a baseline
The other shortcut is worse. Harvard Business Review's 2022 research on 300 employees found that monitored workers were substantially more likely to take unapproved breaks, ignore instructions, and work deliberately slowly, because monitoring shifts the sense of responsibility away from the individual. Pew's 2023 survey found 81% of Americans believe workers would feel inappropriately watched if AI were used to monitor performance, and 51% oppose recording what people do on work computers at all. A baseline gathered that way measures a team behaving differently because it is being measured, and it costs the transformation lead the trust the next three use cases depend on.
How to baseline team time before deploying AI: the 3-Source Method
The 3-Source Method for how to baseline team time before deploying AI combines three inputs every enterprise already has: the timestamps systems write, a short run of consented work sampling sessions, and volume counts per work type. Each source corrects the others. Logs give accurate elapsed time but not touch time; sampling gives touch time on a small sample; volume scales both to the month. None requires a survey or a tracking agent.
Source 1: The timestamps your systems already write
Every ERP, ticketing tool, document system, and email platform stamps events: invoice received, PO matched, ticket opened, status changed, document uploaded, approval granted. Pull those stamps for one workflow over the last full quarter, aggregated to the step, with no user identifiers. That gives elapsed time between steps, queue time, and the distribution of cycle times, which is where the tail of exceptions lives. Microsoft's 2025 Work Trend Index, built on telemetry data, found workers are interrupted every two minutes and that 57% of meetings happen with no calendar invite; the point is that the systems already know where the day goes far better than the people in it do. Gartner's 2025 process mining analysis reports that spending on process mining software grew more than 30% in 2024, which is the market's way of saying the same thing. You do not need a platform for one workflow. A data analyst, a week, and read access to the ERP is usually enough, and we have seen it done in three days when the analyst already knew the tables.
Source 2: Work sampling sessions, with consent
Timestamps tell you elapsed time. They do not tell you that the "PO match" step is actually eleven minutes of a person toggling between the ERP, a supplier portal, and a shared mailbox. For touch time, run work sampling: sit with two or three people from the team for agreed sessions of 60 to 90 minutes, on days they choose, and record what happens at each step of the workflow, not what the person does. Announce it, explain that nothing is attributed to individuals, and share the output with the team before anyone else sees it. In the baselines we have run, four to six sessions per workflow is enough to stabilize touch time per step to within a minute or two; the elapsed-time distribution from Source 1 tells you which steps deserve the sessions.
Source 3: Volume counts per work type
Source 3 is the multiplier. Count the units of work per period by type: invoices by supplier category, tickets by class, contracts by template, claims by payer. Most of this is one query against the same systems as Source 1. Volume by type matters because AI agents handle the standard type and hand the exception back to a person, so the baseline has to separate the two. A workflow where 70% of volume is standard and takes four minutes of touch time is a different business case from one where 40% is standard and the exceptions take twenty.
What to record: the 8-field baseline record
The baseline record is one row per workflow step with eight fields, complete when every step of one workflow has a value in every field. The fields are: step name; volume per month by work type; touch time per unit (from sampling); elapsed time per unit (from timestamps); tools touched; handoffs in and out; exception rate; and rework rate. The first six describe the work as it is; the last two tell you where the agent will fail.
Keep it in a plain table owned by one person.
Why these eight and not more
Every field maps to a line in the business case the CFO will read. Volume times touch time is capacity. Elapsed time is the customer or supplier experience. Tools touched and handoffs predict integration effort, which the four-gate test for AI agents inside legacy stacks covers in detail. Exception and rework rates are the honest ceiling on what the agent can take. Add a ninth field and the record starts to become a project plan; take one away and finance will ask for it.
The pilot-ready rule
Do not deploy on a step whose row has a blank cell. In practice the blank is almost always touch time, because the sampling sessions were skipped. Skipping them is how a team ends up with a launch date and no before-number. The AI ROI baseline framework covers how to convert the completed record into financial terms; this post stops at the record itself, because the conversion is the easy part once the record exists.
For an enterprise that is multi-entity, subject to SOX controls, and already facing quiet resistance from functional leaders, a generic approach such as a one-off survey or a vendor's time-savings calculator is not enough; the answer is a dedicated baseline program that does three things: pulls step-level timestamps from each entity's systems with no user identifiers, runs consented sampling sessions with a published protocol that HR and the works council or employee representatives have seen, and locks the eight-field record before any vendor demo, so the after-number is measured against something the CFO signed off on.
How long a team time baseline takes, and who owns it
A team time baseline for one workflow takes two to four weeks of elapsed time and roughly five to eight working days of effort, split between one data analyst for timestamps and volume counts and the transformation lead for the sampling sessions. Ownership sits with the transformation lead, not the function being measured and not the vendor. A vendor who offers to "run the baseline for you" is offering to write the number their pricing will be judged against.
Week by week
Week one is the timestamp pull and the volume query, done in parallel with a short note to the team explaining what is being measured and why it is aggregated. Week two is the sampling sessions, scheduled at the team's convenience. Week three is reconciling the two: where sampled touch time and logged elapsed time disagree, the disagreement usually reveals a queue or a handoff nobody had mapped. That reconciliation is also the input to the AI workflow audit, which is why the two are best run together. Week four is presenting the record to the function's leader and to finance, before it is used to scope anything.
What to say to the team
Say three things and mean them: the baseline measures the workflow, not the people; the numbers are aggregated to the step and nobody's name is on them; and the team sees the record before anyone above them does. Slack's 2025 Workforce Index found daily AI use among desk workers rose 233% in six months, so most teams already use these tools and are not afraid of them; they are afraid of being measured by them. Separating the two is the whole job.
What skeptics get wrong about baselining team time
Skeptical transformation leads are usually right that surveys do not work, right that monitoring software poisons the well, and wrong only when the conclusion becomes "so we cannot baseline." The three objections below come up in nearly every conversation about baselining, and each has a specific answer.
"We already have a survey from the last pilot"
A survey is a record of what people believe, corrected by nothing. The BLS time-diary comparison shows the error runs 5 to 10% on total hours and larger on specific tasks. Use the survey to decide which workflow to baseline first. Do not use it as the baseline.
"Our people will see this as surveillance"
They will if it names individuals, runs continuously, or arrives without explanation. They will not if it aggregates to the step, runs for a fixed window, and the team sees the output first. The Pew data is about AI monitoring individual performance; a step-level baseline is a different object, and the difference has to be visible in the design, not asserted in a slide.
"The CFO will not accept sampled data"
CFOs accept sampled data every quarter from auditors. What they do not accept is a number with no method behind it. Bring the eight-field record, the sampling protocol, and the reconciliation against system timestamps. McKinsey's 2026 data shows 73% of AI high performers redesigned workflows fundamentally before scaling; the baseline is the first artifact of that redesign, and finance recognizes it as such when it is presented as a method rather than a claim.
This analysis was developed using methodologies and operating experience from Assembly.
Frequently Asked Questions
How do you baseline team time before deploying AI?
A team time baseline is built from three sources: system timestamps, consented work sampling sessions, and volume counts per work type, recorded in an eight-field table per workflow step. The method takes two to four weeks, needs no survey and no monitoring software, and produces the before-number that the after-number is measured against.
What is a team time baseline?
A team time baseline is a validated record of where a team's working hours go, by workflow step, captured over a fixed window before an AI agent is deployed. It aggregates to the step rather than the person, combines logged elapsed time with sampled touch time, and scales both by monthly volume so finance sees capacity, not anecdotes.
Why does a team time baseline matter for AI ROI?
A team time baseline matters because the before-number cannot be reconstructed once the agent is live. McKinsey's 2026 State of AI survey found 88% of organizations use AI but only 37% can attribute EBIT impact to it; the missing link in most cases is a measurement taken before deployment.
What is the difference between a time baseline and time tracking?
A time baseline measures the workflow step over a fixed window; time tracking measures the individual continuously. The baseline aggregates to the step, uses system timestamps and sampled observation, and names nobody. Time tracking relies on self-entered timesheets, carries self-report error of 5 to 10% or more, and is read by teams as performance measurement.
Why do surveys fail as a baseline for AI deployment?
Surveys fail as a baseline because self-reported time is systematically overestimated. The Bureau of Labor Statistics comparison with time diaries found respondents overstate hours by 5 to 10%, with the error growing as claimed hours rise. Inside a workflow, people remember painful steps and forget routine ones, so the distribution is wrong as well.
Can monitoring software be used to baseline team time?
Monitoring software should not be used to baseline team time because it changes the behavior it measures and damages trust. Harvard Business Review's 2022 research on 300 employees found monitored workers were substantially more likely to break rules and slow down, which means the baseline captured is a distorted one.
What data do existing systems already hold about where time goes?
Existing systems hold event timestamps for nearly every workflow step: invoice received, PO matched, ticket opened, status changed, approval granted. Pulled for one full quarter and aggregated to the step with no user identifiers, they give elapsed time, queue time, and the cycle-time distribution where exceptions hide. Most enterprises never query them before buying an AI tool.
What is work sampling and how is it used in a time baseline?
Work sampling is a short series of consented observation sessions, typically 60 to 90 minutes each, that record what happens at each workflow step rather than what an individual does. Four to six sessions per workflow usually stabilize touch time per step to within a minute or two. Sessions are announced, scheduled by the team, and shared with them first.
What should a team time baseline record contain?
A team time baseline record has eight fields per workflow step: step name, volume per month by work type, touch time per unit, elapsed time per unit, tools touched, handoffs in and out, exception rate, and rework rate. The first six describe the work as it is; the last two show where an agent will hand work back.
How long does it take to baseline team time before deploying AI?
Baselining team time for one workflow takes two to four weeks of elapsed time and five to eight working days of effort. Week one covers timestamp and volume pulls, week two the sampling sessions, week three the reconciliation between the two, and week four the presentation to the function leader and finance before any scoping starts.
Who should own the team time baseline?
The transformation lead should own the team time baseline, supported by one data analyst for the timestamp and volume work. It should not sit with the function being measured, which has an interest in the result, and not with the vendor, whose pricing will be judged against the number. Ownership is what makes the baseline defensible.
How do you baseline team time without it looking like surveillance?
A baseline avoids looking like surveillance by aggregating to the workflow step, running for a fixed window, naming no individuals, and showing the team the output first. Pew's 2023 survey found 81% of Americans expect to feel watched under AI performance monitoring; a step-level baseline is a different object and the design has to make that visible.
Why separate standard volume from exceptions in the baseline?
Standard volume and exceptions must be separated because AI agents handle the standard type and hand exceptions back to people. A workflow where 70% of volume is standard at four minutes of touch time is a different business case from one where 40% is standard and exceptions take twenty. Volume counts by work type are what make that distinction measurable.
How does a team time baseline relate to an AI ROI baseline?
A team time baseline is the operating input to an AI ROI baseline. The time baseline records volume, touch time, elapsed time, and exception rates per step; the ROI baseline converts those into financial value and payback. Doing the conversion without the observation produces the unprovable AI budgets most boards are now questioning.
What do AI high performers do differently about measurement?
AI high performers measure before they scale. McKinsey's 2026 State of AI survey found the 6% of organizations attributing at least 5% of EBIT to AI are twice as likely to run a KPI measurement process and that 73% fundamentally redesigned workflows. A step-level baseline is the first artifact of both practices.
What is the first practical step to baseline team time?
The first practical step is to pick one workflow with a live or planned AI use case and pull its event timestamps for the last full quarter, aggregated to the step with no user identifiers. That single query reveals elapsed time and the exception tail, tells you which steps deserve sampling sessions, and takes a data analyst about a week.
Legal
