Supervised agents run in the foreground: a person triggers the work, reviews the output and approves it. Background automation runs unattended and only surfaces exceptions. Lleverage's view is that the mode is a property of the process rather than of the technology, and that most back-office work starts supervised and earns its way into the background.
The distinction matters because it is the question every operations lead actually has to answer, and almost nobody phrases it that way. The vendor conversation is usually about categories: agents against workflows, AI against robotic process automation, autonomy against rules. Those are useful definitions and they settle nothing on a Tuesday morning, when you have an order desk drowning in PDFs and you need to decide whether a person still checks every one. This guide covers what each mode is, which processes belong in which, how to run the decision as a repeatable test, and how a process graduates from one mode to the other.
We build agents for manufacturers, wholesalers and distributors whose order desks and finance teams look much like yours, so our interest here is plain. It is also the reason the examples below are named companies with numbers attached rather than composites. If you want to see the shape of this on your own paperwork, look at how order intake runs when an agent owns it, or book a demo.
What is the difference between supervised agents and background automation?
A supervised agent does the work and hands it to a person before anything is committed. A background automation commits the work itself and escalates only what it cannot resolve. Both can be the same underlying agent with the same business rules. What changes is who holds the final decision and how often a human touches the process.
The confusion in the market comes from treating these as two different products. They are not. In our experience the same agent, reading the same order emails against the same customer rulesets, is deployed in supervised mode at one company and in background mode at another. What differs is what each business is willing to have happen without a person watching. A wholesaler whose customers reorder from a settled article list can let orders post themselves. A precision engineering firm quoting bespoke assemblies cannot, and should not want to.
There is a second, quieter difference that only shows up after go-live. Supervised work scales with attention: if the volume doubles, the review burden doubles, and eventually you are back to hiring. Background work scales with exceptions, so if the volume doubles but the exception rate stays flat, the team handles twice the work at roughly the same cost. That is the whole economic argument for moving processes into the background, and it is also why moving them too early is expensive.
| Supervised agent | Background automation | |
|---|---|---|
| Who commits the result | A person, on every run | The agent, on every run |
| When a human is involved | Every time | Only on exceptions |
| Typical trigger | Someone starts it | An inbound email, file or schedule |
| Cost curve as volume grows | Rises with volume | Flat until exceptions rise |
| Best fit | Judgement work, high variance, high consequence | Repeatable decisions, known formats, defined exception path |
| What it needs to work | A person who owns the output | An exception queue someone owns |
| Failure mode | Review becomes a formality | Errors accumulate silently |
| What it feels like to the team | An expert colleague on demand | Work that has already been done |
Read the last row carefully, because it is the one that predicts adoption. A supervised agent is something people use. A background automation is something people notice only when it stops. Those are different change-management problems, and the second one needs more evidence up front than the first.
Which processes need a supervised agent?
Supervised agents fit work where a person is accountable for the answer, the inputs vary in ways nobody can fully enumerate, and the cost of a confident mistake is higher than the cost of a review. Technical troubleshooting, quoting non-standard configurations, and answering a customer question that touches a commitment all sit here. The agent removes the search, not the judgement.
Oude Reimer is the clean example. The service team supports machines built by many different manufacturers, and every troubleshooting call used to depend on which technician picked up the phone and what they happened to remember. Lleverage built a searchable knowledge base over 170 manuals, and at Oude Reimer technical triage that once meant scrolling through hundreds of pages now returns a referenced answer in 70 seconds. The technician still decides what to tell the customer. That is the point.
"The answers come back with references to the exact manual and section. That is what builds trust — the technician can verify the information themselves before passing it to a customer." Remco Hooft, Technical Owner, Oude Reimer
Customer support sits in the same category for a different reason. At J. Kisch & Zonen, a furniture supply business trading for 130 years, roughly 80% of the daily support questions are repeats of each other. The agent now handles the majority of that repetition, so the team can spend its time on the genuinely complex cases. The volume is high enough to justify automation, but the replies go out under a person's name to a customer with a live relationship, so a human stays on the send button.
The pattern underneath all three cases is worth naming, because it is more useful than any list of use cases. Supervise the work when the person doing it would be uncomfortable defending an answer they had not read. That discomfort is not resistance to change. It is an accurate reading of where the accountability sits, and any deployment that overrules it tends to get quietly abandoned within a quarter.
There is one more category that belongs here and is routinely overlooked: anything you have never measured. If nobody can tell you the current error rate of a process, you cannot tell whether an agent running it unattended is an improvement. Supervised mode generates that baseline as a by-product, because every correction a reviewer makes is a data point.
Which processes should run as background automation?
Background automation fits high-volume work where the decision is repeatable, the inputs arrive in recognisable shapes, and there is a defined path for the cases that do not fit. Order intake, invoice matching, document classification and master data updates are the usual candidates. What decides it is whether the exceptions can be named in advance.
Topa Bathroom Products is the reference case. Before the agent, four and a half people spent their days typing incoming orders into Microsoft Dynamics 365 Business Central. Now over 90% of incoming orders at Topa are processed straight into Business Central with no manual input. Customers receive their confirmation within 30 seconds of sending the order, and the equivalent of nearly 4 FTEs moved onto after-sales support and service planning. None of that requires a person to look at an order before it posts.
What makes it work is not the accuracy figure. It is the exception design. When a commission number is missing, the agent drafts the response email rather than guessing. When an article code is unrecognised, it flags the order for human review instead of picking the nearest match. Every processed message is labelled in Outlook so the team can see what the agent touched. The agent is trusted with the routine precisely because the boundaries of the routine are drawn explicitly, and everything outside them is handed back.
That design principle generalises. A background automation is only as safe as its escalation rules. The useful scoping question is therefore not about capability at all. It is about what the agent does when it runs out of rules. If the honest answer is that it would proceed on a best guess, the process is not ready for the background yet, whatever the demo showed.
Volume matters too, though less than people expect. A process running 40 times a day is an obvious candidate. A process running twice a month almost never is, because the review cost is trivial and the cost of building and maintaining an unattended path is not. Between those poles, the deciding factor is usually whether the work arrives predictably or in bursts. Burst arrival is exactly when supervised review turns into a backlog, and when background processing starts to pay for itself.
How do you decide which mode a process needs?
Run the process through five questions: who is accountable for the output, how variable the inputs are, whether the exceptions can be named in advance, what a wrong answer costs before anyone notices, and whether the volume justifies the exception path. Three or more answers pointing the same way settle it. A split verdict means start supervised.
The five questions
- Who signs off today? If a named person is accountable to a customer, an auditor or a regulator for this specific output, start supervised. Accountability that cannot be delegated to a person cannot be delegated to an agent either.
- How variable are the inputs? Count the distinct formats and the number of customer-specific rules. Variety is not automatically a blocker, and at Xpol a single agent codifies 25 customer-specific rulesets covering unit conversions, weekday-specific label text and multi-depot splits. Variety you cannot enumerate is the blocker.
- Can you name the exceptions? Write them down. If the list is short, concrete and stable, the background path is designable. If it ends with "and anything unusual", it is not, and that vagueness is the actual finding.
- What does a wrong answer cost before detection? A mispriced quote caught at approval costs a minute. The same quote sent to a customer costs a margin, and possibly the relationship. Time-to-detection matters more than error rate.
- Does volume justify the exception path? Unattended processing needs monitoring, an owner and a queue. Below a few hundred runs a month that overhead rarely pays back, and supervised mode is the cheaper answer as well as the safer one.
A worked example: a wholesaler's order desk
Take a distributor receiving around 400 orders a week across email, PDF attachments, Excel files and a customer portal. Question one: the internal sales team signs off today, but they are transcribing rather than deciding, and no customer commitment is made at the point of entry. That points to background. Question two: four formats and perhaps twenty customer-specific conventions, all documented once someone sits down to write them out. Background again.
Question three is where the process splits in two. Known articles, known customers and clean quantities are nameable and safe. New articles, missing commission numbers and quantities that break a pack size are nameable as exceptions, which means they can be routed rather than guessed. That is a background process with a supervised tail, not a choice between the two modes for the whole desk.
Question four confirms it, because a mis-keyed order caught at picking costs a re-pick, while one caught at delivery costs a return and a credit note. Question five settles the economics at 400 orders a week. The answer for this desk is background processing with an explicit exception queue. That is exactly the shape Koninklijke Dekker landed on when it automated intake of orders arriving as PDFs, spreadsheets and plain text emails. If you want the full intake picture behind that decision, our guide to B2B order management works through the format problem in detail.
How does a process move from supervised to background?
Processes graduate by evidence, not by calendar. The agent runs supervised until its corrections drop to a level the team is comfortable with, the exception list stops growing, and the reviewers start agreeing with it more often than they change it. Then supervision narrows to the exceptions. Our reading is that this is the only sequence that survives contact with a real operations team.
Xpol shows the sequence working. A fresh flower supplier of about 25 people faced a summer of rising order volume with its most senior order specialist about to retire, which is a deadline no pilot can negotiate with. The agent reads any customer's order format, applies those 25 rulesets, and pre-fills Business Central for human review. At Xpol that replaces a manual process taking roughly 20 minutes on a large order, across about 150 orders a week, and it absorbed both the retirement and the growth without the one or two planned hires.
Note what has not happened there. The agent pre-fills and a person still confirms, because the business decided that trust is earned in public and in stages.
"It's a matter of building trust in the organisation with these kinds of initiatives. You can't just throw something like this over the fence." Cees Maaskant, General Manager, Xpol
The three stages, and what ends each one
Stage one, shadow. The agent processes real work and produces an output nobody acts on, while a person does the job as usual. You are comparing, not deploying. This stage ends when the agent's output matches the human's often enough that the comparison stops being interesting, and it typically surfaces two or three business rules nobody had written down.
Stage two, supervised. The agent's output becomes the draft and the person becomes the reviewer. Throughput improves immediately, which is why teams like this stage, and the corrections made here are the real training data for the exception list. This stage ends when reviewers are approving without changes on the routine cases and their edits cluster into a small number of nameable categories.
Stage three, background with a supervised tail. Those nameable categories become escalation rules, the routine runs unattended, and the reviewer's job changes from checking everything to owning a queue. Most of our customers stop here permanently. A fully unattended process with no exception queue is rare and usually a sign that the process was narrower than anyone admitted.
Skipping stage two is the most common and most expensive mistake. It saves a few weeks and costs the thing that makes the whole programme work, which is a team that believes the numbers because they watched them accumulate. Where an agent needs to hold business rules that span systems, the agent layer and the company brain behind it is what makes those rules explicit rather than buried in someone's head.
What does supervision look like once an agent is live?
Supervision after go-live means three concrete things: an exception queue with a named owner, a visible record of what the agent did and why, and a review path for anything it declined to handle. It is an operating routine, not a feeling of oversight. Teams that skip the queue owner end up with unattended processing by accident.
The record matters more than most people expect at scoping time. Topa's practice of labelling every processed email in Outlook is a small implementation detail with a large effect, because it means anyone can reconstruct what happened to a given order without asking. Oude Reimer's answers carry references to the exact manual and section for the same reason. In both cases the traceability is what lets a person disagree with the agent quickly, and the ability to disagree quickly is what keeps people using it.
The queue itself needs a genuine owner rather than a rota. Exceptions are, by definition, the cases where the business rules ran out. The person handling them is doing the most interesting work in the process, and is also the source of the next round of rule changes. Treating that as leftover admin wastes the signal. The teams that do this well review the exception categories monthly and convert the recurring ones into rules, which steadily shrinks the queue.
One warning from repeated deployments. Watch for review becoming a formality, which shows up as approval times collapsing and correction rates falling to nearly nothing while nobody has changed anything. That is not the process graduating. It is supervision decaying, and the fix is either to move the process into the background properly, with real escalation rules, or to sample rather than review everything. Pretending to check is worse than either.
What goes wrong when you pick the wrong mode?
Both errors are recoverable, and they fail differently. Over-supervision fails slowly and visibly: the team keeps working, throughput barely improves, and the business case quietly dies. Under-supervision fails suddenly and invisibly: everything looks fine until a batch of wrong records reaches a customer, an auditor or an ERP.
| Over-supervised | Under-supervised | |
|---|---|---|
| Symptom | Approval queue grows, savings do not appear | Clean dashboards, angry customers |
| Detected by | The business case review | A complaint, a credit note or an audit |
| Time to detection | Weeks | Often months |
| Root cause | No graduation criteria were ever defined | Exceptions were never named |
| Fix | Define stage-three criteria and narrow supervision to exceptions | Return to supervised, rebuild the escalation rules |
| Cost of the fix | A few weeks of unrealised benefit | Rework, credit notes and lost trust |
Over-supervision is by far the more common of the two, and the reason is structural rather than cultural. Nobody is ever criticised for adding a review step, so the review step outlives the uncertainty that justified it. The remedy is to write the graduation criteria at the start, before the agent is live. Name three things: the correction rate, the stability of the exception list and the review period. Together they define the moment this process moves to the background. Criteria written afterwards are always negotiable, and therefore never met.
Under-supervision usually comes from a demo that went too well. An agent that handles fifty test documents flawlessly is not evidence that it handles the fifty-first. The difference between the two is whether anyone asked what the agent does when it is unsure. If your finance processes are the ones in question, the same logic applies to invoice matching and collections as it does to order intake. The escalation rules matter more there, because the money has already moved.
For a fuller picture of how these two modes fit into a back office that runs on agents rather than on individual automations, read our definition of the autonomous back office. If you want the underlying technical distinction between an agent and a workflow, the difference between AI agents and AI workflows covers it directly.
Frequently asked questions
Is a supervised agent the same as human-in-the-loop AI?
Broadly yes, though supervised describes the operating mode rather than the architecture. Human-in-the-loop usually refers to a person approving or correcting model output. Supervised, as we use it, means a person owns the commit step for every run, which is an operational commitment about accountability rather than a statement about how the agent is built.
Can the same agent run in both modes at once?
Yes, and in practice most mature deployments do. The routine cases run unattended while a named set of exceptions routes to a person, which is the stage-three pattern described above. The rules deciding which path a given case takes are part of the agent's configuration, so the split can be adjusted as confidence grows without rebuilding anything.
How long does a process take to move from supervised to background?
It depends on volume rather than time, because graduation is driven by how many cases the reviewers have seen. A process running a few hundred times a week generates enough evidence in weeks. One running a few dozen times a month can take a quarter or more, which is usually an argument for leaving it supervised permanently.
Does background automation mean nobody checks the work?
No. It means nobody checks every run. Someone still owns the exception queue, monitors the volumes and reviews the categories of exception periodically, and in a well-run process that person also converts recurring exceptions into new rules. Removing that role is how unattended processing turns into unnoticed errors.
Which mode should a first AI project use?
Start supervised, unless the process is high-volume with genuinely nameable exceptions and someone has already measured the current error rate. Supervised mode produces the baseline, the exception list and the internal trust that a background deployment depends on, and it does so while already saving time.
The decision is rarely as difficult as the vendor conversation makes it sound, and it is almost never permanent. Pick the mode the process needs today, write down what would justify changing it, and revisit when the evidence arrives. If you want to work through that decision against your own processes rather than in the abstract, book a demo and bring the paperwork you argue about most.
