Blog · Article

From a Dozen AI Pilots to One Autonomous Back Office

Lennard Kooy·Sep 4, 2026·16 min read
From a Dozen AI Pilots to One Autonomous Back Office

Most companies do not have an AI problem. They have a dozen AI pilots and no route into production. At Lleverage we think the fix is a change of unit, not a fourteenth experiment. Stop buying pilots per department. Build one autonomous back office instead, where agents run real work inside the systems you already own.

Walk into any manufacturer or wholesaler of a few hundred people and you will find the evidence in a slide deck. A chatbot on the intranet that three people use. A document reader that handles one supplier and breaks on the rest. A copilot licence rolled out to finance in March that nobody renews the training for. Each was reasonable on its own. Together they have produced almost no change in how the work actually gets done, and the board is starting to ask why. This guide sets out what goes wrong between a promising demo and a running process, and the sequence we use to close that gap.

We are usually called in after the third or fourth stalled experiment, which is the vantage point this is written from. Lleverage builds agents that run order intake and invoice matching inside ERPs such as Business Central, SAP and Exact, with your team approving what matters. If you would rather see the difference than read about it, book a demo.

Why do most AI pilots never reach production?

Pilots die in the gap between demonstrating a capability and owning a process. A demo proves that a model can read an order. Production is a harder bar. Every order gets read, every exception gets routed, every posting lands in the ERP, and somebody is accountable on the Tuesday after go-live. Almost no pilot is scoped to answer the second question.

The numbers are stark, and they are not a vendor's numbers. MIT's Media Lab published a report in July 2025 called The GenAI Divide. Drawing on 52 executive interviews, 153 survey responses and 300 public AI deployments, it found that 95% of the pilots studied delivered no measurable impact on profit or loss. The researchers put the enterprise spending behind those pilots at 30 to 40 billion dollars. Gartner reached a compatible conclusion from a different angle in June 2025. It forecast that more than 40% of agentic AI projects will be cancelled by the end of 2027, on cost, unclear value or missing risk controls.

What both findings describe is not a technology failure. The models in these pilots usually worked. What was missing was everything around them. A connection to the system of record. A rule for what happens when the document is ambiguous. A named owner, and a budget line that survives the quarter in which the sponsor changes job. A pilot is designed to answer "can it?". Production has to answer "who runs it, what does it cost, and what happens when it is wrong?", and those three questions are where most programmes quietly stop.

There is a second, subtler reason, and it is the one we see most often in mid-sized companies. Pilots are commissioned per department. Finance buys one thing, the order desk trials another, IT evaluates a third. Each is small enough not to need a business case and small enough not to change anything. Nobody is looking at the whole administrative load of the company. So nobody notices that the same document is being read three times by three different products, and that none of them writes anything back into the ERP.

What does AI pilot fatigue actually cost?

The direct cost of a stalled pilot is small, which is exactly why it is dangerous. A few thousand euros of licences and a month of somebody's attention rarely triggers a post-mortem. The real cost is the credibility spent. After the third experiment that changed nothing, the operations team stops volunteering processes, and the next proposal has to fight the memory of the last one.

Meanwhile the administrative work does not go anywhere. In the operations we walk into, a mid-sized wholesaler will have three to five people whose entire day is reading inbound documents and retyping them. Topa Bathroom Products, a Dutch bathroom wholesaler serving around 700 customers, had four and a half people doing nothing but keying orders into Business Central. That is about 3.8 full-time equivalents, and it was the position before any of this started. That is the cost that keeps running while the pilots come and go.

"We had four and a half people — about 3.8 full-time equivalents — sitting there all day long, manually entering every single order into our Business Central ERP." – Bryan van Ingen, Operations Director, Topa Bathroom Products

The wider picture is that adoption is climbing while impact is not. Eurostat reported that 20.0% of EU enterprises with ten or more employees used AI technologies in 2025, up from 13.5% in 2024, with large enterprises at roughly 55%. Adoption and transformation have come apart. More companies are using AI every year, and the share of them that can point to a process running differently because of it has barely moved.

What is the difference between a dozen pilots and one autonomous back office?

A pilot is a capability shown in isolation. An autonomous back office is a set of agents that read the inbound documents, take the routine decisions and complete the work inside the company's existing systems. People supervise the exceptions. The distinction is not scale. It is whether the work finishes.

The term is ours, and we set out what we mean by it in our definition of the autonomous back office. The shortest version: an agent that drafts an email for a person to send has not finished the work. A human still has to do the last mile, every single time. An agent that posts a validated sales order into Business Central and confirms it back to the customer has finished it, and the desk's capacity actually changes.

Three differences follow from that, and they are the ones that decide whether a programme compounds or stalls.

A dozen pilotsOne autonomous back office
Unit of workA capability, demonstratedA process, completed in the ERP
Where it runsAlongside the systems of recordInside them, over the existing integrations
Who owns it after go-liveThe project sponsor, until they move onThe operations or finance lead who owns the process
What the second one costsThe same as the firstLess, because the rules already exist
What accumulatesSlide decksRules, exceptions and corrections your company owns
How success is measuredModel accuracy on a sampleHours removed from a named desk

The last row of that table is the one we would argue hardest for. A pilot reports accuracy because accuracy is what a demo can show. A production process reports hours, exceptions and cycle time, because those are what the person paying for it feels. If a proposal on your desk cannot name the desk it will change and the hours it will remove, it is a pilot regardless of what it is called.

How do you pick the first process to take to production?

Pick the process where the volume is boring, the format is messy, and the outcome lands in one system. High volume gives the agent enough examples to be corrected against. Messy formats are where rule-based automation has already failed, so the comparison is honest. A single destination system means success is unambiguous: the record is in the ERP or it is not.

In practice that first process is almost always one of four in the companies we work with: inbound sales order intake, supplier order confirmation matching, purchase invoice matching, or repetitive customer questions about order status. They share a shape. A document or message arrives in a format the sender chose. Somebody interprets it against rules that live in a colleague's head, and the result has to end up as a record in the ERP.

What to avoid on the first attempt is just as useful:

  • Processes with no system of record at the end. If the output is a judgement that lives in an email thread, there is nothing to measure and nothing to write back.
  • Processes that touch six systems. Every additional system is another integration, another owner and another reason for the go-live date to slip.
  • The lowest-volume, highest-prestige process. The month-end board pack is visible, rare, and a poor teacher for an agent.
  • Anything where the current process is already clean and rule-based. If a straightforward integration would fix it, do that instead and keep the credibility.

One more filter, which matters more than it looks. Choose a process whose current owner wants it changed. Xpol, a Dutch fresh-flower wholesaler of about 25 people, started its order-intake work in the run-up to a senior specialist's retirement, with volume climbing and no replacement in sight. The desk wanted help. Where the desk does not want help, the exceptions never get corrected and the agent never improves.

How does a pilot earn its way into production?

Autonomy is granted in stages, on evidence, by the people who carry the risk. We run it as five steps: scan, foreground, shadow, background, compound. Each step ends with a decision your team takes on the record, based on what the log shows rather than on a promise from us.

Scan: find the hours before writing any rules

Two to three days with the operations and finance leads, mapping where the manual work and the exceptions actually sit. The output is not a strategy document. It is a short list of processes with a volume, a current handling time and a named owner for each. Most companies are surprised by which process comes top; it is rarely the one the pilots were pointed at.

Foreground: the agent proposes, a person approves everything

The agent reads live documents and prepares the result, and a person checks and approves every single one. This is slow on purpose. It produces the evidence, and more importantly it produces the corrections, which are the raw material for everything that follows. A controller who has approved 400 order postings and seen the reasoning behind each one is in a position to widen the mandate. One who has seen a demo is not.

Shadow: running alongside the old way

The agent handles the flow end to end while the existing process continues in parallel, and the two are compared. This is the step most programmes skip, and skipping it is why so many go-lives get reversed in week two. Shadow running is where you find the customer who sends orders as a photograph of a fax. It is also where you find the supplier whose confirmations use a different unit of measure, and the twelve other things that never show up in a sample of fifty documents.

Background: routine cases run on their own

Once the deviations are understood and encoded, the routine cases run without approval and the exceptions route to a person with context attached. The agent is now doing the work. Supervision has not disappeared; it has concentrated on the cases that need judgement, which is where the operations team was useful all along.

Compound: the second process costs less than the first

Every correction, rule and preference captured in the first process is available to the second. The customer-specific rules the order desk taught the intake agent are the same rules the invoice agent needs. This is the step that changes the economics, and it is the one a portfolio of separate pilots can never reach, because each pilot starts from zero.

"It's a matter of building trust in the organisation with these kinds of initiatives. You can't just throw something like this over the fence." – Cees Maaskant, General Manager, Xpol

What does the second agent get for free?

The second agent inherits the company's accumulated rules, its exception history and its integration into the ERP, so it reaches production faster and at lower cost than the first. This is the compounding effect, and it is the strongest practical argument against running a portfolio of unrelated pilots.

Consider what the first process actually produces beyond its own output. It produces a working connection to the ERP and the permission model around it. It produces a written record of customer-specific quirks. Which customer uses which unit, which one always omits the delivery date, which supplier's reference number needs reformatting. It produces a record of what the team corrected and why. Xpol's intake agent codifies 25 customer-specific rulesets that previously lived in the heads of senior staff, including unit conversions, weekday-specific label text and multi-depot splits.

None of that transfers between disconnected pilots, because each is built on a different foundation with a different vendor and a different data model. Inside one intelligence layer, all of it transfers. The practical effect on a mid-sized company is that the first process takes weeks and the third takes days. The business case for process four then writes itself, from the evidence of processes one to three.

There is a defensive argument here too, and it is one operations leaders raise with us more than they used to. Knowledge that lives in a system the company owns does not walk out of the door when a senior colleague retires. That was Xpol's explicit reason for starting.

What has to be true inside the company, not just the technology?

Three things, and none of them are technical. There has to be a named process owner in operations or finance rather than in IT. There has to be a decision rule for widening autonomy. And there has to be a budget that is attached to the process rather than to the experiment.

The ownership point is the one that decides most outcomes. When the owner is a project sponsor, the agent's mandate stops widening the moment that person's attention moves, because nobody else has standing to approve the next step. When the owner is the person who runs the order desk, the widening is continuous and self-interested: they benefit directly from every case that stops needing their approval.

Governance is the second, and it is more prosaic than the word suggests. In practice it means three things. An audit trail of what the agent did and why, approval thresholds set by the business, and exceptions routed to a named person rather than into a queue. On the control side this is a product question. The part that fails is usually organisational. Nobody decided who signs off a variance above a certain amount, so everything gets escalated and the promised time saving never materialises.

The third is money, and it is the quiet reason pilots stay pilots. An experiment is funded from an innovation budget that is renewed annually and defended weakly. A process is funded from the operating budget of the department whose hours it removes. Our own pricing reflects that view. We charge per agent per month, from 2,000 euros for a standard mid-market deployment. An agent doing a job should sit on the same budget line as the people doing that job today. Whatever the vendor, the test is the same. If the funding line disappears when the innovation programme is reviewed, the agent is still a pilot.

A worked example: from three stalled experiments to a running order desk

The pattern below is a composite of deployments we have run, with the published figures attributed to the companies they belong to. It is written the way these projects actually unfold, not the way a case study summarises them.

The starting position

A wholesaler of roughly 200 people, running Business Central. Three AI initiatives in the previous eighteen months. A chatbot on the website. A document-reading trial that handled two suppliers and was abandoned when a third format appeared. Copilot licences for finance. Order intake is handled by four people reading emails, PDFs, Excel attachments and the occasional scanned form, and keying the results in by hand. Volume is growing about 15% a year and the desk is the constraint on taking on new customers.

What the scan found

The scan takes two days and produces an uncomfortable finding: the highest-value process is not the one any of the three pilots addressed. Order intake absorbs roughly 60% of the desk's time, and the errors it produces generate a second, invisible workload in customer service and credit notes. The document-reading trial had actually been pointed at the right kind of work. It failed on coverage rather than on capability, which is the most common way these trials fail.

The first eight weeks

Weeks one and two are foreground: the agent reads live orders, prepares the Business Central entry, and a person approves every one. Roughly one in six needs a correction, which sounds like a poor result and is in fact the point. Weeks three to five are shadow running. The coverage problem from the old trial reappears, in the form of two customers whose formats nobody had written down. This time they get encoded rather than abandoned. By week six the routine cases are running in the background and approvals have concentrated on price deviations and unknown article codes.

Where it lands

The published outcomes from comparable deployments give the shape of the result. Topa now has more than 90% of inbound orders posting into Business Central automatically, with confirmations back to the customer inside 30 seconds. Close to four full-time roles moved onto after-sales and service planning. Xpol saves around 20 minutes per large order across roughly 150 orders a week and avoided one to two planned hires. The second agent at both companies took a fraction of the effort of the first, because the rules and the ERP connection were already there.

How do you measure whether this is working?

Measure hours removed from a named desk, the share of cases completing without a human touch, and the time from a process being chosen to it running in production. Model accuracy is a diagnostic, not a result. If a metric would not appear in an operations review, it is not the metric to run this on.

Four measures we would put on the first review, deliberately few:

  1. Straight-through rate. The percentage of cases the agent completes with no human touch. Expect it to start low during foreground running and climb as exceptions get encoded. A rate that plateaus early usually means corrections are not being captured.
  2. Hours returned, by desk. Not "hours saved" in the abstract, but the change in what a named team spends its week on. This is the number that survives contact with a CFO.
  3. Exception mix. Which cases still need a person, and whether that list is shrinking or just changing shape. A stable exception list is a healthy end state; a growing one means the process was scoped too broadly.
  4. Time to production for the next process. The compounding test. If process three takes as long as process one, the intelligence layer is not accumulating anything and you are running pilots again under a different name.

What we would leave off the list is anything that cannot be tied to a person's week. Accuracy on a held-out sample, tokens processed, documents ingested: useful for debugging, useless for deciding whether to widen the mandate.

Frequently asked questions

How many AI pilots is too many?

More than two running at once, in a company of a few hundred people, is usually a sign that nobody is comparing them. The problem is rarely the count itself. It is that separate pilots cannot share rules, integrations or corrections, so the eleventh costs exactly what the first one did and teaches the organisation nothing new.

Should we cancel our existing pilots before starting?

Not necessarily. Keep any pilot that is already changing a named process and can say who owns it after go-live. Stop the ones that cannot answer that question, and reuse what they taught you. In our experience the useful residue of a failed trial is the list of document formats it choked on. That list is exactly what the next attempt needs.

How long does it take to get a first process into production?

For a single, well-chosen process in a company running a mainstream ERP, weeks rather than months. The scan takes days. Foreground running takes two to three weeks to build enough evidence, and shadow running another two to four, depending on how varied the inbound formats are. Programmes that quote nine months are usually scoping several processes at once.

Does this replace our ERP?

No. The agents run inside the ERP you already have, over its existing interfaces, and write records into it the way a person would. Business Central, SAP, Exact, AFAS and Infor are the systems we see most often. Replacing the system of record is a different project with a different risk profile, and confusing the two is a good way to stall both.

What if our processes are too specific for AI?

Specific is the argument for this approach rather than against it. Generic products fail on specific work because they have no opinion about what a sales order means in your business. The rules that make your order desk hard to automate are the same rules an agent learns once and then applies consistently. That is why the first process is worth the corrections it takes.

Turn the experiments into a running process

If you have a folder of AI pilots and no change in how the work gets done, the next move is not another trial. It is a scan of where the hours sit and one process taken all the way into the ERP. That is where we would start, and it is what we do with customers in the first fortnight.

Book a demo and bring your own documents: the messy ones, the customer who sends orders as a photograph, the supplier whose confirmations never match. Those are the cases that decide whether an agent belongs in production, and they are the ones a demo should be tested against.

Give your back office an AI workforce