← All posts Article · Dec 9, 2025

How to Build an Internal AI Agent: A Step-by-Step Guide

Tom van Wees·Dec 9, 2025·14 min read·Updated September 2026
How to Build an Internal AI Agent: A Step-by-Step Guide

How to build an internal AI agent, in short: pick one process that starts with somebody reading something, write down what the agent may decide and what it must escalate, choose build or buy, then run it supervised until the measured error rate earns it more autonomy. At Lleverage we think the sequence matters more than the technology.

Almost every failed internal agent we have seen failed in the first two steps rather than the last five. The process chosen was too broad, or nobody wrote down where the agent's authority ended, and six months later there is a demo nobody uses. The engineering is rarely the hard part now. Deciding what the thing is allowed to do, and who checks it, is.

This guide walks the seven steps in the order they actually happen, with the failure modes at each one. We build these agents for manufacturers, wholesalers and logistics operators, usually starting with order management, and the pattern below is drawn from that work rather than from a reference architecture. If you want it applied to one of your own processes, book a demo.

What is an internal AI agent, and how is it different from automation?

An internal AI agent is a system that reads unstructured input, decides what it means against your own data and business rules, acts inside your systems of record, and escalates what it cannot resolve. Conventional automation executes steps somebody specified in advance. The agent works out what the step should be.

The distinction matters because it changes what you can automate. A rules engine needs the decision made before it runs, which is why it handles the orders that arrive clean and leaves the desk with the awkward ones. An agent takes the reading and deciding as its job, which is the part that consumes the hours.

Conventional automationInternal AI agent
InputFixed format, known field positionsEmail, PDF, scan, spreadsheet, free text
DecisionRules written in advanceInferred from the document and your history
Unexpected inputFails or routes to a personHandles it, or escalates with its reasoning
EvidenceExecution logThe result plus the source it came from
Who maintains itDevelopers, after each changeThe business, by correcting the agent
Right forHigh volume, zero varianceHigh variance, judgement-heavy intake

Our view is that "agent" has been stretched to cover anything with a language model in it, and that the useful test is narrower. If it cannot act inside your ERP and cannot tell you why it did what it did, it is a chat assistant. We have written separately on supervised agents versus background automation, which is the other axis worth being precise about.

Step 1: Which process should your first internal AI agent handle?

Pick a process that begins with somebody reading something, runs at least twenty times a week, and has a person who already owns the outcome. Order intake, supplier invoices, specification requests and customer queries fit. Anything with a single monthly run, no named owner, or no historical record to learn from does not.

The volume floor matters more than it looks. Below roughly twenty runs a week you cannot tell in a reasonable period whether the agent is right, because there is not enough traffic to see the error pattern. You end up arguing from anecdotes. Above that, a fortnight of live running tells you exactly where it struggles.

The reading requirement is the real filter. If the input already arrives structured from a source you control, you do not need an agent, and the honest recommendation is a connection or an existing bot. The agent earns its place where the input is a mess: four hundred customers, each with their own purchase order layout, half of them scanned.

History is the third condition, and the one most often missed. SIG Benelux report that 88% of incoming order lines are matched to the correct article on the first attempt, including lines that arrive with no article number at all, because years of order history in Dynamics 365 do the translation. That history was the asset. No specification document could have replaced it.

The audit that takes an afternoon

  1. List every process where a person opens an attachment before anything happens. That list is your candidate set. Nothing else belongs on it yet.
  2. Count the weekly volume and time one run end to end, splitting reading and deciding from keying and filing. On most desks the ratio is around four to one in favour of reading.
  3. Check what history exists. Two years of processed records in the ERP is a strong base. A folder of PDFs and nothing structured is a weak one.
  4. Name the person who will review the agent's output. If nobody will own that, stop here and pick a different process.

Step 2: How do you scope what the agent is allowed to decide?

Write down three lists before any building starts: what the agent decides on its own, what it drafts for a person to approve, and what it must never touch. Scope is an authority question, not a capability question. Most internal agent projects stall because this was left implicit and discovered in production.

Be specific to the point of pedantry. "The agent handles order intake" is not a scope. "The agent creates a draft sales order when it identifies the customer, matches every line to an article and finds no price deviation above two per cent; otherwise it flags the order with the reason" is a scope. The second version can be tested. The first cannot.

The escalation rules deserve the same care as the happy path, because they are what makes the agent safe to run at all. Decide what happens when confidence is low, when two documents disagree, when a customer is new, and when a value sits outside its usual range. Each of those needs a named destination, not a generic exception queue that nobody reads.

Put the three lists somewhere the reviewer can see them, not in a project document nobody opens after go-live. A page pinned next to the process, in the words the desk actually uses, beats a formal specification. The person correcting the agent on a Tuesday is the one who has to know where its authority ends. Revisit it when you widen the scope, and date each version, so an argument about what the agent was supposed to do has an answer.

Scope also sets the measurement. If the agent is allowed to draft but not post, the metric is how many drafts survive review unchanged. SIG Benelux report that 73% of order emails reach Dynamics 365 as a draft order on their own, with the desk opening and approving it. That is a scope, a metric and an operating model in one sentence.

Step 3: Should you build the agent yourself or buy one?

Build when the process is genuinely unique to you, you have engineers to maintain it for years, and the integration surface is small. Buy when the process is a common shape, when the cost of being wrong is high, and when you need it running this quarter rather than next year. Most mid-sized companies should buy the first one.

Most advice on how to build an internal AI agent opens with this question. In our experience it belongs third, because the answer changes once you know the process and the scope. The build case is stronger than vendors admit. If you employ a capable engineering team and the agent touches one system you wrote yourself, building is reasonable and you keep the knowledge in house. The case weakens on the second year, when whoever built it has moved on and the model landscape has changed twice.

The buy case is not about capability, it is about the unglamorous middle. Connector maintenance, retries, audit trails, permission models, the behaviour when a document is a photograph taken at an angle. That layer is most of the work and none of the interesting part. We have set out the full comparison in build versus buy for back-office AI agents, including where we think building is genuinely the right call.

There is a third option people forget: buy the reading, keep the writing. An agent resolves the incoming document and hands over a clean structured result, and whatever you already own writes it into the system of record. Our integrations cover the common mid-market ERPs directly, but keeping an existing bot on the write step is a perfectly good design.

Step 4: What does an internal AI agent actually need to work?

Four things: access to the input where it arrives, a source of truth to check against, a write path into the system of record, and a reviewer. Everything else is detail. The component most often missing is the second, because companies assume the model supplies the knowledge when the knowledge is in their own data.

The input access is usually a shared mailbox or a folder, and it is worth getting right early. An agent reading a personal inbox will be blocked by holiday, permissions and mail rules. A shared address that the team already uses for orders is the stable version of the same thing.

The source of truth is where most of the quality comes from. Customer records, article masters, price agreements, past orders, specification libraries. Oude Reimer unified 170 equipment manuals from more than 15 manufacturers into one searchable base, and report technical triage answered with references in 70 seconds where it previously meant scrolling hundreds of pages. The agent did not know any of that. The library did.

The write path is the part of how to build an internal AI agent that nobody draws in a diagram, and it decides whether the work is finished or merely prepared. An agent that produces a beautiful summary somebody then retypes has moved the work rather than removed it. Topa Bathroom Products report that over 90% of incoming orders now go straight into Business Central with no manual input, and that customers get a confirmation within 30 seconds. That only counts because the write path exists.

The reviewer is a role, not a job title. Somebody sees what the agent proposed, approves or corrects it, and the corrections feed back. That loop is how the agent improves, and it is also the mechanism by which the team comes to trust it.

Step 5: How do you get an internal AI agent from prototype to production?

Run it in parallel on real work before it touches anything. Feed it the live input, have the agent produce its result, and compare that against what the desk did, without the agent writing anywhere. Two weeks of that tells you the real accuracy and the actual exception pattern, which a test set never does.

Parallel running is the step most teams skip when working out how to build an internal AI agent, and skipping it is why go-live feels like a gamble. Your test documents are the ones somebody chose. Real input includes the order that arrives as three forwarded emails, the scan with the coffee ring, and the customer who writes their quantities in the subject line. You want those in the sample.

Once the comparison holds, go live narrow. One customer, one document type, one person reviewing everything. Widen by adding customers rather than by adding autonomy, because a wider input set with a human still checking is a much smaller risk than full autonomy on a narrow one.

It's a matter of building trust in the organisation with these kinds of initiatives. You can't just throw something like this over the fence.
— Cees Maaskant, General Manager, Xpol

That sequencing is also what makes the change management work. People accept an agent that has visibly been right for a fortnight on their own orders far more readily than one introduced by a presentation. Xpol report around 20 minutes saved on each large order across 150 orders a week, and describe twenty-five customer-specific rulesets that had previously lived in the heads of senior staff.

Step 6: How do you supervise an agent once it is live?

Watch four numbers weekly: the share of runs completed without a person, the share escalated, the share corrected after approval, and the time from arrival to finished record. The third number is the one that matters. Corrections after approval mean the reviewer is rubber-stamping, which is worse than a high escalation rate.

Set the thresholds before go-live rather than after, so nobody has to argue about what good looks like while the agent is running. A useful starting position is that escalations above twenty per cent mean the scope is too wide, and corrections after approval above two per cent mean review has become a formality.

Supervision should get cheaper over time, and if it does not, something is wrong with the design. The pattern we look for is a process moving from every run reviewed, to spot checks, to exceptions only. Microtechniek report around sixty purchase orders a week written into Ridder IQ without anyone typing them, measured across August 2026, and more than 500 hours a year returned to the admin desk, with the administrator stepping in only where the workflow is unsure.

Keep the evidence trail regardless of how autonomous the agent becomes. Every result should be traceable to the document and the rule behind it. SPL Treatments route aerospace parts against a library of more than 800 cross-referencing standards, and report that introducing a first-time part now takes 2.5 minutes with every recommendation cited back to the section it came from. In a regulated process that citation is not a nice extra, it is the reason the output is usable at all.

Step 7: How do you go from one agent to a back office that runs itself?

Add the adjacent process, not the exciting one. If the first agent handles order intake, the next is order confirmation or the invoice that follows, because it shares the customer data, the reviewer and the trust you have already built. Agents that share context compound. Agents scattered across unrelated departments do not.

The compounding is mostly about knowledge rather than code. Once the article-matching knowledge exists, the quotation agent inherits it. Once customer-specific rules are written down for order intake, the delivery agent can use the same ones. J. Kisch en Zonen report 80% less repetition in customer support after the recurring questions were handled by an agent drawing on that shared base.

There is a practical ceiling worth naming. Roughly three or four agents in, the constraint stops being automation and becomes data quality: article masters with duplicates, customers recorded twice, price agreements that live in a spreadsheet. That is the right time to deal with it, and not before, because the agents will have shown you exactly which records cause the trouble.

What goes wrong when companies build internal AI agents?

Six failures account for most of it. A process chosen for visibility rather than volume. Scope written as a sentence rather than three lists. No historical data to learn from, and no named reviewer. Going live on autonomy instead of on breadth. Measuring keystrokes saved rather than hours returned.

Starting too big is the most common and the most expensive. A programme that begins with "automate the back office" produces an architecture diagram and no working agent. A programme that begins with one document type from one customer produces something live in weeks, and the architecture emerges from it.

Underestimating the data requirement is second. Companies assume the model brings the knowledge, when in fact the model brings the reading and your records bring the answers. If your article master is a mess, the agent will surface that immediately, which is useful but not what anyone budgeted for.

Going live on autonomy rather than on breadth is the failure that damages trust fastest. An agent given full authority over a narrow slice will eventually meet an input nobody anticipated. No person was in the loop, so the first anyone hears of it is a wrong order at the customer. The same agent running supervised across a wider slice meets that input too, and a person catches it.

The subtlest failure is measuring the wrong thing. Keystrokes saved is a proxy that flatters any automation. The number that decides whether this was worth doing is hours returned to a named person, and whether those hours went into something better than typing. Topa Bathroom Products report 4 FTE moved off order entry into after-sales and service planning, which is a much harder claim to make and a much more honest one.

Frequently asked questions

How long does it take to build an internal AI agent?

For a common process on an existing product, a first agent handling real work in a narrow scope typically takes four to six weeks. Most of that is scoping, data access and parallel running rather than building. A custom build on your own infrastructure is a multi-quarter commitment, and the maintenance starts immediately after.

Do we need specialist AI engineers to build an internal AI agent?

Not for a bought agent. You need somebody who knows the process deeply, somebody who can grant access to the mailbox and the ERP, and a named reviewer. For a custom build you need engineers who will still be there in two years, because the maintenance burden is continuous rather than one-off.

What data does an internal AI agent need?

Historical records of the process being automated, plus the reference data it must check against: customers, articles, prices, specifications. Two years of processed records is a strong base. Without that history the agent has nothing to infer from, and you are back to writing rules by hand.

Can an internal AI agent work with an old ERP?

Usually yes, including systems with no documented connection, though the write path takes longer to build. SPL Treatments run a legacy desktop ERP with no connection available, and the agent still delivers the decision and the evidence for a person to enter. Read-only starting points are a legitimate first step.

Should the first internal AI agent be fully autonomous?

No. Start with every run reviewed by a named person and widen the input set before widening the authority. Autonomy should be earned against a measured error rate, not granted at go-live. The companies that got this wrong are the ones whose agents were switched off after the first bad week.

Build the first one against your own documents

The step that decides whether an internal AI agent works is the one before any building starts: choosing a process that begins with reading, and writing down where the agent's authority ends. Everything after that is execution, and execution is the part that has become straightforward.

Send us three weeks of real orders, invoices or specification requests in whatever state they arrive. We will show you which portion an agent resolves on its own, which portion it escalates, and what the first scope should be. Book a demo and we will work through your documents rather than a sample set.

Give your back office an AI workforce.