Build vs buy AI agents is the wrong first question for a back office. We think the real question is who owns the work after go-live, because a prototype that reads invoices takes a weekend and a production agent that reads them reliably takes integrations, exception handling, governance and a named owner.
Most teams arrive at this decision holding a demo that worked. Someone in finance or operations wired a model to a mailbox and watched it pull the right numbers out of 20 purchase orders. There is now a credible proof of concept and a budget conversation to open. The gap between that and an agent the order desk relies on every morning is where the money goes.
At Lleverage we build and run those agents for manufacturers, wholesalers and logistics operators, so we see both sides of this decision regularly, including the in-house builds that worked. This piece is the decision framework rather than the argument. For the case that the economics have shifted towards building, read our earlier piece on the build vs buy shift; to see what a running agent looks like, take a look at the platform or book a demo.
Should you build or buy AI agents for your back office?
Build when the process is your competitive advantage, the data is yours alone, and you can staff the maintenance indefinitely. Buy when the process is a well-understood back-office function that hundreds of companies run the same way, and where the time to production matters more than owning the code. Most back-office work is the second case.
That is a blunt rule, and it hides the interesting part. Buying in this market rarely means buying a finished product. Nothing on a shelf reads your customers' order emails, knows your article numbers and posts into your ERP configuration. The real choice is between building the whole thing with your own people, or buying an agent layer that someone else keeps running while you supply the process knowledge.
Our view is that this middle option is what most SMEs end up wanting once they price the alternatives honestly, because it puts the part that is genuinely yours, the rules and the exceptions, in your hands, and the part that is nobody's advantage, model plumbing and ERP connectors, in someone else's.
What does build actually mean for an AI agent?
Building means your organisation owns the whole stack: the model calls, the document parsing, the integration into every system the process touches, the exception routing, the audit trail, the evaluation harness that tells you whether last week's change made things better, and the people who maintain all of it. The prototype is roughly 10% of that list.
The parts teams underestimate are the unglamorous ones. Document intake has to survive the supplier who sends a photograph of a delivery note. The ERP connection has to handle a locked record, a failed post and a partial write without creating a duplicate order. Somebody has to decide what the agent does at 70% confidence. Somebody else has to be told when it took that decision.
There is also a staffing shape to this that gets missed during the enthusiasm phase. An agent in production needs a person who understands both the process and the system, available on the day it starts behaving differently because a customer changed their template. In a 200-person manufacturer that person already has a full-time job. The agent becomes their fifth priority in the week it needs to be their first.
What the build column actually contains
- Model access and cost control, including what happens when a long document blows the context budget.
- Document intake for every format the process really receives, not the 3 you tested with.
- Integration into the ERP, the mailbox, the document store and whatever else the process touches.
- Exception handling, approvals and the routing that decides which human sees what.
- An audit trail that satisfies your auditor, not just your curiosity.
- Evaluation: a way to tell whether a change improved results across your real backlog.
- Ongoing maintenance, which in our experience outweighs the original build within about 2 years.
What does buy mean when nobody sells your process?
Buying means acquiring the runtime, the integrations, the governance and the operating team, then configuring your process on top. What you are not buying is a finished agent for your order desk, because your order desk is not the same as anyone else's. The honest description is a managed agent layer with your business rules in it.
This matters for how you evaluate vendors. A demonstration of an invoice agent tells you almost nothing, because reading the invoice is the smaller half of the problem. Ask to see the exception path, the ERP write, the approval step and the audit record, on a document from your own backlog. Then ask what happens when the agent is unsure. Treat a confident answer to an ambiguous document as a red flag, not a feature.
The second thing you are buying is a rate of improvement. Agents get better because corrections accumulate, so the relevant question is whether a correction made on Tuesday changes the result on Wednesday, across every similar case, or only fixes that one document. That distinction decides whether year 2 costs less than year 1 or the same.
AI is a topic everywhere. The real challenge is creating concrete value with it. Together we tackled a recognisable problem and translated it quickly into a practical solution that works right away.
Robert Broeckmans, Finance and IT Director at SIG Benelux, on a project that automated order intake into Dynamics 365 in 4 weeks.
What does each option really cost?
Build costs are mostly people and mostly recurring, while buy costs are mostly subscription and mostly predictable. The comparison only becomes useful when you count the second and third year, because a build reaches parity on paper in year 1 and then carries a maintenance load that a subscription absorbs.
| Cost line | Build in-house | Buy an agent layer |
|---|---|---|
| Time to first production process | MIT NANDA puts enterprise pilot-to-implementation at 9 months or longer | MIT NANDA puts the fastest mid-market companies at 90 days |
| Up-front spend | Engineering time, model credits, infrastructure | Implementation scoped into the engagement |
| Recurring spend | Salaries for the people who keep it running | Monthly price per agent, integration included |
| Odds of reaching production | MIT NANDA 2025: internal builds fail roughly twice as often | MIT NANDA 2025: external partnerships see twice the success rate |
| Who fixes it at 08:00 | Your engineer, if they are not on leave | The vendor's team, under an agreement |
| Cost of process 2 | Lower than process 1, if you reused the plumbing | Lower, because the knowledge layer is already filled |
| Risk you carry | Model changes, integration drift, key-person departure | Vendor viability, contractual lock-in |
Our own commercial model is per agent rather than per seat, with integration, AI usage and ongoing improvement inside the monthly price; the pricing page sets out how an engagement is scoped. We mention the shape because build cases are routinely argued against a per-seat licence number, and per-seat is not how agent work is priced.
Those timings are not ours. MIT NANDA's "The GenAI Divide: State of AI in Business 2025", built on 52 structured interviews, 153 survey responses and 300 public implementations, found the top mid-market performers averaging 90 days from pilot to full implementation while enterprises took 9 months or longer. Our reading is that the difference is scope discipline rather than engineering skill, and scope discipline is available to anyone.
The number that decides most cases is not in the table at all. It is how many processes you intend to automate. A single process is the worst possible economics for a build, because all the plumbing gets amortised across one workflow. At 5 or 6 processes the calculation changes, and a company with a real engineering function and a multi-year plan can make building pay.
Why do back-office agents fail after the prototype?
Back-office agents fail in production for 4 recurring reasons: the real document mix is messier than the test set, the system of record resists being written to, exceptions have no owner, and the knowledge the agent needs was never written down. None of them are model problems, which is why a strong prototype predicts so little.
The attrition is not a rounding error. MIT NANDA's 2025 report found that among organisations evaluating enterprise-grade AI systems, 60% evaluated, 20% reached a pilot and just 5% reached production, and it attributes the losses to brittle workflows, lack of contextual learning and misalignment with day-to-day operations rather than to model quality. Gartner, in a June 2025 press release, predicted that more than 40% of agentic AI projects will be cancelled by the end of 2027 on escalating costs, unclear business value and inadequate risk controls. Both describe the distance between a demo and a process, which is what this decision is actually about.
The document mix is the first shock. At SIG Benelux, roughly 3% of inbound order lines carried an article number and about 47% arrived as a description only, with the remainder arriving as messages, photographs or sentences that only made sense to someone who knew the customer. An agent tested on clean purchase orders meets that on its first morning.
The system of record is the second. At SPL Treatments the ERP is a legacy desktop application with no interface to write to at all, so the workflow was designed to produce an output a person carries across rather than to post directly. That constraint is common in industrial operations, and it has to be designed around rather than discovered in month 4.
The third is exception ownership. An agent that handles 80% of a process creates a new queue containing the other 20%, and that queue needs a named owner and a route. Our position is that an agent should surface what it cannot resolve rather than guess, and that the measure of a good implementation is what it does with the hard cases, not the share it clears silently. That is the difference between supervised agents and background automation, and both belong in a serious back office.
The fourth is knowledge, and MIT NANDA names it as the single biggest barrier to scaling: not infrastructure, regulation or talent, but learning, because most systems do not retain feedback, adapt to context or improve over time. The rules that make a process work usually live in the heads of 2 or 3 long-serving people, and neither a build nor a buy succeeds without extracting them. The recurring good news is that the history often holds them already: at SIG Benelux the years of order history in Dynamics 365 turned out to contain the translations the sales desk was doing by hand, which is why 88% of order lines now match the correct article on the first try.
When does building in-house make sense?
Building makes sense when the process is a genuine differentiator, when you have an engineering function that will still be there in 3 years, and when you plan enough processes for the shared work to pay off. A pricing engine that encodes a commercial strategy nobody else has is worth owning. An accounts payable agent almost never is.
Two further conditions matter more than they get credit for. The first is data gravity. If the process runs on data that cannot leave your estate for regulatory or contractual reasons, and no vendor can operate inside it, building may be the only route. The second is organisational patience. An in-house build judged on quarterly delivery gets cancelled in month 5, halfway through the integration work that makes it useful.
Be honest about the alternative use of that engineering time as well. In most manufacturers and distributors the engineering function is small and already committed to the product or to keeping the ERP estate alive. A build is not free just because the salaries are already on the payroll; it is paid for with whatever those people were going to do instead.
There is a halfway position worth naming, because it is where several of the in-house builds we have seen actually succeed. The company buys the runtime and the connectors, then staffs a small internal team to own the process design, the rules and the exception queue. They are not writing document parsers or ERP adapters, but they are unambiguously the owners of how the process behaves, and they can change it on a Tuesday afternoon without raising a ticket. If your reason for wanting to build is control rather than cost, that arrangement gives you the control without the maintenance bill.
When is buying the right call?
Buying is right when the process is standard back-office work, when the time to production matters, and when you would rather own the business rules than the runtime. That covers most order intake, invoice processing, document validation, order confirmation and customer-question handling, which is to say most of the work that keeps a back office busy.
The strongest external evidence points the same way. MIT NANDA's 2025 report lists "external partnerships see twice the success rate of internal builds" as one of the 4 patterns defining its GenAI Divide, and states it in reverse as well: internal builds fail twice as often. We read that as a finding about who carries the integration and the learning loop, not as evidence that vendors are cleverer than your engineers. It is still the number a build case has to argue against.
| Signal | Points towards build | Points towards buy |
|---|---|---|
| The process is a competitive differentiator | Yes | No |
| You plan 1 or 2 processes total | No | Yes |
| The data cannot leave your estate under any arrangement | Yes | No |
| The ERP is old and awkward to integrate with | No | Yes |
| You need something running this quarter | No | Yes |
| The rules change often and operations must change them | Either, with the right design | Yes |
Speed is the signal people discount and then regret. A 4-week engagement that puts one process into production changes the internal conversation in a way a 9-month build does not. The organisation sees the thing working before it is asked to believe in it. The SIG Benelux work went live on real orders in June 2026 after a 4-week project scoped to one location, one process and one mailbox.
There is an integration argument too. Most of the difficulty in a back-office agent is on the ERP side, and connectors are the least differentiating code you will ever write. Our integrations cover SAP, Business Central, Exact, AFAS and Infor, and the work that goes into keeping a native connection healthy through vendor upgrades is the sort of thing that is cheap to share and expensive to own alone.
How do you decide?
Run the decision as 6 questions about one specific process, not as a strategy debate about AI. The answers usually point the same way within an afternoon, and the exercise is worth doing per process rather than once for the company, because the right answer for pricing strategy is frequently not the right answer for invoice matching.
- Is this process something customers would notice us doing differently from a competitor? If not, nobody gains from you owning the code.
- Who is on call the morning it breaks, by name, and what else is that person responsible for?
- What does the real document mix look like across 3 months, not 3 samples?
- Can we write into the system of record, and what happens on a partial or failed write?
- Who owns the exception queue, and what is their turnaround commitment?
- How many processes are we going to do this to in the next 2 years?
Question 6 is the one that most often flips the answer. A company automating 1 or 2 processes should almost always buy, because the fixed cost of a build lands on a very small base. A company with 8 processes and a real engineering team has a defensible case either way, and the deciding factor becomes how quickly the first one has to be live.
A worked example
A 180-person industrial manufacturer wants to stop retyping purchase orders. Volume is about 60 orders a week, arriving as PDFs and email prose from 40 regular customers. The ERP is on-premise with a usable interface. The differentiator test fails immediately: nobody buys from them because of how they type orders. The on-call test fails too. Their one internal developer already maintains the ERP integrations that exist.
That points to buying, and the shape of the result is familiar. At Microtechniek, a mechanical engineering firm of 130 people, between 45 and 71 purchase orders a week are now written into Ridder IQ without anyone typing them, giving the admin desk back more than 500 hours a year. Orders that arrive on a Friday evening are in the system before Monday, which is a benefit no business case predicted and everyone noticed.
Change one variable and the answer changes with it. If that same manufacturer had 6 processes queued behind this one, an engineering team of 8, and a 3-year mandate, building a shared layer and using this process as the first tenant becomes a reasonable decision. What does not work is deciding to build because the prototype was impressive.
Frequently Asked Questions
Is it cheaper to build or buy AI agents?
Buying is usually cheaper for the first 1 or 2 processes, because a build has to pay for integrations, exception handling and governance before any process benefits. Building can be cheaper across several processes if the shared work is genuinely reused and the maintenance staffing holds. Count year 2 and year 3, and weigh them against MIT NANDA's finding that internal builds fail roughly twice as often.
How long does it take to build an AI agent in-house?
A working prototype takes days. Production is the slow part: MIT NANDA's 2025 report puts enterprise pilot-to-implementation at 9 months or longer, against 90 days for the fastest mid-market companies. Most of that elapsed time goes into integration and the unhappy paths rather than the model work, which is why a convincing demo tells you very little about the schedule.
What is the biggest risk of building AI agents in-house?
Key-person dependency. In-house agents are commonly built by 1 or 2 enthusiasts, and when they move on the organisation is left running a production process nobody fully understands. The second risk is drift: models, document formats and ERP versions all change, and an unmaintained agent degrades quietly rather than failing loudly.
Do we lose control of our business rules if we buy?
Not if the rules live in a layer you can edit. Ask where the business rules are stored, who can change them without a vendor ticket, and whether corrections made by your own team change future results. The rules and the exception knowledge are yours in any sensible arrangement; the runtime and the connectors do not need to be.
Which back-office process should we start with?
Start with the one that is high volume, painful and bounded, which is usually order intake or invoice processing. Avoid starting with the process that is most interesting or most broken, because both tend to be broken for reasons an agent cannot fix. A narrow first process that reaches production beats a broad one that reaches month 7.
Where we land
Build if the process is genuinely yours and you can staff it for years; buy if it is ordinary back-office work and you want it running this quarter. Most of the work in a back office is ordinary by design, which is why we think the buy case wins more often than the build enthusiasm in the room suggests, and why the companies that build successfully are usually the ones doing it for their fifth process rather than their first.
One more finding belongs in the budget conversation. MIT NANDA reports that roughly 50% of GenAI budgets go to sales and marketing while back-office automation, which often yields better returns, stays underfunded, and it puts that down to easier metric attribution rather than to actual value. Our reading is that the build-versus-buy argument is frequently being had about the wrong process to begin with.
If you want to test the decision against a real process rather than a slide, bring one: the document mix, the ERP, the exception rules and the volume. We will tell you where it lands, including when the answer is that you should build it yourself. Book a demo and bring the awkward documents.
