Part of our guide hub: AI solutions and automation
Every business we talk to in 2026 has been pitched an AI agent. Very few have been told where agents fail. That gap is expensive, because the failure mode is not a crash: it is an agent doing the wrong thing confidently, at scale, for three weeks before anyone notices.
Here is the map we use with clients before scoping any agent work.
The property that decides everything
Agents perform well on tasks where verification is cheaper than execution. If checking the output takes a fraction of the effort of producing it, an agent is a strong bet: a human stays in the loop at low cost, and the economics work. If verifying the output costs as much as doing the work yourself, an agent adds a review burden instead of removing a workload.
Draft a supplier email. Verification is a five-second read. Reconcile a month of bank transactions. Verification means redoing the reconciliation. That single test predicts most pilot outcomes we have seen.
Where agents are genuinely working now
- First-pass document extraction. Invoices, purchase orders, CVs, delivery notes, ID documents. Accuracy is high, errors are visible, and a confidence threshold routes the uncertain cases to a person. This is the single most reliable ROI in operations today.
- Tier-one support triage. Not answering everything: classifying, tagging, pulling the relevant account context and drafting a reply for an agent to approve. Handle time drops materially; resolution quality stays human.
- Internal knowledge retrieval. Answering staff questions from your own policies, SOPs and past tickets. Wrong answers are caught immediately by the person who asked, which is exactly the cheap-verification property.
- Code and test scaffolding. Boilerplate, test cases, migration scripts. Our own engineering teams use it daily; the compiler and test suite are the verifier.
- Sales research and outreach drafting. Account briefs, personalisation, follow-up sequencing. With a human sending.
Where they quietly break
- Anything that must reconcile to a number. Financial close, stock counts, payroll. Agents produce plausible numbers, and plausible is worse than wrong. Wrong gets caught.
- Multi-step processes with no checkpoint. Error compounds. An agent chaining eight actions with 95% step accuracy is right about 66% of the time end to end. Insert human checkpoints or cap the chain length.
- Decisions with regulatory or safety consequence. Clinical, credit, legal, HR terminations. Not because the model cannot reason, but because you need an accountable decision-maker and an audit trail a regulator accepts.
- Anything depending on tacit context nobody wrote down. If the correct answer lives in the head of one operations manager and has never been documented, the agent will not find it. Document first, automate second.
- Low-volume, high-variance work. If it happens eleven times a year and looks different each time, the automation will cost more than the work.
The four questions before any pilot
- What is the volume? Under a few hundred instances a month, the build rarely pays back. Automate the thing that happens 4,000 times, not the thing that happens 40 times.
- What does a wrong answer cost? Multiply by expected error rate. If that number is uncomfortable, you need a human checkpoint, not a better prompt.
- Who owns the output? Every agent needs a named human accountable for its decisions. If no one will sign, do not deploy.
- How will you know it degraded? Model behaviour, your data and your process all drift. Without sampling and monitoring you will find out from a customer.
How to pilot without risk
Run the agent in shadow mode for three to four weeks. It processes real work and produces real output, but a human still does the job and nobody sees the agent output except you. At the end you have a measured accuracy rate on your data instead of a vendor benchmark, and the decision to go live becomes arithmetic rather than faith.
Shadow mode also surfaces the thing demos never show: the 15% of your real inputs that are messy, incomplete or formatted in a way nobody mentioned.
The honest summary
AI agents are excellent at producing a first draft of structured work and poor at being the final answer. Design the process around that and they pay for themselves. Design the process assuming they are the final answer and you will spend more on cleanup than you saved.
Ezitech builds applied AI systems into ERP, support and document workflows for clients across 10+ industries. If you want a pilot scoped honestly (including being told not to build it), talk to our team.
