Custom AI agent development, for the parts of your business where somebody currently has to read something, decide, and then go and change a system. Our engineers build the agent inside your own tools, and you set exactly how far it can go before it needs a person to say yes.
You choose the models it runs on and the ground it runs on. Open weights on your own hardware, a closed API, or both inside the same build.
You do not need to be a HiQBot customer to hire the team. Last updated .
Ninety seconds, on a Sunday night. The enquiry that fitted nothing on the books waited for a person on Monday, with the working attached.
Custom AI agent development is the practice of building an AI system that makes a decision and then carries it out inside your own software, instead of answering a question and leaving the work to a person. An agent of this kind reads whatever arrives, a message, a document, a form, a call, works out what it means against your rules and your records, calls the systems that hold the answer, checks what came back, and then either finishes the job or stops and asks someone. The difference from a chatbot is the action: a chatbot can tell a customer how to get a refund, an agent issues it, updates the order and writes the note. The difference from ordinary automation is judgement. A script needs the input to arrive in the shape it expected. An agent does not.
HiQBot is our own product. The agent inside it handles live customer conversations on website chat, WhatsApp and Instagram: it reads the customer’s documents, calls tools that book meetings and take orders, and hands the thread to a person at the point it should stop.
Retrieval, tool calls, memory, guardrails and the handover rules are things we have had to get right with real customers watching, not things we have read about. The engineers who would build yours are the ones who built that and carry the pager for it.
You can go and look at the product before you speak to anyone here, which is more than most agencies can offer you.
This is the map of the agentic AI work businesses are actually commissioning in 2026, grouped by the department that feels the pain. Most real builds are two or three of these wired together. What separates an agent that saves an hour a week from one that takes a whole process off your desk is how deep it is allowed to reach into the systems that hold the answer.
Agents that go and find the work, not agents that wait for it.
Researches the account properly, decides who is worth contacting and when, writes outreach that references something real, runs email, LinkedIn and WhatsApp together, handles the reply, and books the meeting. The CRM is updated because the agent updated it.
Reads the intent behind an enquiry, asks the two questions that decide it, then routes on territory, product, language or who actually has capacity today. Instant booking beats a callback queue every time.
Watches tenders, funding rounds, hiring, planning applications, job changes, whatever your trigger is, then builds the list and drafts the approach while the signal is still warm.
Turns an emailed request into a priced quote against your catalogue, stock and discount rules, chases the missing spec, and puts anything outside the rules in front of a person.
Briefs the rep before the call, drafts the proposal after it, writes the CRM notes nobody writes, and chases the follow-up that quietly decides half your pipeline.
Including the part of search that is now answers rather than links.
Finds the queries and the citation gaps where your brand is missing from AI answers, builds the brief, drafts against it, handles internal linking, and reports on both rankings and how often the assistants quote you.
Repurposes one piece into the formats each channel wants, localises it, keeps it inside your brand voice, and pushes it to the CMS instead of a document somebody has to move.
Segments on behaviour rather than a static list, writes the variant for each segment, watches what happens, and changes the next send based on it.
Monitors mentions, reviews and communities, drafts the reply in your voice, and escalates the ones where a founder should answer personally.
Deflection-only bots settle around 30 to 50 per cent of contacts. Agents wired into the systems that hold the answer reach 70 to 85 per cent on well-scoped work.
Not a knowledge-base search with a chat window. It reads the account, the order and the history, then issues the refund, reships the item, changes the booking or cancels the subscription inside the rules you set, and opens a ticket with the full working when it should not act alone.
One agent behaving the same on your site, WhatsApp, Instagram or in-product, with a handover that carries the whole thread so nobody asks the customer to start again.
Takes the call or places it, copes with interruptions and accents, and files the outcome against the record. Warm transfer hands a person the context, not a transcript to skim.
Works out which product, plan, policy or route actually fits from what the customer said, explains why, and stays inside what you are allowed to recommend.
Works a shared mailbox the way a good coordinator does: sorts what came in, answers what it can, files the rest against the right account, and surfaces anything with a deadline.
Traditional QA samples one to three per cent of interactions, so 97 to 99 per cent are never looked at. An agent scores all of them.
Scores every chat, call, email and ticket against a rubric you own, on voice and text alike, and shows you the ones worth watching instead of a random sample.
Turns those scores into something a team lead can use on Monday: the clip, the moment it went wrong, and what good looked like on the call that went well.
Watches for the disclosure that was not read, the promise nobody is allowed to make, and the personal data that ended up somewhere it should not be.
Sits beside your human team in real time with the answer, the next step and the wording, rather than replacing them.
Queues somebody works through by hand, where the agent touches the record and a wrong move costs money.
Invoices, claims, applications, contracts, packing lists. Reads whatever shape it arrives in, checks it against your system of record, and flags the disagreement rather than averaging over it.
Matching, reconciliation, chasing what is overdue, coding the transactions your bookkeeper recodes every month, and the procurement round trip from request to purchase order.
Watches the exceptions on the order, the load or the job, works out what to move and who to tell, and updates the systems that are supposed to agree with each other.
Screens against the requirement rather than the keyword, answers the policy questions HR retypes weekly, schedules across three calendars, and runs a joiner or a leaver through the checklist.
Access requests, provisioning a new starter, triaging the ticket, running the runbook step nobody remembers, and escalating the ones that touch production.
Notes that turn into actions with owners, follow-ups that actually get sent, and scheduling across calendars that never agree.
Jobs measured in days. It waits on an approval, survives a system going down over the weekend, retries what failed, and never quietly drops an item on the floor.
The questions that currently need the one person who has been there seven years.
Answers across contracts, tickets, wikis, drives and the shared folder nobody has opened since 2019, with a citation behind each claim and a plain "I do not know" when the answer is not there.
Your own staff ask the business a question and get a real answer: the numbers, the policy, the last three cases like this one, and a draft of whatever they were about to type.
Watches the sources you would read if you had the time, regulations, competitors, prices, tenders, and tells you what changed and what it means for you.
Answers a question against the warehouse, shows the query it ran, and flags the number that moved before anyone opens the dashboard.
The part most quotes leave out, and the part that decides whether the thing survives its second month.
Rules on what an agent may reach and what it may do alone. Anything above the line stops and waits for a named person in Slack or email, with the reasoning laid out.
The cases, the grader and the scores that tell you an agent got worse before your customers do, and tell you when it has earned a longer leash.
When one job needs a researcher, a decider and a doer, they run as one system with one trace and one owner. Not four bots emailing each other.
The two figures above are 2026 industry benchmarks rather than our own numbers: resolution rates by integration depth are from Lorikeet’s 2026 benchmark study, and QA sampling coverage from Mihup’s analysis of automated agent scoring. We quote them because they set a realistic expectation, not because they flatter us. An agent given read-only access lands at the bottom of that range, and integration depth is what moves it.
Some of this you can have without a build at all. Our own platform ships an AI sales agent with lead qualification on every plan, and live chat with human handover plus the omnichannel inbox from the second tier up. If a plan covers the job, we will point you at the plan and take no money for a build. An engagement is for the work the plan was never shaped for.
A model with no reach can draft an answer. Give it your systems and it can finish the job. Read access first, write access once the numbers say it has earned it.
We build against anything with an API, anything that speaks MCP, and a fair number of systems that only expose a database. If yours is a twenty-year-old package with a green screen and no documentation, say so on the call. It is usually still doable, and it is always more work.
This is the part that decides whether a build survives contact with real customers. Every action an agent can take sits in one of three places, and where the line falls is a setting you control rather than something you find out about afterwards. It is also the pattern that production deployments converged on through 2026: autonomous execution for routine work, a named human approving anything high-stakes, and a queryable audit log behind both, as surveys of agent control planes describe it.
| Level | What sits here | Who decides |
|---|---|---|
| It acts on its own | The routine majority: answering, filing, booking, updating a record, issuing the refund that sits inside the policy. Reversible, low value, high volume. This is where the money is. | The agent |
| It proposes, a person approves | Anything expensive, contractual or awkward. The agent does the work and presents the decision with its reasoning, and a named person clicks yes in Slack or in email. Faster than doing it yourself, safer than letting it run. | A named approver |
| It stops and hands over | Low confidence, an angry customer, or anything outside what it was given. The person receives the whole thread and the working, not a ticket saying the bot could not help. | Your team |
Underneath all three, every run keeps a trace: what came in, which tools the agent called with what arguments, which rule allowed or blocked each call, who approved what, and what it produced. That is what you show an auditor, and it is what you read at 9am when something went sideways at 2. Where the line sits is not a personality you hope the agent has, it is a rule per action, and it moves. Most builds start with almost everything requiring approval, because that is how the team learns whether they trust it, and the evaluation numbers are what justify widening it later. Tightening it again after a bad week is a settings change rather than a redeploy, which matters more than it sounds like it should.
Something arrives, somebody reads it, decides, and updates a system. That is the same job in a lettings office and a claims department. What we ask for is one person who knows how the decision is really made, including the parts nobody wrote down.
Qualifies an enquiry on budget, area and timing, matches it against what is actually on the books, and puts the viewing in the diary of whoever covers that patch. Chases the ones that go quiet, and keeps the landlord and tenant threads apart.
Screens against the requirement rather than the keyword, answers the candidate questions your team retypes every week, schedules around three calendars, and walks a joiner through onboarding without anyone chasing a form.
Triages a request against your own protocol, books it into the correct clinic slot, handles the reschedules and the reminders, and puts anything clinical in front of a human with the notes already gathered.
Reads the application or the claim with its attachments, checks it against policy, clears the clean ones inside the limits you set, and hands the edge cases to an underwriter with the reasoning written down and auditable.
Reads the contract, the tender or the brief, pulls out obligations, dates and unusual clauses, compares against your standard position, and drafts the response for someone qualified to sign off.
Works out dates, party size and budget before a person is involved, holds the option, upsells where it is welcome, and rebuilds an itinerary at midnight when a leg gets cancelled.
Watches exceptions on the load or the job, works out what to reschedule and who to move, updates the systems that need to agree, and tells the customer before the customer calls you.
Answers against the live order rather than a generic policy page, decides refunds, exchanges and reships inside the rules you wrote, and escalates the ones where a person is better placed.
Turns an emailed request for quote into a priced response against your catalogue and stock, chases the missing spec, and flags the orders that will not ship on time before the delivery date does.
Handles admissions questions, applications and timetable clashes, tracks who has not completed what, and gives a tutor the summary before the meeting instead of after it.
A scheduled script with twenty lines of conditions is cheaper to run, cheaper to fix at 3am, and it never invents anything. Agents earn their cost on messy input and judgement calls: a document that turns up in a different shape every time, a customer who buries the real question in paragraph four, an exception nobody wrote a branch for, a decision that depends on three systems disagreeing with each other. The useful question is not whether AI could do the task, it is whether the task varies. Where the input arrives in the same shape every time and the rules fit on a page, ordinary automation wins on cost, on predictability and on the amount of sleep your team gets. We will say so on the first call if what you have described is a workflow rather than an agent, and sometimes the honest answer is an n8n flow we set up in a week that you never pay for again.
We are not loyal to one of these, and you should be suspicious of a shop that is. If your team already lives in n8n, we build into your n8n rather than standing a second system up beside it. If the work wants the multi-agent shape CrewAI is good at, or a graph in LangGraph, or one of the vendor agent SDKs, that is a scoping decision we make with you and explain.
A fair amount of this work ends up as ordinary Python with a state machine and no framework at all, and the boring option wins more often than people expect. Where a system you already run speaks MCP, we connect through that instead of writing one more adapter for you to maintain.
Which one handles which step is a cost and accuracy question, and we settle it with measurements rather than opinions.
| Closed modelsClaude, GPT, Gemini | Open-weight modelsLlama, Qwen, Kimi, GLM, DeepSeek, Mistral | |
|---|---|---|
| Hardest reasoning | Still ahead on the genuinely ambiguous cases, which is why we keep one for the step where judgement decides the outcome. | Trails, and the gap has narrowed with every release cycle. Fine for extraction, classification and routing. |
| Where your data goes | Out to the provider, under a contract. Enough for most commercial work, not enough for some regulators. | Wherever you put it, including a rack in your own building. Nothing crosses a border you did not choose. |
| Cost shape | Per token. Cheap to start, and it scales with your success whether you like it or not. | GPU time. Higher floor, and the cheaper shape once volume is high. |
| What you operate | Nothing. No hardware to buy, no inference to keep up. | Real infrastructure. Somebody has to own the GPUs, the serving stack and the upgrades. |
| Changing your mind | A configuration change and a re-run of the evaluation set. | The same, because the model sits behind our interface rather than through the code. |
Most builds end up mixed rather than picking a side. An open model does the high-volume extraction, classification and routing, where the work is repetitive and the cost per call is what decides whether the project pays for itself, and a closed model is kept for the one step where the judgement is genuinely hard. Because the model sits behind our own interface rather than threaded through the code, swapping one for another is a configuration change and a re-run of the evaluation set instead of a rewrite. That matters more than which model is ahead this quarter: prices fall, licences change, a provider has a bad month, and a better open release lands every few weeks. The build should be able to take advantage of all of that without you paying for it twice.
Where the agent runs is your decision and the first thing we settle in discovery, because it changes what the rest of the build can look like.
Personal data is masked before anything reaches a model. You set retention, and zero is one of the settings. What an agent may read is scoped the same way as what it may do, so it cannot answer from a folder it was never meant to open. Access is by role, and your material is not training data for us or for anyone we call.
An agent that demos beautifully and falls over in week three is the normal outcome. What prevents it is writing down what a correct answer looks like before the build starts, then running every change against those cases.
You get the cases, the grader and the scores. When somebody edits a prompt six months from now, your team can tell whether the agent got better or only got different. It is also what tells you when the agent has earned a longer leash, and what lets you move to a cheaper model without crossing your fingers.
Most AI agent development quotes are a guess wearing a number. Four steps, and you can stop after any of them.
Tell us what the process is, which systems it touches, and what a wrong answer costs you. You leave that call knowing whether this should be an agent at all.
We map the decision, the data behind it and the integration points, agree where the approval line sits, and write the evaluation cases with your team. You get a fixed scope and a price before anyone writes agent code.
A working agent on your real data, measured against those cases every time it changes. A pilot group is usually running it within weeks, with the agent on a short leash while the numbers come in.
We roll it out, widen what it is allowed to do as it earns the trust, and keep it healthy. Or we hand over the repository, the prompts, the evaluation set and the deployment scripts, and your engineers carry on without us.
An AI agent is a system that takes a goal, works out the steps itself, calls real software to carry them out, and checks the result before deciding what to do next. In a business it usually sits where a person currently reads something and then updates a system: an enquiry, an invoice, a claim, a ticket, a call. The test is whether it can complete the task, not whether it can describe it.
RPA follows a recorded path and breaks when the screen, the form or the wording changes. An agent reads the input, works out what it means, and picks the path, so a document arriving in a shape nobody anticipated is a normal Tuesday rather than an incident. RPA is cheaper and more predictable where the input never varies. Where it does vary, RPA becomes a maintenance bill.
A chatbot picks a reply. An agent picks an action, runs it against a real system, looks at what came back, and decides what to do next. That is the difference between telling a customer how to get a refund and issuing it, and it is why deflection bots plateau around a third of contacts while agents wired into the backend get past two thirds.
That is your decision and we build it as a setting, not a personality. Each action gets a rule: do it, propose it and wait for a named person, or refuse. Value thresholds, customer tier and confidence all feed into it, and you can tighten the whole thing in an afternoon if something goes wrong.
Through your APIs where they exist, through MCP where a system already speaks it, and through the database where a vendor never built an API. Read access first, write access when the evaluation numbers say it has earned it. The integration depth is what decides how much the agent can actually finish.
Discovery runs one to two weeks. A first agent doing real work for a pilot group is usually a few weeks after that. Getting to full production is mostly a question of how many systems it touches and how careful the sign-off has to be, so that part varies.
We price the build after discovery and you see the number before we start. What moves it is the number of systems the agent integrates with and how expensive a wrong answer is, because that decides how much review, testing and guarding goes in.
Probably. The shape repeats: something arrives, someone reads it, decides, and updates a system. The vocabulary and the rules change, the shape does not. What we ask for is one person who knows how the decision is really made, including the parts that are not written down anywhere.
Yes, agentic AI is the umbrella term and a custom agent is what you get at the end of it. The word matters less than the test behind it: can the thing choose an action, run it against a real system, look at the result and decide what to do next. If it can only pick a sentence to say, it is a chatbot with better marketing.
No. Under the API terms of the providers we use, business inputs are not used to train their models, and we do not train on your material either. When a contract is not enough for your regulator, we run open-weight models on hardware you control and nothing goes out to an API at all.
Yes. Open-weight models such as Llama, Qwen, Kimi, GLM or Mistral run inside your network, with the agent, the vector store and the logs sitting there alongside them.
LangGraph, LangChain, CrewAI, n8n, MCP and the vendor agent SDKs, picked per project instead of pushed at everybody. Plenty of builds end up as ordinary Python with a state machine, because a framework is not always worth the ceiling it puts on you. If your team has already standardised on something, we work in that.
You find out from the evaluation set rather than from a customer, which is the whole point of shipping one. Every run keeps a trace of its input, the tools it called with what arguments, the rule that allowed or blocked each call, who approved what, and what it produced, so a bad outcome is a readable sequence rather than a mystery. Tightening what the agent may do alone is a settings change, not a redeploy.
Against cases written before the build starts, with your team, describing what a correct outcome looks like for your process. Every change runs against them, so you can see whether an edit made the agent better or only different. Resolution or completion rate, escalation rate and the cost per completed task are the numbers that decide whether it stays switched on.
Yes, and that is usually where the value is rather than in the model. We integrate through your APIs where they exist, through MCP where a system already speaks it, and directly against the database where a vendor never built an API. Read access comes first and write access follows once the evaluation numbers say it has earned it.
Yes. The code, the prompts and the evaluation set are yours at handover.
Half an hour, an engineer rather than a salesperson, and a straight answer on whether this is worth building. If it is not, we will say that too.
Tell us what you want built. We reply within one business day.