Custom AI agents that read the case, decide, and do the work.

Custom AI agent development, for the parts of your business where somebody currently has to read something, decide, and then go and change a system. Our engineers build the agent inside your own tools, and you set exactly how far it can go before it needs a person to say yes.

You choose the models it runs on and the ground it runs on. Open weights on your own hardware, a closed API, or both inside the same build.

You do not need to be a HiQBot customer to hire the team. Last updated .

One agent run, an estate agency at 11pm
01Read the enquiry3 bed, 450k, wants to view this week
02Matched 4 live listings2 inside budget and the catchment
03Checked the diaryAdam covers that patch, free Tuesday at 4
04Sent both, booked the viewingwrote it back to the CRM

Ninety seconds, on a Sunday night. The enquiry that fitted nothing on the books waited for a person on Monday, with the working attached.

What custom AI agent development actually means

Custom AI agent development is the practice of building an AI system that makes a decision and then carries it out inside your own software, instead of answering a question and leaving the work to a person. An agent of this kind reads whatever arrives, a message, a document, a form, a call, works out what it means against your rules and your records, calls the systems that hold the answer, checks what came back, and then either finishes the job or stops and asks someone. The difference from a chatbot is the action: a chatbot can tell a customer how to get a refund, an agent issues it, updates the order and writes the note. The difference from ordinary automation is judgement. A script needs the input to arrive in the shape it expected. An agent does not.

The loop a custom AI agent runs on every item of workFour stages in sequence: read the input, interpret it against your rules and records, act by calling your systems, then verify the result. Verification either loops back to interpret, or escalates the case to a person.ReadInterpretActVerifynot right yet, or new information arrivedoutside its remit, escalate to a person

We already run one of these in production.

HiQBot is our own product. The agent inside it handles live customer conversations on website chat, WhatsApp and Instagram: it reads the customer’s documents, calls tools that book meetings and take orders, and hands the thread to a person at the point it should stop.

Retrieval, tool calls, memory, guardrails and the handover rules are things we have had to get right with real customers watching, not things we have read about. The engineers who would build yours are the ones who built that and carry the pager for it.

You can go and look at the product before you speak to anyone here, which is more than most agencies can offer you.

Anything that comes down to a decision somebody keeps making by hand.

This is the map of the agentic AI work businesses are actually commissioning in 2026, grouped by the department that feels the pain. Most real builds are two or three of these wired together. What separates an agent that saves an hour a week from one that takes a whole process off your desk is how deep it is allowed to reach into the systems that hold the answer.

Revenue and pipeline

Agents that go and find the work, not agents that wait for it.

AI SDR and outbound agents

Researches the account properly, decides who is worth contacting and when, writes outreach that references something real, runs email, LinkedIn and WhatsApp together, handles the reply, and books the meeting. The CRM is updated because the agent updated it.

Inbound qualification and routing

Reads the intent behind an enquiry, asks the two questions that decide it, then routes on territory, product, language or who actually has capacity today. Instant booking beats a callback queue every time.

Business development and signal agents

Watches tenders, funding rounds, hiring, planning applications, job changes, whatever your trigger is, then builds the list and drafts the approach while the signal is still warm.

Quoting and deal-desk agents

Turns an emailed request into a priced quote against your catalogue, stock and discount rules, chases the missing spec, and puts anything outside the rules in front of a person.

Sales copilots

Briefs the rep before the call, drafts the proposal after it, writes the CRM notes nobody writes, and chases the follow-up that quietly decides half your pipeline.

Marketing and search

Including the part of search that is now answers rather than links.

SEO and GEO agents

Finds the queries and the citation gaps where your brand is missing from AI answers, builds the brief, drafts against it, handles internal linking, and reports on both rankings and how often the assistants quote you.

Content operations agents

Repurposes one piece into the formats each channel wants, localises it, keeps it inside your brand voice, and pushes it to the CMS instead of a document somebody has to move.

Campaign and lifecycle agents

Segments on behaviour rather than a static list, writes the variant for each segment, watches what happens, and changes the next send based on it.

Listening and reputation agents

Monitors mentions, reviews and communities, drafts the reply in your voice, and escalates the ones where a founder should answer personally.

Customer conversations

Deflection-only bots settle around 30 to 50 per cent of contacts. Agents wired into the systems that hold the answer reach 70 to 85 per cent on well-scoped work.

Support resolution agents

Not a knowledge-base search with a chat window. It reads the account, the order and the history, then issues the refund, reships the item, changes the booking or cancels the subscription inside the rules you set, and opens a ticket with the full working when it should not act alone.

Live chat and omnichannel agents

One agent behaving the same on your site, WhatsApp, Instagram or in-product, with a handover that carries the whole thread so nobody asks the customer to start again.

Voice agents

Takes the call or places it, copes with interruptions and accents, and files the outcome against the record. Warm transfer hands a person the context, not a transcript to skim.

Advisory and recommendation agents

Works out which product, plan, policy or route actually fits from what the customer said, explains why, and stays inside what you are allowed to recommend.

Inbox agents

Works a shared mailbox the way a good coordinator does: sorts what came in, answers what it can, files the rest against the right account, and surfaces anything with a deadline.

Quality, coaching and compliance

Traditional QA samples one to three per cent of interactions, so 97 to 99 per cent are never looked at. An agent scores all of them.

Conversation evaluation agents

Scores every chat, call, email and ticket against a rubric you own, on voice and text alike, and shows you the ones worth watching instead of a random sample.

Coaching agents

Turns those scores into something a team lead can use on Monday: the clip, the moment it went wrong, and what good looked like on the call that went well.

Compliance and risk agents

Watches for the disclosure that was not read, the promise nobody is allowed to make, and the personal data that ended up somewhere it should not be.

Agent assist

Sits beside your human team in real time with the answer, the next step and the wording, rather than replacing them.

Back office, finance and people

Queues somebody works through by hand, where the agent touches the record and a wrong move costs money.

Document agents

Invoices, claims, applications, contracts, packing lists. Reads whatever shape it arrives in, checks it against your system of record, and flags the disagreement rather than averaging over it.

Finance operations agents

Matching, reconciliation, chasing what is overdue, coding the transactions your bookkeeper recodes every month, and the procurement round trip from request to purchase order.

Order and exception agents

Watches the exceptions on the order, the load or the job, works out what to move and who to tell, and updates the systems that are supposed to agree with each other.

HR and recruiting agents

Screens against the requirement rather than the keyword, answers the policy questions HR retypes weekly, schedules across three calendars, and runs a joiner or a leaver through the checklist.

IT and internal service desk agents

Access requests, provisioning a new starter, triaging the ticket, running the runbook step nobody remembers, and escalating the ones that touch production.

Meeting and admin agents

Notes that turn into actions with owners, follow-ups that actually get sent, and scheduling across calendars that never agree.

Long-running workflow agents

Jobs measured in days. It waits on an approval, survives a system going down over the weekend, retries what failed, and never quietly drops an item on the floor.

Knowledge and decision support

The questions that currently need the one person who has been there seven years.

Retrieval agents

Answers across contracts, tickets, wikis, drives and the shared folder nobody has opened since 2019, with a citation behind each claim and a plain "I do not know" when the answer is not there.

Internal copilots

Your own staff ask the business a question and get a real answer: the numbers, the policy, the last three cases like this one, and a draft of whatever they were about to type.

Research and monitoring agents

Watches the sources you would read if you had the time, regulations, competitors, prices, tenders, and tells you what changed and what it means for you.

Analytics agents

Answers a question against the warehouse, shows the query it ran, and flags the number that moved before anyone opens the dashboard.

The layer that keeps the rest trustworthy

The part most quotes leave out, and the part that decides whether the thing survives its second month.

Guardrails and approval routing

Rules on what an agent may reach and what it may do alone. Anything above the line stops and waits for a named person in Slack or email, with the reasoning laid out.

Evaluation harnesses

The cases, the grader and the scores that tell you an agent got worse before your customers do, and tell you when it has earned a longer leash.

Orchestration across several agents

When one job needs a researcher, a decider and a doer, they run as one system with one trace and one owner. Not four bots emailing each other.

The two figures above are 2026 industry benchmarks rather than our own numbers: resolution rates by integration depth are from Lorikeet’s 2026 benchmark study, and QA sampling coverage from Mihup’s analysis of automated agent scoring. We quote them because they set a realistic expectation, not because they flatter us. An agent given read-only access lands at the bottom of that range, and integration depth is what moves it.

Some of this you can have without a build at all. Our own platform ships an AI sales agent with lead qualification on every plan, and live chat with human handover plus the omnichannel inbox from the second tier up. If a plan covers the job, we will point you at the plan and take no money for a build. An engagement is for the work the plan was never shaped for.

An agent is only as useful as what it is plugged into.

A model with no reach can draft an answer. Give it your systems and it can finish the job. Read access first, write access once the numbers say it has earned it.

Email and calendar
Google Workspace, Microsoft 365, Exchange
CRM and sales
Salesforce, HubSpot, Pipedrive, Zoho, your own
Support and tracking
Zendesk, Freshdesk, Intercom, Jira, Linear
Finance and ERP
Xero, QuickBooks, NetSuite, SAP, Dynamics
Commerce
Shopify, WooCommerce, Magento, marketplace feeds
Data
Postgres, MySQL, SQL Server, Snowflake, BigQuery
Files and documents
Drive, SharePoint, Dropbox, S3, DocuSign
Messaging and telephony
Slack, Teams, WhatsApp, Twilio, SIP

We build against anything with an API, anything that speaks MCP, and a fair number of systems that only expose a database. If yours is a twenty-year-old package with a green screen and no documentation, say so on the call. It is usually still doable, and it is always more work.

How much rope the agent gets is your decision, per action.

This is the part that decides whether a build survives contact with real customers. Every action an agent can take sits in one of three places, and where the line falls is a setting you control rather than something you find out about afterwards. It is also the pattern that production deployments converged on through 2026: autonomous execution for routine work, a named human approving anything high-stakes, and a queryable audit log behind both, as surveys of agent control planes describe it.

How agent autonomy narrows as the cost of a wrong answer risesThree bands across one axis. On the left, low-value reversible actions the agent takes on its own. In the middle, expensive or contractual actions the agent proposes and a named person approves. On the right, low-confidence or out-of-scope cases the agent hands to a human.Agent actsAgent proposes, person approvesPerson decidesReversible, low value, high volumeExpensive, contractual, irreversibleCost of a wrong answer, increasing
The three autonomy levels an action can be assigned, and who makes the decision at each
LevelWhat sits hereWho decides
It acts on its ownThe routine majority: answering, filing, booking, updating a record, issuing the refund that sits inside the policy. Reversible, low value, high volume. This is where the money is.The agent
It proposes, a person approvesAnything expensive, contractual or awkward. The agent does the work and presents the decision with its reasoning, and a named person clicks yes in Slack or in email. Faster than doing it yourself, safer than letting it run.A named approver
It stops and hands overLow confidence, an angry customer, or anything outside what it was given. The person receives the whole thread and the working, not a ticket saying the bot could not help.Your team

Underneath all three, every run keeps a trace: what came in, which tools the agent called with what arguments, which rule allowed or blocked each call, who approved what, and what it produced. That is what you show an auditor, and it is what you read at 9am when something went sideways at 2. Where the line sits is not a personality you hope the agent has, it is a rule per action, and it moves. Most builds start with almost everything requiring approval, because that is how the team learns whether they trust it, and the evaluation numbers are what justify widening it later. Tightening it again after a bad week is a settings change rather than a redeploy, which matters more than it sounds like it should.

The vocabulary changes by industry. The shape does not.

Something arrives, somebody reads it, decides, and updates a system. That is the same job in a lettings office and a claims department. What we ask for is one person who knows how the decision is really made, including the parts nobody wrote down.

Real estate and property

Qualifies an enquiry on budget, area and timing, matches it against what is actually on the books, and puts the viewing in the diary of whoever covers that patch. Chases the ones that go quiet, and keeps the landlord and tenant threads apart.

Recruitment and HR

Screens against the requirement rather than the keyword, answers the candidate questions your team retypes every week, schedules around three calendars, and walks a joiner through onboarding without anyone chasing a form.

Clinics and healthcare admin

Triages a request against your own protocol, books it into the correct clinic slot, handles the reschedules and the reminders, and puts anything clinical in front of a human with the notes already gathered.

Finance, lending and insurance

Reads the application or the claim with its attachments, checks it against policy, clears the clean ones inside the limits you set, and hands the edge cases to an underwriter with the reasoning written down and auditable.

Legal and professional services

Reads the contract, the tender or the brief, pulls out obligations, dates and unusual clauses, compares against your standard position, and drafts the response for someone qualified to sign off.

Travel and hospitality

Works out dates, party size and budget before a person is involved, holds the option, upsells where it is welcome, and rebuilds an itinerary at midnight when a leg gets cancelled.

Logistics and field service

Watches exceptions on the load or the job, works out what to reschedule and who to move, updates the systems that need to agree, and tells the customer before the customer calls you.

E-commerce and retail

Answers against the live order rather than a generic policy page, decides refunds, exchanges and reships inside the rules you wrote, and escalates the ones where a person is better placed.

Manufacturing and distribution

Turns an emailed request for quote into a priced response against your catalogue and stock, chases the missing spec, and flags the orders that will not ship on time before the delivery date does.

Education and training

Handles admissions questions, applications and timetable clashes, tracks who has not completed what, and gives a tutor the summary before the meeting instead of after it.

Half the time you do not need an agent. If a rule can make the decision, a rule should make it.

A scheduled script with twenty lines of conditions is cheaper to run, cheaper to fix at 3am, and it never invents anything. Agents earn their cost on messy input and judgement calls: a document that turns up in a different shape every time, a customer who buries the real question in paragraph four, an exception nobody wrote a branch for, a decision that depends on three systems disagreeing with each other. The useful question is not whether AI could do the task, it is whether the task varies. Where the input arrives in the same shape every time and the rules fit on a page, ordinary automation wins on cost, on predictability and on the amount of sleep your team gets. We will say so on the first call if what you have described is a workflow rather than an agent, and sometimes the honest answer is an n8n flow we set up in a week that you never pay for again.

The framework is a decision we make per project.

We are not loyal to one of these, and you should be suspicious of a shop that is. If your team already lives in n8n, we build into your n8n rather than standing a second system up beside it. If the work wants the multi-agent shape CrewAI is good at, or a graph in LangGraph, or one of the vendor agent SDKs, that is a scoping decision we make with you and explain.

A fair amount of this work ends up as ordinary Python with a state machine and no framework at all, and the boring option wins more often than people expect. Where a system you already run speaks MCP, we connect through that instead of writing one more adapter for you to maintain.

What we build with
  • LangGraph
  • LangChain
  • CrewAI
  • n8n
  • MCP
  • Claude Agent SDK
  • OpenAI Agents SDK
  • LlamaIndex
  • Qdrant
  • pgvector
Models we build on
  • Claude
  • GPT
  • Gemini
  • Llama
  • Qwen
  • Kimi
  • GLM
  • DeepSeek
  • Mistral

Which one handles which step is a cost and accuracy question, and we settle it with measurements rather than opinions.

Open weights or a closed API. You pick, and you can change your mind later.

Closed API models compared with open-weight models across reasoning, data location, cost shape and operational burden
Closed modelsClaude, GPT, GeminiOpen-weight modelsLlama, Qwen, Kimi, GLM, DeepSeek, Mistral
Hardest reasoningStill ahead on the genuinely ambiguous cases, which is why we keep one for the step where judgement decides the outcome.Trails, and the gap has narrowed with every release cycle. Fine for extraction, classification and routing.
Where your data goesOut to the provider, under a contract. Enough for most commercial work, not enough for some regulators.Wherever you put it, including a rack in your own building. Nothing crosses a border you did not choose.
Cost shapePer token. Cheap to start, and it scales with your success whether you like it or not.GPU time. Higher floor, and the cheaper shape once volume is high.
What you operateNothing. No hardware to buy, no inference to keep up.Real infrastructure. Somebody has to own the GPUs, the serving stack and the upgrades.
Changing your mindA configuration change and a re-run of the evaluation set.The same, because the model sits behind our interface rather than through the code.

Most builds end up mixed rather than picking a side. An open model does the high-volume extraction, classification and routing, where the work is repetitive and the cost per call is what decides whether the project pays for itself, and a closed model is kept for the one step where the judgement is genuinely hard. Because the model sits behind our own interface rather than threaded through the code, swapping one for another is a configuration change and a re-run of the evaluation set instead of a rewrite. That matters more than which model is ahead this quarter: prices fall, licences change, a provider has a bad month, and a better open release lands every few weeks. The build should be able to take advantage of all of that without you paying for it twice.

Your data, on ground you choose.

Where the agent runs is your decision and the first thing we settle in discovery, because it changes what the rest of the build can look like.

Your cloudWe deploy into your own AWS, GCP or Azure account. You own the infrastructure, the logs and the bill, and you can lock us out the day the work ends.
Our cloudWe host it and keep it running, isolated per customer, and you keep the same view of what the agent did and why it did it.
Your own hardwareOpen-weight models inside your own network, air-gapped if that is what your regulator wants. The prompts do not leave the building either.

Personal data is masked before anything reaches a model. You set retention, and zero is one of the settings. What an agent may read is scoped the same way as what it may do, so it cannot answer from a folder it was never meant to open. Access is by role, and your material is not training data for us or for anyone we call.

The evaluation set is part of what you get.

An agent that demos beautifully and falls over in week three is the normal outcome. What prevents it is writing down what a correct answer looks like before the build starts, then running every change against those cases.

You get the cases, the grader and the scores. When somebody edits a prompt six months from now, your team can tell whether the agent got better or only got different. It is also what tells you when the agent has earned a longer leash, and what lets you move to a cheaper model without crossing your fingers.

A price before the code, not after it.

Most AI agent development quotes are a guess wearing a number. Four steps, and you can stop after any of them.

01

Scoping call

Tell us what the process is, which systems it touches, and what a wrong answer costs you. You leave that call knowing whether this should be an agent at all.

02

Discovery

We map the decision, the data behind it and the integration points, agree where the approval line sits, and write the evaluation cases with your team. You get a fixed scope and a price before anyone writes agent code.

03

Build

A working agent on your real data, measured against those cases every time it changes. A pilot group is usually running it within weeks, with the agent on a short leash while the numbers come in.

04

Run it, or take it

We roll it out, widen what it is allowed to do as it earns the trust, and keep it healthy. Or we hand over the repository, the prompts, the evaluation set and the deployment scripts, and your engineers carry on without us.

Questions we get asked before the call

What is an AI agent in a business context?

An AI agent is a system that takes a goal, works out the steps itself, calls real software to carry them out, and checks the result before deciding what to do next. In a business it usually sits where a person currently reads something and then updates a system: an enquiry, an invoice, a claim, a ticket, a call. The test is whether it can complete the task, not whether it can describe it.

What is the difference between an AI agent and RPA?

RPA follows a recorded path and breaks when the screen, the form or the wording changes. An agent reads the input, works out what it means, and picks the path, so a document arriving in a shape nobody anticipated is a normal Tuesday rather than an incident. RPA is cheaper and more predictable where the input never varies. Where it does vary, RPA becomes a maintenance bill.

How is a custom agent different from a chatbot?

A chatbot picks a reply. An agent picks an action, runs it against a real system, looks at what came back, and decides what to do next. That is the difference between telling a customer how to get a refund and issuing it, and it is why deflection bots plateau around a third of contacts while agents wired into the backend get past two thirds.

How much are you willing to let it do on its own?

That is your decision and we build it as a setting, not a personality. Each action gets a rule: do it, propose it and wait for a named person, or refuse. Value thresholds, customer tier and confidence all feed into it, and you can tighten the whole thing in an afternoon if something goes wrong.

How does it connect to our systems?

Through your APIs where they exist, through MCP where a system already speaks it, and through the database where a vendor never built an API. Read access first, write access when the evaluation numbers say it has earned it. The integration depth is what decides how much the agent can actually finish.

How long does a build take?

Discovery runs one to two weeks. A first agent doing real work for a pilot group is usually a few weeks after that. Getting to full production is mostly a question of how many systems it touches and how careful the sign-off has to be, so that part varies.

What does it cost?

We price the build after discovery and you see the number before we start. What moves it is the number of systems the agent integrates with and how expensive a wrong answer is, because that decides how much review, testing and guarding goes in.

Do you work in our industry?

Probably. The shape repeats: something arrives, someone reads it, decides, and updates a system. The vocabulary and the rules change, the shape does not. What we ask for is one person who knows how the decision is really made, including the parts that are not written down anywhere.

Is this the same thing as agentic AI?

Yes, agentic AI is the umbrella term and a custom agent is what you get at the end of it. The word matters less than the test behind it: can the thing choose an action, run it against a real system, look at the result and decide what to do next. If it can only pick a sentence to say, it is a chatbot with better marketing.

Do you train models on our data?

No. Under the API terms of the providers we use, business inputs are not used to train their models, and we do not train on your material either. When a contract is not enough for your regulator, we run open-weight models on hardware you control and nothing goes out to an API at all.

Can it run on our own servers?

Yes. Open-weight models such as Llama, Qwen, Kimi, GLM or Mistral run inside your network, with the agent, the vector store and the logs sitting there alongside them.

Which frameworks do you build on?

LangGraph, LangChain, CrewAI, n8n, MCP and the vendor agent SDKs, picked per project instead of pushed at everybody. Plenty of builds end up as ordinary Python with a state machine, because a framework is not always worth the ceiling it puts on you. If your team has already standardised on something, we work in that.

What happens when the agent gets something wrong?

You find out from the evaluation set rather than from a customer, which is the whole point of shipping one. Every run keeps a trace of its input, the tools it called with what arguments, the rule that allowed or blocked each call, who approved what, and what it produced, so a bad outcome is a readable sequence rather than a mystery. Tightening what the agent may do alone is a settings change, not a redeploy.

How do you measure whether an agent is actually working?

Against cases written before the build starts, with your team, describing what a correct outcome looks like for your process. Every change runs against them, so you can see whether an edit made the agent better or only different. Resolution or completion rate, escalation rate and the cost per completed task are the numbers that decide whether it stays switched on.

Can an agent work with our CRM, ERP and internal database?

Yes, and that is usually where the value is rather than in the model. We integrate through your APIs where they exist, through MCP where a system already speaks it, and directly against the database where a vendor never built an API. Read access comes first and write access follows once the evaluation numbers say it has earned it.

Do we own what you build?

Yes. The code, the prompts and the evaluation set are yours at handover.

Describe the process. We will tell you whether it needs an agent.

Half an hour, an engineer rather than a salesperson, and a straight answer on whether this is worth building. If it is not, we will say that too.

Or send us the brief

Tell us what you want built. We reply within one business day.

By sending this, you agree to our Privacy Policy.