AI Customer Service Chatbot That Won't Invent Prices
Darya NikolaevaDigital Marketing · 7 October 2026
Only 27% of customers will try a company's chatbot again after one bad experience, according to a Gartner survey of 3,566 customers published in September 2026. For a small business in the US or Australia, the fastest way to spend that one chance is an AI customer service chatbot that quotes a price nobody agreed to.
We recently built a support assistant for a moving company in the United States. In moving, a wrong number turns into an argument on moving day, or into a customer who never calls back.
So the owner gave us one rule. The bot can be slow, it can say "let me check with the team", but it never names a price it can't prove. This is how we built that, including the parts we got wrong on the way.
The short version
Every number in a draft reply is checked against the company's own knowledge base. If it can't be traced, the reply is held and a person is called.
A customer who asks for a human gets one on a timer: reminders from minute five, a second contact on the fifth reminder, a form at minute fifteen.
The bot keeps answering while it waits. Only the operator's first real reply switches it off.
Every knowledge-base update runs against 48 reference questions, and a version that scores worse doesn't ship.
Tests never message real staff. We learned that one by opening nine real threads in a working chat.
If the AI provider goes down, the chat turns into a request form instead of an error. 13 of 13 checks passed with the model switched off.
What the market data says about AI customer service chatbots
Customers are open to chatbots, but only on conditions. Half say service is easier when a company uses GenAI, yet 87% call access to a human agent essential (Gartner, 2026). And 61% of service leaders already have a backlog of knowledge articles waiting to be fixed (Gartner, December 2024).
Customers who would have used a chatbot if offered
49%
Gartner, 2026
Customers who actually used one last time
7%
Gartner, 2026
Customers willing to retry a chatbot after one bad experience
27%
Gartner, 2026
Customers who say access to a human is essential
87%
Gartner, 2026
Service leaders with a backlog of knowledge articles
61%
Gartner, 2024
That 27% is the number I'd pin above any chatbot project: one confident wrong answer can cost you that customer's chatbot use for good. The gap between 49% willing and 7% actually using a bot says the same thing from the other side.
Why bots make up prices
A language model is trained to produce a fluent answer, and fluency keeps going after the facts run out. The gap gets filled with something plausible. Researchers call it hallucination; a customer calls it being lied to.
What do the cases say? In February 2024 a Canadian tribunal ordered Air Canada to honour a bereavement-fare refund that its support chatbot had made up (Moffatt v. Air Canada). The airline argued the chatbot was responsible for its own words.
The tribunal disagreed: the company was.
In April 2025 the support bot of the code editor Cursor, signing its emails as "Sam", invented a login policy that didn't exist, and users cancelled before staff stepped in with refunds (Cursor). Klarna went further and moved much of its support to an AI assistant; in May 2025 its CEO said cost had been "a too predominant evaluation factor", quality dropped, and the company went back to hiring people (Entrepreneur).
If a number can't be traced to the business's own records, the bot doesn't get to say it.
How we check every number
The assistant writes a draft reply. Before that draft goes anywhere, a separate check pulls out every amount and looks for it in the knowledge base the assistant was given to read.
If every number matches, the reply goes out. If one doesn't, the reply is held: the customer sees a short "let me confirm this with the team", and the conversation goes to a person with the draft attached, so the operator sees exactly what the bot was about to say. A reply cut off mid-sentence is held the same way, because half a sentence about a price is worse than none.
Problemthe first version of the check would have held almost every correct answer. A reply like "that costs $1,200." ends with a full stop, and "$1,200." didn't match "$1,200" in the base. Unit tests passed; a live conversation caught it.
Fixwe changed the matching and put the case into the reference set, so the bug can't quietly come back.
The quieter version of the same problem starts with the customer naming a price first, something like "your colleague said it would be $900". A polite model agrees. That number isn't in the base either, so the reply is held.
Repeating a customer's figure is still inventing a price. It just sounds friendlier.
The decision lives in one place in the code and every channel shares it: the website widget now, Instagram, Messenger and WhatsApp once they're connected. A new channel can't arrive without the check.
Do it yourself. Ask your current chatbot five price questions whose answers aren't on your website, then one more where you name a wrong price yourself. Count the confident numbers you get back. If the count is above zero, the bot needs a gate before it needs a better prompt.
When should an AI chatbot hand a customer to a human?
As soon as the customer asks, and then on a clock. In Gartner's 2024 survey the top fear about AI support was that reaching a person gets harder, and by 2026 87% of customers called that access essential (Gartner). The hand-off is the part that has to hold up on a busy afternoon.
Minute five without a reply: the first reminder to the operator, then one every minute.
Fifth reminder: a second contact, the owner or whoever is on duty, through a mention that breaks silent mode.
Minute fifteen: the customer gets a short form, so the request is captured even if nobody is free.
Minute twenty-five: the conversation returns to the assistant, which keeps helping with what it can answer.
Outside working hours nobody gets pinged. Ten reminders at 2 a.m. won't produce an answer, so the customer is told when the team is next online.
Problemin the first test, one deleted chat thread crashed the whole reminder run, and nobody got reminded about anything. Forty open conversations at once also hit the messaging app's rate limit.
Fixeach conversation is processed on its own, a lost thread is replaced with a new one, and reminders go out at most eight per pass.
The bot doesn't go quiet while it waits
Our first version switched to human mode the moment a customer asked for a person. The operator hadn't even seen the request, and every next question got the same reply, "I've passed this to the team", for up to 25 minutes.
Now human mode starts with the operator's first real reply. Until then the assistant keeps answering from the knowledge base, and if a person takes over while it is still typing, its reply is cancelled so the customer never gets two answers at once.
Do it yourself. Write down, in minutes, how long a customer may wait for a person before someone else is pulled in, and who that someone is at night. If the honest answer is "it depends", your customers are already waiting on that.
How do you test an AI customer service chatbot before every update?
With a fixed set of questions that have known right answers. Ours has 48: prices and facts, questions the bot should answer with "I don't know", topics it mustn't touch, replies in the customer's language, and questions that qualify the request. In five of the 48 the correct move is to call a person instead of answering.
Every new version of the knowledge base runs against the set. If it scores worse than the version that's live, it doesn't ship.
Where the knowledge base itself was wrong
Collecting the knowledge took longer than writing the code, and we are not unusual: 61% of service leaders in Gartner's survey have a backlog of knowledge articles waiting to be fixed. In the owner's own materials we found two phone numbers that differed by a single digit, and a price in the FAQ that was off by a factor of ten. The bot would have repeated both, politely and with total confidence.
Tests never message real people
During development the assistant ran on the real connections to the team's chat and CRM. One run of the reference set opened nine real threads in the team's working chat, and a colleague started answering them.
Nothing broke. We still stopped relying on whoever runs the tests to be careful: a sandbox switch now cuts both outgoing paths on every local run and prints what would have been sent.
Do it yourself. Collect the 30 questions your team answered most often last month, with the answers you would sign off on. That list is your first reference set, and building it will show you which answers nobody on the team actually agrees on.
The assistant isn't live on the client's site yet, so we can't give you conversion numbers. What we can give you is what it did under test, on the full funnel from the first message to the request landing with the team.
Test suite
What it checks
Result
Acceptance
The whole funnel against the client's brief
39 of 39, three runs in a row with no differences
Integrations
Requests reaching the CRM and the team chat
15 of 15
Widget
History after reload, greeting, suggestions, picking up contact details
24 of 24
AI provider down
A conversation with the model switched off
13 of 13
Our own test logs, September 2026. Run against a staging copy, not live customers.
When the AI is down and nobody notices
Every model provider has outages. We ran the full conversation with the model's key deliberately broken, and the widget still greeted the customer, took the question, saved the contact details and called a person. All 13 checks passed.
An outage at 9 p.m. costs you nothing if the chat quietly turns into a request form. It costs you the customer if the chat shows an error.
Do it yourself. Ask whoever runs your chatbot one question: what does a customer see if the AI provider is down right now? If nobody knows, you have your answer.
What an AI customer service chatbot costs to run
Less than most owners expect. The model is the cheapest line: caching the knowledge base, which Anthropic prices at a tenth of the normal input rate on every re-read, made a conversation 2.4 times cheaper in our measured runs.
The expensive part is people's time: collecting the knowledge and checking it. A spending ceiling per company sits on top, so a traffic spike or a deliberate flood can't turn into a surprise bill.
What each safeguard prevents
Failure
Public or our own example
What it cost
Safeguard
Invented refund rule
Air Canada, 2024
About CA$812 in damages and fees, plus worldwide coverage
Number check, reply held
Invented account policy
Cursor, 2025
Cancelled subscriptions and refunds
Same check, hand-off to a person
Cost-first automation
Klarna, 2025
Lower quality, people hired back
Human on a timer
Customer's own price repeated
Our test run
Caught before launch
Held like any untraced number
Tests reaching real staff
Our test run
Nine real threads, one confused colleague
Sandbox switch
AI provider outage
Our test run
Chat would have shown an error
Fallback to a request form
Air Canada figure from the tribunal decision; the other costs are as reported, without amounts.
What we'd do differently
If you have budget for one thing, put it into the knowledge base before the model. Most failures in that table start with a bot that was never told what it doesn't know.
We'd also start the reference set on the first day, earlier than we did. The one number we don't have yet is conversion on a live site. Ask us in three months.
If you're relying on automations someone else built and can't see into, this piece on borrowed automation covers how they fail quietly.
FAQ
Can an AI chatbot give customers wrong prices?
Yes. A language model fills gaps with plausible guesses, and a price is where a guess costs most. The fix is a check outside the model that blocks any number it can't trace to your records.
Is a business responsible for what its chatbot says?
In the Air Canada case a Canadian tribunal said yes, and the airline had to honour the refund its chatbot invented. Treat every bot reply as something your company said, because customers will.
When should a chatbot hand over to a human?
When the customer asks, when the bot can't trace a fact or a price, and when the topic is on your stop list. The hand-off itself should run on a timer, with a second person pulled in if the first doesn't answer.
How much does an AI customer service chatbot cost to run?
For a small business the model bill is usually the smallest line, provided the knowledge base is cached; caching made our measured setup 2.4 times cheaper. The bigger cost is the time it takes to collect and check the knowledge.
Teaching software to say "let me check"
The odd part of this project was how much of the work went into teaching the bot to hold back. A good human support agent already knows the move: "let me check and come back to you." We had to build that one sentence into the software, with a timer and a second person behind it.
If you're weighing an assistant for your own site, send me the ten questions your customers ask most. I'll send back the reference-set template we use, filled in with your ten, and tell you which ones a bot can safely take and which should stay with your team.
Comments
No comments yet. Be the first.