AI Chatbot Wrong Answers: 5 Checks Before You Trust One
Sublex Digital · 2 September 2026 · 14 min read

AI chatbot wrong answers are not a rare glitch. They are the ordinary behaviour of a system built to sound helpful, running on a website where nobody is checking. The assistant does not know it is wrong. It produces a fluent, confident sentence, the customer believes it because it appeared on the business's own site, and the business finds out days later when somebody arrives expecting a price that was never offered.
Preventing AI chatbot wrong answers is the question worth settling before putting an assistant on a website, and it is not the question most buyers ask. Almost every demo is judged on whether the assistant answers. The better test is what it does when it should not.
Where AI chatbot wrong answers actually come from

They are not all the same failure, and the difference matters, because three of the four are fixable by design and the fourth is the one nobody tests for.
Nothing to go on. A customer asks about something the business never wrote down anywhere. Weekend opening. Whether a deposit is refundable. Whether you deliver to the next town. A language model with no source and no permission to refuse will fill the gap, because filling gaps is what it is for. This is the failure people picture when they worry about AI, and it is the easiest one to prevent.
The wrong page. The business has the answer, but the assistant reaches for the wrong part of the site. A price from last season's list. A policy written for trade customers pulled into an answer for a member of the public. The sentence is true somewhere on the site, which makes it far harder to catch than an invention, and it is just as wrong to the person reading it.
Half the page. The answer exists, it is correct, and it has a condition attached. The assistant reports the general rule and drops the condition. This is the dangerous one, because the answer survives casual checking. Read it back and it looks right. It is only wrong to the specific customer the condition was written for, which is exactly the customer who asked.
The stale page. The material was right when it was loaded, and the business has since changed a price or an opening time on the door but not on the website. No assistant can fix this, and any vendor claiming otherwise is selling something. What a good one can do is make the gap visible, which is the missed questions section further down.
The case that turned this into a boardroom question
In November 2022 a man named Jake Moffatt asked the chatbot on Air Canada's website about bereavement fares. The chatbot told him he could book a normal ticket and claim the discount within 90 days afterwards. That policy did not exist. The airline's real policy, sitting on a different page of the same website, did not allow claims after travel.
When Air Canada refused the refund, Moffatt took it to the British Columbia Civil Resolution Tribunal. In February 2024, in Moffatt v. Air Canada, 2024 BCCRT 149, the tribunal found the airline liable for negligent misrepresentation and ordered it to pay $650.88.
The amount is trivial. The finding is not. Air Canada argued the chatbot was a separate entity responsible for its own statements, and the tribunal rejected the argument. What the assistant says is what the business said. The legal side of this is worked through properly in a longer piece on AI chatbot liability, including what a disclaimer does and does not do, which is less than most people assume.
For a small business the exposure looks different from an airline's. Nobody is taking a corner shop to a tribunal over six hundred dollars. The cost arrives as refunds honoured because arguing is not worth it, appointments that were never available, staff time spent apologising, and a customer who tells other people that the website lies.
Why a better model does not fix it
The instinct is to reach for a smarter model, on the assumption that AI chatbot wrong answers are a quality problem that scale will eventually solve. They are not.
A language model is a system for producing likely text. Asked about a delivery charge it has never seen, the likely text is a delivery charge, because that is what an answer to that question looks like. It has no internal marker separating something it retrieved from something it constructed, so it cannot warn anybody which just happened. A stronger model produces a more convincing wrong answer, not fewer of them.
What changes the outcome is not the model's ability. It is what the model is allowed to see, and whether it is permitted to say no.
The design that works: one source, and permission to refuse
An assistant that cannot produce AI chatbot wrong answers about a business is one that has nothing to produce them from.
Sublex Chat reads the business's website, its documents, its price list, the question and answer pairs it was given, and the description the business wrote of itself. It may draw on those and nothing else. There is no third source, so it cannot repeat something a model happens to believe about hotels or clinics or garages in general. Text on a crawled page is treated as quoted material rather than as an instruction, which matters on any site carrying wording a stranger could edit.
The second half is the part that gets skipped. Restricting the source is worthless unless the assistant is allowed to come back empty. If a system must always produce an answer, restricting its material only moves where the invention starts. The permission to refuse is the feature. Everything else is scaffolding around it.
In practice that means a short list of things it will not do at all. It will not confirm a price that is not in the material. It cannot see a calendar, so it never promises an appointment or says when somebody will call. It gives no medical, legal or financial advice, however the question is phrased. Those limits are not a shortfall in the product. They are the product.
A refusal is not a failure, it is a lead
The objection is obvious. A customer who gets "I do not have that" is a customer who has not been helped.
True, and still better than the alternative, because those are not the two options. The two options are an admission or an invention, and the invention costs money later. But a refusal on its own does waste the visit, which is why what happens next is the part worth judging a product on.
When Sublex Chat cannot answer, three things follow. It says so plainly. It asks for a way to reach the person, so the enquiry survives instead of bouncing. And it records the question, in the customer's own wording, on a screen in the dashboard.
That last one turns the weakness into the most useful thing the product does. A month of missed questions is a list of everything the website forgot to mention, written by the people who wanted to know. Type an answer into the row and it is known from then on. Most businesses find the same four or five questions at the top, and they are rarely the ones anybody expected.
The hard case: when the honest answer has a condition
The questions with no answer are the easy ones. The dangerous ones are the questions with a nearly complete answer.
Picture a garage whose own page says every service includes a courtesy car, except the same-day express service, which does not. Ask "do I get a courtesy car?" and a summarising assistant will happily answer "yes, every service includes one" and stop.
It has invented nothing. It has dropped the second half of a sentence. And the customer who booked the express service turns up expecting a car that was never part of it.
Carrying the condition with the fact, every time, even when only the general part was asked about, is the hardest thing in this category to get right, and it is the least likely thing to be demonstrated to you. It is worth its own test, and worth running that test more than once, worded differently each time. A single well-behaved answer in a sales demo tells you very little.
Five checks before you trust an assistant on your website
Ten minutes, on any product, before any money changes hands. Each one is aimed at a different source of AI chatbot wrong answers.

1. Ask something plausible that is not on your site. Not nonsense. Something a real customer would ask that you happen never to have written down. Weekend hours, a deposit, delivery to a particular town. Watch whether it admits the gap or produces something reasonable-sounding. This is the whole test in one question, and a surprising number of products fail it inside a minute.
2. Ask something with a condition attached, more than once. Use a rule of yours that has a "but" in it. Check the "but" survives. Then ask again, worded differently. This failure is not consistent, which is precisely why it reaches customers and never shows up in a demo.
3. Ask it to do something outside its job. Ask for medical advice, or a legal opinion, or a discount. Push a little. An assistant that will improvise a discount because a customer pressed for one will improvise other things when nobody is watching.
4. Look at what it does with the failure. When it cannot answer, does the enquiry survive? Is the question recorded somewhere you will actually look, in the customer's words rather than a category? A product that refuses cleanly and then drops the customer has solved the wrong half of the problem.
5. Check it says what it is. The assistant should identify itself as an assistant. Regulators are moving this way, and beyond compliance it is simply how a customer decides how much weight to give an answer. Nobody should have to work out whether they are talking to a person.
What to do if yours is already getting things wrong
Plenty of businesses arrive at this after the fact, with an assistant already on the site and a customer already annoyed. The order to work in is not obvious, so here it is.
Turn it off before improving it. An assistant that is off costs enquiries. An assistant that is confidently wrong costs money and gets quoted back at you.
Work out which of the four failures it was. Take the answer that was wrong and go looking for it on your own site. If the sentence exists somewhere, this is a retrieval problem and your material is fine. If it exists nowhere, the product invented it, and no amount of editing the website will fix that. If it exists with a condition that got dropped, that is the third failure, and it will be intermittent.
Clean the source before blaming the tool. A large share of AI chatbot wrong answers turn out to be right answers taken from a page the business forgot was still published. Old price lists, a policy from two seasons ago, a duplicate page nobody links to any more. Clearing those out is unglamorous and often solves the problem outright.
Then judge the product on what is left. Once the material is clean, anything still wrong is the product's doing, and the five checks above will tell you quickly whether it is worth keeping.
What the two failures cost, side by side
Both kinds of failure have a price, and AI chatbot wrong answers are much the more expensive of the two.

An assistant that says "I do not have that, can I take your number" costs one enquiry that needed a person, and hands over a contact and a question worth answering.
An assistant that invents costs the refund or discount honoured to avoid an argument, the staff time spent on it, the customer who does not come back, and, on the Air Canada precedent, a finding that the business is answerable for what its assistant said. The second cost arrives later than the first, which is exactly why it gets discounted at the moment the buying decision is made.
Cost is the other half of this decision, and it is decided by the billing model rather than the headline price. What the four pricing models really cost works the arithmetic through at 500 conversations a month.
A business does not need an assistant that always has an answer. It needs one whose answers can be believed at eleven at night, when there is nobody at the desk to correct it.
Where to start
No account is needed to see this working. Give Sublex Chat a web address and watch it read the pages and answer questions from what it finds, then try to catch it out using the five checks above. It takes about two minutes.
When that is settled, what it does covers the inbox, the human takeover and the missed questions screen, what it costs is one price with the AI included rather than billed beside it, and how it compares sets it against the products most people look at first, using their own published prices. There is a free plan, and it stays free.
Frequently asked questions
What causes AI chatbot wrong answers? Four things, and only one of them is invention. The assistant may have no source for the question, reach for the wrong part of the site, report a rule while dropping the condition attached to it, or repeat something that was true when it was loaded and has since changed. Restricting what it may draw on prevents the first. The second and third are a matter of how carefully it reads. The fourth is housekeeping no product can do for you.
Can an AI chatbot be stopped from making things up? Largely, by design rather than by tuning. If the assistant may only use material the business supplied, and is permitted to say it does not know, the room for invention nearly closes. If it must always produce an answer, no amount of restriction helps, because it will construct one from whatever it has.
Is a business responsible for what its chatbot tells a customer? On the evidence so far, yes. In Moffatt v. Air Canada the airline argued the chatbot was responsible for its own statements and the tribunal rejected it. Treat anything the assistant says as something the business said.
Will a disclaimer protect me? Less than most people expect. A notice saying answers may be inaccurate does not obviously undo a specific, confident statement a customer relied on. It is worth having, and it is not a substitute for an assistant that does not invent.
What happens to the questions it cannot answer? On Sublex Chat they are captured and listed in the dashboard in the customer's own wording, with how many times each has been asked and a button to answer it. That list is usually the most useful thing a business gets in its first month, because it is a record of what its website failed to say.
Does refusing to answer lose customers? It loses fewer than inventing does. A refusal that captures a phone number costs one follow-up call. An invented price costs the discount, the argument and the customer. The number worth watching is not how often the assistant answered, but how often it answered correctly.
See Sublex Chat answer from your own website
Give your web address and watch it read your pages, then answer questions about your business. No account needed.