Knowledge base

Test a website chatbot before visitors use it

Use a fixed chatbot test set to check approved facts, unknown questions, human handoff and mobile behaviour. Record blockers and retest before a launch decision.

A chat window that opens and produces fluent text is not an acceptance test. Before visitors use it, check approved facts, unknown topics, human handoff and how the widget behaves on the actual page.

This guide helps you assemble cases, define expected outcomes and record a release decision. The English chatbot test worksheet is a blank template. No OmniTechs or customer chatbot was tested to produce this article, and no accuracy or compliance result is claimed.

Build four groups of test cases

  1. Approved facts: questions answered by a current authoritative source, including relevant qualifications.
  2. Variation: paraphrases, informal wording, misspellings and several questions in one message.
  3. Unknown or restricted topics: missing facts, internal information and a different audience or country than the source covers.
  4. Handoff: complaints, urgency, requests for exceptions, direct requests for a person and attempts to bypass boundaries.

Use protected examples based on real contact reasons, with fictional personal details. For each case, name the source and expected action before looking at the response. An answer's existence is not evidence that it is correct.

Choose enough cases to cover these groups and few enough that an informed person can assess every answer. The content owner should participate alongside the technical team.

Check facts and useful uncertainty

For supported questions, compare substance with the source, missing conditions, invented amounts or promises and the linked next step. Repeat as a paraphrase and record meaningful differences.

For unknown topics, a good response can be a useful refusal: explain what cannot be confirmed, avoid improvisation or disclosure and offer the appropriate next route. Stop unproductive clarification loops rather than repeating the same failed question.

Test the human route with its context

Ask directly for a staff member, raise a complaint and exercise repeated failure. Check the received original question, concise summary, prior answers and escalation reason. Verify what the customer is told about response time and channel.

A working button does not establish that a person receives enough context or that the promised availability is real. The broader division of responsibilities is covered in AI customer service.

Test the widget on the actual mobile page

Use a narrow view around 360 pixels wide alongside relevant other sizes. Check that the opening control does not cover the contact number, menu, main action or consent controls. Test page scrolling, keyboard operation, touch controls and layout movement.

The website should remain usable if the chatbot loads slowly or fails. An isolated chat test cannot establish that it cooperates with the rest of the page.

Record failures and define a no-go rule

For each deviation, record question, source, expected result, observed result, severity, owner and retest date. Define blocking failures before the run: fabricated facts, inappropriate disclosure, failed complaint handoff or a widget blocking contact can all justify no-go.

Fix and repeat affected cases, plus related boundary cases. Review Dutch chatbot answers separately where that is the visitor language; correct navigation can coexist with an overconfident or misleading answer.

Illustrative release decision

A fictional shop uses ten checks: three supported facts or paraphrases, three unknown cases, two handoffs and two mobile checks. In this tabletop exercise one passes, five fail critically and four need smaller corrections. The decision is no-go while any agreed blocker remains.

After fixes, nine repeated cases and three new boundary cases give twelve checks. Eleven meet expectations and one has a minor wording issue. The fictional team accepts that limited configuration with the wording on its follow-up list. These numbers are not evidence about a real bot.

What a go decision actually proves

It establishes the agreed behaviour for the tested cases, configuration and time. It is not a universal accuracy percentage, legal certification or guarantee about future questions. Keep observation by contact reason and a way to limit or disable the route.

Prevent test emails, payments or bookings from affecting real people. Store only necessary evidence with agreed access and retention. See business process automation when the wider workflow needs redesign, rather than treating a chatbot as the whole process.