Correct Dutch does not necessarily mean a correct answer. A chatbot can sound fluent while omitting a condition, applying the wrong source or failing to offer human help. Evaluate each response against the customer's task and approved knowledge.
This English guide is for teams assessing Dutch-language conversations. The English evaluation worksheet has English instructions; Dutch question examples remain labelled in their original language where they test wording.
Use representative wording with synthetic details
Take patterns from actual contact reasons and remove or replace personal data. Include short questions, informal phrasing, misspellings, regional wording and several requests in one message. Do not copy customer names, order numbers or unrestricted chat histories into the worksheet.
Create several formulations for one contact reason while keeping the expected facts the same. Include missing conditions where the correct next step is clarification, rather than an immediate answer.
Establish the source and minimum answer
For every case, name the current authoritative source or an explicit unknown-answer rule. Write required facts, conditions, prohibited claims and handoff conditions before evaluating output.
Test an unknown Dutch price question
For example, the Dutch question “Wat kost spoed?” asks about urgency pricing. If no approved source provides a price, a specific amount is a failure regardless of how naturally it is phrased. This is an illustrative test question, not a statement about an actual service price.
Score dimensions separately
| Dimension | Acceptable result | Failure |
|---|---|---|
| Accuracy | Facts trace to the approved source | Incorrect or invented information |
| Completeness | Necessary conditions and next step are present | A required qualification is missing |
| Language | Understandable, natural Dutch for the audience | Literal translation or confusing terminology |
| Handoff | Appropriate human route with context | No usable route or unnecessary repetition |
An answer can be accurate but incomplete. Name the omitted condition rather than concealing it in a combined score. An average style score cannot compensate for a critical wrong fact.
Review tone without changing certainty
Check form of address, technical vocabulary, long sentences and whether the customer understands what to do. Editing language must not strengthen an uncertain answer or alter the underlying fact.
Where useful, ask a subject owner and a Dutch-speaking reader to assess independently. Resolve differences through the source and agreed rule, not preference alone. Test a complaint and a direct request for a staff member as well as ordinary questions.
Fictional twelve-question review
In a tabletop example, seven responses are correct and complete, three are correct but incomplete and two contain wrong facts. Seven plus three plus two is twelve. An agreed zero-critical-error rule makes the round a no-go despite fluent wording.
The team corrects sources or instructions and repeats failed cases with new paraphrases. The exercise provides no measured accuracy for an existing chatbot. A pass belongs to that tested set, version and time.
Retest the rule as well as the sentence
Keep failures open until both the original case and related variation pass. Stop when sources have no owner or critical errors recur.
Use AI customer service to define responsibility, then check the whole website chatbot journey. A good sentence cannot make a broken human handoff acceptable. For a broader workflow, see business process automation.
