A fluent conversation during a demo is only the start of assessing AI reception. Your practice also needs to know what appears in its systems afterwards, what the receptionist sees and what happens when a task cannot be completed.
The scenarios below are practical acceptance tests to discuss with your practice and supplier. They do not describe confirmed capabilities of a particular product. Adapt each test to the agreed scope and use a test environment with fictional data.
Agree on what “done” means
Create a simple record for each test: conversation goal, inputs, expected outcome, actual outcome, evidence and any correction required. Decide whether the system should make a booking, pass a proposal to reception or only collect a callback request. These are three different outcomes with different acceptance criteria.
NIST AI RMF 1.0 covers defining system tasks, documenting measurements and assigning oversight responsibilities. The scenarios below are an editorial application of those principles to telephone reception. They are not an official NIST checklist or certification.
1. A straightforward request for an appointment
The caller makes an unambiguous request. Check both the spoken response and the result in the place where reception is expected to work. Do the confirmed time, appointment type and practitioner match the agreement? If the implementation only collects callback requests, it should not present one as a completed booking.
2. No appointment meets the request
Prepare a calendar with no matching availability. Check that the caller receives accurate information and an agreed alternative: another date, a handoff to reception or a callback request. Record who is responsible for the next step.
3. Correcting a name or phone number
Deliberately give an incorrect detail and then correct it. Check that the final record contains the corrected value and that the team does not receive two conflicting versions. Use a fictional identity and a number designated for testing.
4. Changing a choice during the conversation
The caller first chooses Monday, then asks for Wednesday. Check the final result and make sure the first proposal has not left an unintended booking behind. Both the caller and reception should be able to understand the outcome.
5. Two conversations requesting the same slot
If the system is expected to write to the calendar, run a controlled test in which two conversations select the same slot. The acceptance criteria concern calendar consistency and an accurate explanation to the second caller. Ask the supplier to demonstrate the final state, not just play a recording of one conversation.
6. A request to change an existing appointment
First agree how the practice and supplier will establish that a caller is entitled to change the appointment. Test both a valid and an invalid set of details. Check that information about another person is not disclosed and that the original slot is handled according to the agreed process.
7. A question the knowledge base cannot answer
Ask about a service that does not exist, an unconfirmed price or a practitioner's availability outside the recorded rota. The expected outcome may be an acknowledgement that the information is unavailable and a handoff to the team. An invented answer should not count as successful handling.
8. Asking to speak to a person
Ask explicitly for a receptionist. Test the handoff during and outside opening hours. If a transfer is unavailable, the caller should know what will actually happen next. Agree where the request goes and who is responsible for receiving it.
9. A clinical question or a request needing special handling
The medical team should prepare or approve the scenario and expected response. The test checks whether the system follows the agreed boundary between administrative support and clinical advice. It does not replace or establish a clinical procedure.
10. An interrupted call followed by another call
End the call before final confirmation, then call again. Check that the result matches the agreed logic: a booking, an incomplete request or no record. Reception should be able to identify the state without guessing from two partial conversations.
11. The practice system is temporarily unavailable
In a controlled test environment, simulate an integration that does not respond. Check that the voicebot distinguishes an attempted booking from a confirmed one and that information intended for reception is retained. When the system returns, verify that retrying the operation has not created a duplicate.
12. A configuration change
After changing opening hours, an appointment type or a handoff rule, repeat the relevant earlier tests. Keep the date and configuration version with the results. This makes it clear what was tested and whether a later correction changed a previously working path.
How should you decide after testing?
Record each scenario's result and keep a list of issues with an owner and a retest date. Do not hide a critical error in an average across many successful conversations. An incorrect booking or disclosure of another person's details needs its own corrective action and decision.
There is no universal success threshold for every practice. Agree on the scope of independent action, human handoffs and launch criteria before testing. A pilot can start with a limited, tested scope and expand after further acceptance checks.
If you are comparing AI reception services, give suppliers the same set of scenarios. You will have evidence to discuss a specific workflow, rather than judging only how the voice sounds.
Where should your practice start?
Choose the most common reasons people call. For each, prepare an example of a correct result and an exception that needs help from reception. If you are planning a pilot, explore AI reception for a dental practice and the receptionOS Voicebot. Confirm calendar access, data access and handoff capabilities for the specific implementation.

