A friendly demo call proves that the system can talk. A launch test must prove that it behaves correctly when the conversation is ordinary, messy, or incomplete.
Write expected outcomes first
For every scenario, record the facts the agent may use, the information it must collect, the action it may take, and the fallback. A pass is an outcome, not a feeling about the voice.
Test conversation pressure
Interrupt it. Change an answer. Give a partial address. Ask two questions at once. Stay silent. Call back with the same number. Use a service the business does not offer. These tests reveal whether the agent recovers or confidently heads in the wrong direction.
Test connected systems
Try a valid appointment, a conflict, a closed day, a disconnected calendar, and a slot that disappears during the call. Confirm that the calendar record and the business notification agree with what the caller heard.
Test the stop conditions
Ask for an unapproved price, guarantee, legal conclusion, or unsupported service. Request a person. Trigger the escalation rule, then do not answer the transfer. The fallback should be as carefully tested as the happy path.
Keep testing after launch
Review a sample of real outcomes, protect caller information, and turn failures into new tests. NIST's risk-management guidance emphasizes ongoing measurement and management; that is the right posture for any customer-facing AI system.