Every change you make to your agent goes through rehearsal conversations first; if it fails, it cannot be published. And the agent that is live sits a real-scenario exam every night — any decline is caught before you hear about it. Quality here isn't a feeling; it's a measurement renewed nightly.
30
real scenarios in every night's exam — from appointments to complaint handovers
Every night
the exam runs; results go into a versioned report card, declines show before morning
1 click
the cost of rolling the live agent back to its previous version
release gate
You changed your agent's behaviour, knowledge or rules — before that change reaches a customer, it goes through realistic rehearsal conversations. The 'customer' in the rehearsal is an AI too: it asks real questions, raises real objections, changes topic.
Fail = no release
If rehearsal results fall below the threshold, the change cannot go live; you fix it first.
A realistic customer
The rehearsal counterpart asks, objects and digresses like a real person.
One-click rollback
See a problem in the live version? Return to the previous one with a single click.
This gate is on for our own agent too: we can't publish a change that fails its own rehearsal.
nightly exam
Being live doesn't mean the testing is over. Every night your agent sits a 30-scenario exam — booking appointments, price questions, complaint handovers, forbidden-phrase checks. Results go into a versioned report card; anything worse than yesterday is visible.
Versioned report card
Every night's result is recorded with its date; you can see which version changed what.
Decline detection
A scenario that passed yesterday and fails today is flagged before morning.
Say-do checking
The exam measures behaviour, not just words: did the agent actually do what it said it did?
Exam scenarios derive from real business flows; scores live in the panel and are never deleted.
observability
Every live conversation is traceable: which knowledge the agent looked at, which tools it called, why it answered the way it did. A separate AI referee scores answers across axes like clarity and accuracy; moments your operators correct are captured as learning opportunities.
Conversation trace
The source document, the tools called and the decision steps sit right next to the conversation.
AI referee
Answers are scored across multiple axes; low-scoring patterns surface.
Learning from operators
Answers your team corrects are collected and turned into improvement suggestions — publishing is always your call.
Improvement suggestions never publish themselves; the final word is always human.
Every morning you'll find a card like this in your panel — scenario by scenario. (Illustrative; yours is built from your own scenarios.)
…and 24 more scenarios — every night, while you sleep.
A practice conversation between your agent and an AI playing the customer. Your agent is tested with realistic questions and objections before any real customer sees it.
Scenarios derive from your real business flows: appointments, pricing, returns, complaint handovers… defined together at setup, growing over time.
It's flagged on the report card and visible in the panel by morning. You can roll the live version back with one click, or fix and re-publish through rehearsal.
No — the exam runs by itself every night. Your part is glancing at the card in the morning; that's it.
We handle the setup in your first week; your first morning report card will be waiting in your panel.