istGalaxy.
Testing & Rehearsal

An agent that fails the exam never faces your customers.

Every change you make to your agent goes through rehearsal conversations first; if it fails, it cannot be published. And the agent that is live sits a real-scenario exam every night — any decline is caught before you hear about it. Quality here isn't a feeling; it's a measurement renewed nightly.

30

real scenarios in every night's exam — from appointments to complaint handovers

Every night

the exam runs; results go into a versioned report card, declines show before morning

1 click

the cost of rolling the live agent back to its previous version

release gate

Rehearsal first, customers second

You changed your agent's behaviour, knowledge or rules — before that change reaches a customer, it goes through realistic rehearsal conversations. The 'customer' in the rehearsal is an AI too: it asks real questions, raises real objections, changes topic.

Fail = no release

If rehearsal results fall below the threshold, the change cannot go live; you fix it first.

A realistic customer

The rehearsal counterpart asks, objects and digresses like a real person.

One-click rollback

See a problem in the live version? Return to the previous one with a single click.

This gate is on for our own agent too: we can't publish a change that fails its own rehearsal.

nightly exam

The live agent sits an exam every night

Being live doesn't mean the testing is over. Every night your agent sits a 30-scenario exam — booking appointments, price questions, complaint handovers, forbidden-phrase checks. Results go into a versioned report card; anything worse than yesterday is visible.

Versioned report card

Every night's result is recorded with its date; you can see which version changed what.

Decline detection

A scenario that passed yesterday and fails today is flagged before morning.

Say-do checking

The exam measures behaviour, not just words: did the agent actually do what it said it did?

Exam scenarios derive from real business flows; scores live in the panel and are never deleted.

observability

A trace for every conversation, a grade for every answer

Every live conversation is traceable: which knowledge the agent looked at, which tools it called, why it answered the way it did. A separate AI referee scores answers across axes like clarity and accuracy; moments your operators correct are captured as learning opportunities.

Conversation trace

The source document, the tools called and the decision steps sit right next to the conversation.

AI referee

Answers are scored across multiple axes; low-scoring patterns surface.

Learning from operators

Answers your team corrects are collected and turned into improvement suggestions — publishing is always your call.

Improvement suggestions never publish themselves; the final word is always human.

The night exam's report card

Every morning you'll find a card like this in your panel — scenario by scenario. (Illustrative; yours is built from your own scenarios.)

  • Writing an appointment request into the calendar
  • A cited answer to a price question
  • Handing a complaint to a human + summary
  • Stopping an unsupported number
  • Refusing an over-cap discount
  • Forbidden-phrase check

…and 24 more scenarios — every night, while you sleep.

Frequently asked

What exactly is a rehearsal conversation?

A practice conversation between your agent and an AI playing the customer. Your agent is tested with realistic questions and objections before any real customer sees it.

Who writes the exam?

Scenarios derive from your real business flows: appointments, pricing, returns, complaint handovers… defined together at setup, growing over time.

What happens on a decline?

It's flagged on the report card and visible in the panel by morning. You can roll the live version back with one click, or fix and re-publish through rehearsal.

Is this extra work for me?

No — the exam runs by itself every night. Your part is glancing at the card in the morning; that's it.

Don't trust your agent. Trust its report card.

We handle the setup in your first week; your first morning report card will be waiting in your panel.