- The caller's audio is turned into text at the end of their turn.
- A language model reads the transcript, the conversation so far and the employee's instructions, then decides the reply and any approved action such as writing a captured field or booking a slot.
- Text is synthesised to speech and streamed to the caller in chunks, so the first words arrive before the whole answer is finished.
- Summaries, CRM updates, recording finalisation and billing settle outside the live audio path, so they never slow the conversation itself.
Speech becomes text
The caller's audio is turned into text at the end of their turn. Endpoint detection decides when they have finished, which is why a rushed detector creates interruptions and a slow one creates awkward silence.
The turn is understood and decided
A language model reads the transcript, the conversation so far and the employee's instructions, then decides the reply and any approved action such as writing a captured field or booking a slot.
The reply is spoken back
Text is synthesised to speech and streamed to the caller in chunks, so the first words arrive before the whole answer is finished. Barge-in cancels speech that is no longer relevant.
Work continues after the call
Summaries, CRM updates, recording finalisation and billing settle outside the live audio path, so they never slow the conversation itself.



