Voice
Realtime Dialog
| Enable Realtime Audio | Global switch turning voice-call entry on or off for this agent. Gates every other setting on this page. |
| Default Voice Display | Starting display mode for a new voice session: Immersive (full-screen, voice-only, canvas pager) or Transcript (scrolling chat-style bubbles). End users can still switch mid-call. |
| Immersive View Welcome Text | Free text shown in the immersive view before any canvas content has arrived. |
Canvas Display
| Automatic Canvas Content Detection | When on, the AI decides response by response what to speak versus show on canvas. |
| Mandatory Canvas Content Types | Free-text list of content types (e.g. chart, table) that must always render on canvas regardless of automatic detection. |
Voice and Transcription
| Voice Persona | Selects the spoken voice used for assistant responses. The preview button speaks the agent's welcome message, falling back to a built-in sample sentence when that is empty. |
| Transcription Model | Selects the model used to transcribe user speech into text; defaults to gpt-4o-transcribe. |
Voice Activity Detection (VAD)
| Smart Turn-Taking | Model-based turn detection instead of purely energy-based; default off. |
| Stop Assistant on Interruption | Whether the assistant stops speaking the instant the user talks, instead of waiting about a second; default off. |
| Sensitivity (eagerness) | Only shown when smart turn-taking is on. Controls how eagerly the model ends the user's turn. |
| Sensitivity (threshold) | Only shown when smart turn-taking is off. Energy-based activation threshold slider; default 0.7. |
Push-to-talk
| Require Hold-to-talk | Shows a hold-to-talk button in the call UI; user can toggle between this and automatic listening at runtime. Enabled by default. |
| Start Calls In | Only shown when hold-to-talk is enabled. Which mode a session opens in: Automatic (default) or Hold-to-talk. |