Voice Agents

Models & Voices

The voice agent model catalog — speech-to-text, LLM, text-to-speech, and speech-to-speech providers, with voices, cost, and latency.

The AI Models tab is where you choose what your agent hears with, thinks with, and speaks with. Dynamiq ships a curated catalog: every provider and model listed here has been wired up and measured, so you are picking from combinations that are known to work rather than assembling one yourself.

The AI Models tab showing the agent type toggle, the estimate bar, and the STT, LLM, and TTS sections

Agent type

The Agent type toggle picks the architecture:

  • Pipeline (STT → LLM → TTS) — you configure three stages independently. The widest choice, the lowest cost, and the mode where tuning turn-taking pays off most.
  • Realtime (speech-to-speech) — one model takes audio in and puts audio out. Fewer knobs, lower turn latency, more natural prosody, higher per-minute cost.

Switching type discards the configuration of the mode you're leaving; the builder confirms first.

What every stage needs

Providerselectrequired
The vendor. Changing it auto-selects that provider's default model and its first voice.
Modelselectrequired
The model. Each option shows its measured latency, price, and quality — for example "300 ms · $0.005/min · 2.7% WER".
Connectionselectrequired
The Connection holding your API key for that provider. Dynamiq auto-selects a system connection where one exists. Calls fail without one.
Languageselect
Speech-to-text only. Options come from the selected provider; a model that declares its own default language snaps to it when you pick the model.
Voiceselect
Text-to-speech and realtime only. Pick from the provider's catalog voices, or choose Custom voice ID… to type one from your own account.
Temperaturenumber
Shown only for models that actually accept it. Leave blank for the model default.
Base URLtext
Shown for providers that serve your own deployment, and for OpenAI-compatible endpoints. It must start with https:// — the builder does not check this, but saving fails if it does not.

Every model needs a Connection — voice agents call the providers with your credentials and your provider bills you directly.

Cost and latency estimates

Above the stage fields, an estimate bar shows the configuration's Estimated cost (~$x.xx /min, with a bar showing each stage's contribution) and Estimated latency (~NNN ms, the sum of each model's typical first-response time).

Read them as comparisons between configurations, not as a bill:

  • Cost uses provider list prices and assumes typical call traffic. It excludes telephony, and it's billed by your own provider accounts. Realtime configurations show a range rather than a single figure.
  • Latency is model time only. The turn-taking settings in Advanced also affect what a caller experiences.
  • Each estimate is marked independently. ~ is a complete estimate; ≥ is a lower bound, because at least one model in the configuration publishes no figure for that measure. Cost and latency can differ — a configuration can show a ~ price and a ≥ latency. A missing price also adds the caption "Some models have no published price."
  • Speech-to-speech models re-send the conversation each turn, so longer realtime calls cost more per minute than the bar's steady-state number suggests.

Speech-to-text

ProviderModelsDefault
Deepgramnova-3, nova-3-medical, nova-2-phonecall, nova-2-conversationalai, nova-2, flux-general-en, flux-general-multinova-3
AssemblyAIuniversal-3-5-pro, universal-streaming-english, universal-streaming-multilingualuniversal-3-5-pro
Cartesiaink-2, ink-whisperink-2
ElevenLabsscribe_v2_realtimescribe_v2_realtime
Speechmaticsenhanced, standardenhanced
OpenAIgpt-4o-transcribe, gpt-4o-mini-transcribegpt-4o-transcribe
Groqwhisper-large-v3-turbowhisper-large-v3-turbo
Basetenwhisper-large-v3-streaming — your own deployment, needs a Base URLwhisper-large-v3-streaming

Deepgram's nova-3 defaults to multilingual rather than English — it detects the language itself. Pick an explicit language when you know the call will be in one.

Some speech-to-text models detect end-of-turn themselves. Deepgram flux-*, Cartesia ink-2, and the AssemblyAI models decide when a caller has finished speaking as part of transcription, which removes a whole endpointing stage from the pipeline and is the single biggest latency win available. See Turn-taking & latency.

Deepgram and AssemblyAI expose extra controls under the model — keyterms (bias recognition toward names and jargon your callers use), endpointing milliseconds, numeral formatting, and profanity filtering.

Large language models

ProviderModelsDefault
OpenAIgpt-5.6-luna, gpt-5.4, gpt-5.4-mini, gpt-5.2, gpt-5.1, gpt-5, gpt-5-mini, gpt-5-nano, gpt-4.1-mini, gpt-4.1-nano, gpt-4o-mini, gpt-5.6-terra, gpt-5.6-solgpt-5.6-luna
Anthropicclaude-haiku-4-5, claude-sonnet-4-6, claude-sonnet-5claude-haiku-4-5
Googlegemini-3.5-flash-lite, gemini-3.5-flash, gemini-3.6-flashgemini-3.5-flash-lite
Groqopenai/gpt-oss-20b, openai/gpt-oss-120b, llama-3.3-70b-versatile, qwen/qwen3.6-27bopenai/gpt-oss-20b
Cerebrasgpt-oss-120bgpt-oss-120b
Fireworks AIDeepSeek v4 Flash and Pro, GPT-OSS 20B and 120B, Kimi K2 Turbo, GLM Fastaccounts/fireworks/models/deepseek-v4-flash
Together AIopenai/gpt-oss-120b, openai/gpt-oss-20b, Qwen/Qwen3.5-9B, meta-llama/Llama-4-Scout-17B-16E-Instruct, meta-llama/Llama-3.3-70B-Instruct-Turboopenai/gpt-oss-120b
xAIgrok-4.3, grok-4.5, grok-4.20-0309-non-reasoninggrok-4.3
Custom (OpenAI-compatible)none — you type the model id—

For voice, model speed usually matters more than model strength: the caller hears every millisecond of thinking time. Start with the default and move up only if the agent's answers are actually wrong.

Not every model accepts Temperature — the catalog records this per model, verified against each provider's API, and the builder only shows the control where it's supported.

Bringing your own LLM

Select provider Custom (OpenAI-compatible) to point the agent at any OpenAI-compatible chat completions endpoint. The Model field becomes free text — type the model id your endpoint serves — and you supply the Base URL and a Connection holding the API key.

Text-to-speech

ProviderModelsVoices
ElevenLabseleven_flash_v2_5 (default), eleven_flash_v2, eleven_multilingual_v2A curated lineup — Elara, Eddie, Talia, Darian, Maisie, Caleb and more
Cartesiasonic-3.5 (default), sonic-3, sonic-2, sonic-turbo, sonic20 curated voices from Cartesia's public library
Deepgramaura-2 (default), auraThe voice is the model id — aura-2-thalia-en, aura-2-apollo-en, and ~40 more
OpenAIgpt-4o-mini-tts-2025-12-15 (default), tts-1The standard OpenAI voice set
Groqcanopylabs/orpheus-v1-englishOrpheus voices
Basetenorpheus-3b — your own deployment, needs a Base URLOrpheus voices
Kokoro (via DeepInfra)hexgrad/Kokoro-82M — needs a Base URL, prefilled with DeepInfra'sKokoro voices

ElevenLabs Flash v2.5 is the fastest text-to-speech in the catalog and the builder's default. ElevenLabs models expose speed, stability, similarity, style, speaker boost, and a pronunciation dictionary; Cartesia exposes speed, emotion, and (on sonic-3) volume.

Custom voice IDs

Every voice picker ends with a Custom voice ID… option. Choosing it reveals a text field where you paste a voice id from your own provider account — a cloned or fine-tuned voice, for example. Models with no catalog voices show a plain Voice ID field instead.

Realtime (speech-to-speech)

ProviderModelsDefault
OpenAIgpt-realtime-2.1-mini, gpt-realtime-2.1gpt-realtime-2.1-mini
Googlegemini-3.1-flash-live-preview, gemini-2.5-flash-native-audio-preview-12-2025gemini-3.1-flash-live-preview
xAIgrok-voice-think-fast-2.0, grok-voice-think-fast-1.0grok-voice-think-fast-2.0

Realtime agents take a single Provider, Model, Voice, and Connection. There are no separate STT and TTS stages to tune, and interruption handling is always on, because these models manage turn-taking themselves.

Temperature appears in realtime mode only for Google's gemini-2* models; no other realtime model exposes it.

Realtime calls report response latency (time to first token) rather than an end-to-end breakdown, because there are no pipeline stages to attribute time to. The Calls tab labels this honestly — see Calls & observability.

The defaults, and why

A new agent starts on a combination chosen for latency rather than for any single benchmark:

StageDefaultReason
Speech-to-textCartesiaink-2 has native end-of-turn detection, so the pipeline skips an endpointing stage entirely.
LLMOpenAIA product decision rather than a measurement — no first-token benchmark backs this one, so treat it as a starting point and compare against your own calls.
Text-to-speechElevenLabsFlash v2.5 is the fastest text-to-speech measured in the catalog.

In each case the builder picks the provider, and then whichever model that provider's catalog entry flags as its default.

Where to go next

On this page