Models & Voices
The voice agent model catalog — speech-to-text, LLM, text-to-speech, and speech-to-speech providers, with voices, cost, and latency.
The AI Models tab is where you choose what your agent hears with, thinks with, and speaks with. Dynamiq ships a curated catalog: every provider and model listed here has been wired up and measured, so you are picking from combinations that are known to work rather than assembling one yourself.

Agent type
The Agent type toggle picks the architecture:
- Pipeline (STT → LLM → TTS) — you configure three stages independently. The widest choice, the lowest cost, and the mode where tuning turn-taking pays off most.
- Realtime (speech-to-speech) — one model takes audio in and puts audio out. Fewer knobs, lower turn latency, more natural prosody, higher per-minute cost.
Switching type discards the configuration of the mode you're leaving; the builder confirms first.
What every stage needs
ProviderselectrequiredModelselectrequiredConnectionselectrequiredLanguageselectVoiceselectTemperaturenumberBase URLtextEvery model needs a Connection — voice agents call the providers with your credentials and your provider bills you directly.
Cost and latency estimates
Above the stage fields, an estimate bar shows the configuration's Estimated cost (~$x.xx /min, with a bar showing each stage's contribution) and Estimated latency (~NNN ms, the sum of each model's typical first-response time).
Read them as comparisons between configurations, not as a bill:
- Cost uses provider list prices and assumes typical call traffic. It excludes telephony, and it's billed by your own provider accounts. Realtime configurations show a range rather than a single figure.
- Latency is model time only. The turn-taking settings in Advanced also affect what a caller experiences.
- Each estimate is marked independently.
~is a complete estimate;≥is a lower bound, because at least one model in the configuration publishes no figure for that measure. Cost and latency can differ — a configuration can show a~price and a≥latency. A missing price also adds the caption "Some models have no published price." - Speech-to-speech models re-send the conversation each turn, so longer realtime calls cost more per minute than the bar's steady-state number suggests.
Speech-to-text
| Provider | Models | Default |
|---|---|---|
| Deepgram | nova-3, nova-3-medical, nova-2-phonecall, nova-2-conversationalai, nova-2, flux-general-en, flux-general-multi | nova-3 |
| AssemblyAI | universal-3-5-pro, universal-streaming-english, universal-streaming-multilingual | universal-3-5-pro |
| Cartesia | ink-2, ink-whisper | ink-2 |
| ElevenLabs | scribe_v2_realtime | scribe_v2_realtime |
| Speechmatics | enhanced, standard | enhanced |
| OpenAI | gpt-4o-transcribe, gpt-4o-mini-transcribe | gpt-4o-transcribe |
| Groq | whisper-large-v3-turbo | whisper-large-v3-turbo |
| Baseten | whisper-large-v3-streaming — your own deployment, needs a Base URL | whisper-large-v3-streaming |
Deepgram's nova-3 defaults to multilingual rather than English — it detects the language itself. Pick an explicit language when you know the call will be in one.
Some speech-to-text models detect end-of-turn themselves. Deepgram flux-*, Cartesia ink-2, and the AssemblyAI models decide when a caller has finished speaking as part of transcription, which removes a whole endpointing stage from the pipeline and is the single biggest latency win available. See Turn-taking & latency.
Deepgram and AssemblyAI expose extra controls under the model — keyterms (bias recognition toward names and jargon your callers use), endpointing milliseconds, numeral formatting, and profanity filtering.
Large language models
| Provider | Models | Default |
|---|---|---|
| OpenAI | gpt-5.6-luna, gpt-5.4, gpt-5.4-mini, gpt-5.2, gpt-5.1, gpt-5, gpt-5-mini, gpt-5-nano, gpt-4.1-mini, gpt-4.1-nano, gpt-4o-mini, gpt-5.6-terra, gpt-5.6-sol | gpt-5.6-luna |
| Anthropic | claude-haiku-4-5, claude-sonnet-4-6, claude-sonnet-5 | claude-haiku-4-5 |
gemini-3.5-flash-lite, gemini-3.5-flash, gemini-3.6-flash | gemini-3.5-flash-lite | |
| Groq | openai/gpt-oss-20b, openai/gpt-oss-120b, llama-3.3-70b-versatile, qwen/qwen3.6-27b | openai/gpt-oss-20b |
| Cerebras | gpt-oss-120b | gpt-oss-120b |
| Fireworks AI | DeepSeek v4 Flash and Pro, GPT-OSS 20B and 120B, Kimi K2 Turbo, GLM Fast | accounts/fireworks/models/deepseek-v4-flash |
| Together AI | openai/gpt-oss-120b, openai/gpt-oss-20b, Qwen/Qwen3.5-9B, meta-llama/Llama-4-Scout-17B-16E-Instruct, meta-llama/Llama-3.3-70B-Instruct-Turbo | openai/gpt-oss-120b |
| xAI | grok-4.3, grok-4.5, grok-4.20-0309-non-reasoning | grok-4.3 |
| Custom (OpenAI-compatible) | none — you type the model id | — |
For voice, model speed usually matters more than model strength: the caller hears every millisecond of thinking time. Start with the default and move up only if the agent's answers are actually wrong.
Not every model accepts Temperature — the catalog records this per model, verified against each provider's API, and the builder only shows the control where it's supported.
Bringing your own LLM
Select provider Custom (OpenAI-compatible) to point the agent at any OpenAI-compatible chat completions endpoint. The Model field becomes free text — type the model id your endpoint serves — and you supply the Base URL and a Connection holding the API key.
Text-to-speech
| Provider | Models | Voices |
|---|---|---|
| ElevenLabs | eleven_flash_v2_5 (default), eleven_flash_v2, eleven_multilingual_v2 | A curated lineup — Elara, Eddie, Talia, Darian, Maisie, Caleb and more |
| Cartesia | sonic-3.5 (default), sonic-3, sonic-2, sonic-turbo, sonic | 20 curated voices from Cartesia's public library |
| Deepgram | aura-2 (default), aura | The voice is the model id — aura-2-thalia-en, aura-2-apollo-en, and ~40 more |
| OpenAI | gpt-4o-mini-tts-2025-12-15 (default), tts-1 | The standard OpenAI voice set |
| Groq | canopylabs/orpheus-v1-english | Orpheus voices |
| Baseten | orpheus-3b — your own deployment, needs a Base URL | Orpheus voices |
| Kokoro (via DeepInfra) | hexgrad/Kokoro-82M — needs a Base URL, prefilled with DeepInfra's | Kokoro voices |
ElevenLabs Flash v2.5 is the fastest text-to-speech in the catalog and the builder's default. ElevenLabs models expose speed, stability, similarity, style, speaker boost, and a pronunciation dictionary; Cartesia exposes speed, emotion, and (on sonic-3) volume.
Custom voice IDs
Every voice picker ends with a Custom voice ID… option. Choosing it reveals a text field where you paste a voice id from your own provider account — a cloned or fine-tuned voice, for example. Models with no catalog voices show a plain Voice ID field instead.
Realtime (speech-to-speech)
| Provider | Models | Default |
|---|---|---|
| OpenAI | gpt-realtime-2.1-mini, gpt-realtime-2.1 | gpt-realtime-2.1-mini |
gemini-3.1-flash-live-preview, gemini-2.5-flash-native-audio-preview-12-2025 | gemini-3.1-flash-live-preview | |
| xAI | grok-voice-think-fast-2.0, grok-voice-think-fast-1.0 | grok-voice-think-fast-2.0 |
Realtime agents take a single Provider, Model, Voice, and Connection. There are no separate STT and TTS stages to tune, and interruption handling is always on, because these models manage turn-taking themselves.
Temperature appears in realtime mode only for Google's gemini-2* models; no other realtime model exposes it.
Realtime calls report response latency (time to first token) rather than an end-to-end breakdown, because there are no pipeline stages to attribute time to. The Calls tab labels this honestly — see Calls & observability.
The defaults, and why
A new agent starts on a combination chosen for latency rather than for any single benchmark:
| Stage | Default | Reason |
|---|---|---|
| Speech-to-text | Cartesia | ink-2 has native end-of-turn detection, so the pipeline skips an endpointing stage entirely. |
| LLM | OpenAI | A product decision rather than a measurement — no first-token benchmark backs this one, so treat it as a starting point and compare against your own calls. |
| Text-to-speech | ElevenLabs | Flash v2.5 is the fastest text-to-speech measured in the catalog. |
In each case the builder picks the provider, and then whichever model that provider's catalog entry flags as its default.