Media · Generate · List available TTS voices
List voices available to the current workspace, filtered by language.
language_code is the ONLY required parameter — use an exact BCP-47 locale
(en-US, en-GB, fr-FR, hi-IN, ja-JP). Combine with is_premium=true
for a short, high-quality shortlist.
The optional filters (gender, provider, engine, is_premium,
language_name, search) all default to an empty string "", which
means “no filter — return everything”. Leave them at "" unless the user
explicitly asked to filter. Do NOT substitute 0/false for “no filter”:
0 is a REAL value (gender 0 = Male, provider 0 = Google, engine 0 =
Neural), so sending it silently narrows the list. The empty-string default
is the ONLY correct “unset” value.
Response:
voices[]:id(uuid),voice_name,language_code,avatar,voice_sample,gender(0=Male, 1=Female, 3=Neutral),is_premium,description,supports_instructions(bool —trueif the voice honors theinstructionsstyle/pace hint on generateAudio). Useidasvoice_idin generateAudio. To re-resolve a single knownvoice_id(e.g. from a saved preference) without re-listing, usegetVoice.languages[]: distinct{language_code, language_name}pairs.
To resolve a known voice_name (e.g. from an avatar roster) to its id, pass
voice_name + language_code — that pair is unique, and the id differs per
locale, so the name must be re-resolved whenever the language changes. Always
check the entry you use actually has that voice_name rather than trusting
voices[0] — see the voice_name parameter.
Authentication
Header authentication of the form Api-Key <token>
Bearer authentication of the form Bearer <token>, where token is your auth token.
Query parameters
BCP-47 locale, exact match (e.g. en-US, fr-FR). Always required — narrows the list significantly.
Display name, exact match (e.g. English (US)). Use when you only have the display name.
Exact voice name, case-sensitive (e.g. Leda, Puck). NOT unique on its own — a Chirp name exists once per locale (Leda x 53), and each locale has its own distinct id. Always pair with language_code to land a single voice. Use when you hold a name from a curated roster and need its id for generateAudio. ALWAYS confirm the returned entry's voice_name equals the name you asked for before using its id: deployments that predate this filter ignore the parameter and return the whole locale, so trusting voices[0] blindly can hand you a different voice — and a wrong-gender one.
Substring match across language_code and language_name only — it does NOT search voice names. search=Orus returns nothing; to find a voice by name use the voice_name parameter. Use this for genuinely vague locale requests ("any English voice" → search=en).
Filter by TTS provider: 0=Google, 2=ElevenLabs, 3=OpenAI, 4=Microsoft, 5=Google Chirp 3 HD, 6=Gemini Flash TTS. Leave as the default empty string "" to include ALL providers. 0 is a real value (Google), NOT 'any' — only send an integer when the user explicitly names a provider.
Filter by voice gender: 0=Male, 1=Female, 3=Neutral. Leave as the default empty string "" to return ALL genders. 0 is a real value (Male), NOT 'any' — only send an integer when the user explicitly asks for a gender.
Filter by synthesis engine: 0=Neural (higher quality), 1=Standard. Leave as the default empty string "" to return ALL engines. 0 is a real value (Neural), NOT 'any' — only send an integer when the user explicitly asks to restrict by engine.
