Media · Generate · Generate speech from text
Convert text to spoken audio using the voice referenced by voice_id
(get voices from listVoices). storage controls whether the audio is saved.
Set use_ssml: true ONLY when text contains SSML markup like <speak> /
<break> / <emphasis>. For plain prose leave it false — providers handle
punctuation pacing well and SSML support varies.
Storage modes:
default— returns the audio URL only (can expire)asset— also saves it to your asset library for reuse
Authentication
Header authentication of the form Api-Key <token>
Bearer authentication of the form Bearer <token>, where token is your auth token.
Path parameters
Voice UUID from listVoices voices[].id
Request
Text to synthesize. HTML entities will be unescaped server-side.
Set true ONLY when text contains SSML markup like <speak>, <break>, <emphasis>.
default returns the audio URL only; asset also saves a
persistent TldrAsset (library entry) for reuse.
Natural-language style hint controlling tone and delivery
(e.g. "say cheerfully" or audio tags like "[whispers]").
Honored ONLY when the voice's supports_instructions is
true (check via listVoices / getVoice); silently
ignored by voices that don't support it. It nudges delivery
— it is NOT a rate/duration control.
