Skip to navigation

Media · Generate · Generate speech from text

View as Markdown

Convert text to spoken audio using the voice referenced by voice_id (get voices from listVoices). storage controls whether the audio is saved.

Set use_ssml: true ONLY when text contains SSML markup like <speak> / <break> / <emphasis>. For plain prose leave it false — providers handle punctuation pacing well and SSML support varies.

Storage modes:

  • default — returns the audio URL only (can expire)
  • asset — also saves it to your asset library for reuse

Authentication

AuthorizationApi-Key

Header authentication of the form Api-Key <token>

OR
AuthorizationBearer

Bearer authentication of the form Bearer <token>, where token is your auth token.

Path parameters

voice_idstringRequiredformat: "uuid"

Voice UUID from listVoices voices[].id

Request

This endpoint expects an object.
textstringRequired

Text to synthesize. HTML entities will be unescaped server-side.

use_ssmlbooleanOptionalDefaults to false

Set true ONLY when text contains SSML markup like <speak>, <break>, <emphasis>.

storageenumOptionalDefaults to default

default returns the audio URL only; asset also saves a persistent TldrAsset (library entry) for reuse.

Allowed values:
instructionsstringOptional

Natural-language style hint controlling tone and delivery (e.g. "say cheerfully" or audio tags like "[whispers]"). Honored ONLY when the voice's supports_instructions is true (check via listVoices / getVoice); silently ignored by voices that don't support it. It nudges delivery — it is NOT a rate/duration control.

Response

The generated speech audio.

Errors

400
Bad Request Error
401
Unauthorized Error
403
Forbidden Error
429
Too Many Requests Error