Generate Voiceover
Skill ID: generate-voiceover · Install skills · Canonical source
Try it
Find a British-English voice for this approved script and save its narration.
The instructions below describe the workflow. Installed skills and connected tools are separate; use only operations exposed by your authorized connection.
Discover the voice and language before synthesizing. Keep approved spoken meaning intact and return reusable audio with honest timing information.
Scope and generation intent
Use the connected Simplified tools and current schemas. Resolve a named or uncertain workspace/teamspace using api_getWorkspaceInfo and accessible membership from api_listTeamspaces; carry numeric space_id into voice discovery, synthesis, storage, and downstream calls.
A request to create narration authorizes the requested synthesis. “Show voices” or “what would it cost?” is discovery only: do not generate or spend credits. Clarify missing script, language, or constraints that change the result before synthesis. Do not add translations, variants, or repeated paid attempts unless requested.
Select an actual voice
Call api_listVoices with the required exact language_code, such as en-GB or en-US. If the request only says English and no context resolves the locale, obtain or state a suitable locale choice before lookup; do not invent an all-locales call. For a saved UUID, api_getVoice resolves its actual details.
search matches locale/language names, not voice names. Voice UUIDs can differ for the same name across locales; re-resolve the name after a language change. Do not fabricate a voice or silently substitute another one. If the requested name and explicit gender/provider constraints conflict, clarify the choice.
For options, show a short relevant list of names, locales, descriptions, supported instruction behavior, and returned samples. Do not promise a model or price from memory.
Prepare the script and duration
Use the approved script verbatim unless adaptation was requested. Resolve pronunciation or ambiguous markup before spending. Plain prose uses use_ssml: false. Use true only for valid supplied/prepared SSML and an appropriately supported voice; do not wrap plain text merely to set the flag. Validate malformed markup or ask about ambiguous meaning instead of speaking broken tags as text.
The generation schema has no duration or rate field. instructions nudges delivery but is not an exact timing control, and unsupported voices ignore it. Word-count estimates are estimates. For “exactly 30 seconds,” resolve the requirement before a paid job: accept approximate narration or arrange separately supported measurement/editing. Do not promise exact runtime, invent duration, silently change approved copy, or regenerate repeatedly to chase timing.
Synthesize and retain the result
Call api_generateAudio with the verified voice_id, text, use_ssml, and storage: "asset" for retained/reused narration ("default" only for an explicitly temporary result). Add supported instructions only when appropriate. There is no separate language/model/duration argument in this generation request; the voice selects the language/provider.
Inspect the actual returned payload; do not require an image/video response shape. Current middleware waits for completion. If only a pending/timeout task identifier returns, continue that exact task_id with api_getTaskResult where exposed, using reasonable intervals and a bounded wait. Stop on terminal failure, and never synthesize again merely because storage or completion metadata is missing.
Retain a returned verified audio asset UUID. If the result has only a usable URL even though persistence was requested, resolve an identifiable stored result when supported; otherwise use api_createAsset with the existing audio URL, then api_getAsset until ready status: 4. Keep signed query strings intact. Do not mistake a voice UUID, task ID, or arbitrary response id for the audio asset UUID. Follow manage-assets for failed/pending storage.
Return the actual voice/locale, audio link, verified asset UUID/readiness, and measured runtime when available or explicit unverified timing. Script narration does not combine audio with a video by itself. Use only exposed composition operations for that handoff, and do not claim an assembled video or published post from an audio result.
