Skip to navigation

Media · Generate · Generate AI video

View as Markdown

Generate an AI video from a prompt, an image, two frames, several images, or a source video. One call: POST {model, parameters} (capability is optional — derived from which inputs you send). It blocks and returns the final result — do not poll. Use the input slot names below as-is for every model; the server adapts the payload to the model you chose. If something does not fit that model (a ratio, duration or resolution it lacks, too many images for the mode), the error says exactly what to change. getModelFields (type=video) is optional — call it only when you want the model’s duration/resolution options before choosing.

Why task_id is misleading here

The endpoint also returns a task_id, but it tracks the submission Celery task that hands the request off to FAL/Kie — it flips to SUCCESS almost immediately, well before the actual video is rendered. Do not call getTaskResult on this task_id — it tells you nothing useful about generation completion. Always poll getVideoVariation instead. The apikit middleware strips task_id from this response automatically to avoid confusion.

Storage modes

  • default — persists into the AiImageArt gallery (recommended).
  • transient — generated without surfacing in the gallery; useful for one-shot agent runs.
  • asset — saves the rendered video as a standalone reusable TldrAsset (no gallery entry).

Authentication

AuthorizationApi-Key

Header authentication of the form Api-Key <token>

OR
AuthorizationBearer

Bearer authentication of the form Bearer <token>, where token is your auth token.

Request

This endpoint expects an object.
modelstringRequiredDefaults to K25TP

Engine id. Currently registered engines: BOH15 (OmniHuman 1.5 — lipsync: person image + speech audio → talking video dubbed verbatim; output duration = audio duration, ≤60s @720p / ≤30s @1080p; reference_image only; params: image_url, audio_url, optional prompt/resolution/ turbo_mode/mask_url — no duration/aspect_ratio), BSD2 (Seedance 2.0 — text/image/first-last/multi, 480p–1080p, audio), BSD2F (Seedance 2.0 Fast — same modes, 720p max), BSD25 (Seedance 2.5 — text/image/first-last/multi/video→video), K30 (Kling 3.0 — text/image/first-last, std/pro/4K, 3–15s), KV3TT (Kling V3 Turbo text-to-video, 720p/1080p, 3–15s), KV3TI (Kling V3 Turbo image-to-video, 720p/1080p, 3–15s), GT2V (Grok Imagine text-to-video, 6–30s, 480p/720p), GI2V (Grok Imagine image-to-video, 6–30s, 480p/720p), W27T (WAN 2.7 text-to-video, neg-prompt, 720p/1080p, 2–15s), W27I (WAN 2.7 image+first-last-frame, 720p/1080p, 2–15s), MH23P (MiniMax Hailuo 2.3 Pro image-to-video, 768P/1080P, 6s/10s), MH23S (MiniMax Hailuo 2.3 Std image-to-video, 768P/1080P, 6s/10s), GOMNI (Gemini Omni — text/image/multi/v2v, 4–10s, 720p/1080p/4K), K25TP (Kling 2.5 Turbo Pro — text+image, 5s/10s), K25TS (Kling 2.5 Turbo Std — image-to-video, 5s/10s), MH02S (MiniMax Hailuo 02 Std — text/image/first-last, 6s/10s), MH02P (MiniMax Hailuo 02 Pro — text/image/first-last, 6s), MH02F (MiniMax Hailuo 02 Fast — image-to-video, 6s/10s), W25P (WAN 2.5 Preview — text+image, 5s/10s, neg-prompt), VEO31 (VEO 3.1 — text/image/multi/first-last, 4–8s, audio), VEO31F (VEO 3.1 Fast — same capabilities, faster), VEO31L (VEO 3.1 Lite — text/image/first-last). Call getModelFields with type=video and no model_id to get the live list with capabilities and parameter schemas.

parametersobjectRequired

Always a nested object — these keys never go at the top level. Input slots use the same names on every model; each takes an asset UUID from this workspace (from a previous generateImage/createAsset/uploadAsset call), not a raw URL. Engine-specific options (duration, resolution, mode, negative_prompt, generate_audio) are validated against the chosen model; a value it lacks is rejected with the supported list.

capabilityenumOptional

Optional. Derived from the inputs when omitted: a source video → video_to_video; first/last frames → first_last_frame; several images → multiple_images; one image → reference_image; none → prompt. Set it only to force a mode; a mode the model lacks is rejected by name.

Allowed values:
storageenumOptionalDefaults to default

Storage mode for the rendered video — see operation description.

Allowed values:

Response

Video generation completed (apikit blocks here until job_status is terminal — the post-hook polls the AIArtVariation internally). Final body is the AIArtVariation with job_status="DONE"; on terminal failure the body is {error: true, id, art_variation_id, status, message, payload}. Useful keys on success: output.file_url (rendered video), output.thumbnail_cover_image, id (variation id), payload.input (echo of the request).

Errors

400
Bad Request Error
401
Unauthorized Error
429
Too Many Requests Error