Media · Generate · Generate AI video
Generate an AI video from a prompt, an image, two frames, several
images, or a source video. One call: POST {model, parameters}
(capability is optional — derived from which inputs you send).
It blocks and returns the final result — do not poll. Use the
input slot names below as-is for every model; the server adapts the
payload to the model you chose. If something does not fit that model
(a ratio, duration or resolution it lacks, too many images for the
mode), the error says exactly what to change. getModelFields
(type=video) is optional — call it only when you want the model’s
duration/resolution options before choosing.
Why task_id is misleading here
The endpoint also returns a task_id, but it tracks the submission
Celery task that hands the request off to FAL/Kie — it flips to
SUCCESS almost immediately, well before the actual video is rendered.
Do not call getTaskResult on this task_id — it tells you
nothing useful about generation completion. Always poll
getVideoVariation instead. The apikit middleware strips task_id
from this response automatically to avoid confusion.
Storage modes
default— persists into the AiImageArt gallery (recommended).transient— generated without surfacing in the gallery; useful for one-shot agent runs.asset— saves the rendered video as a standalone reusableTldrAsset(no gallery entry).
Authentication
Header authentication of the form Api-Key <token>
Bearer authentication of the form Bearer <token>, where token is your auth token.
Request
Engine id. Currently registered engines:
BOH15 (OmniHuman 1.5 — lipsync: person image + speech audio →
talking video dubbed verbatim; output duration = audio duration,
≤60s @720p / ≤30s @1080p; reference_image only; params:
image_url, audio_url, optional prompt/resolution/
turbo_mode/mask_url — no duration/aspect_ratio),
BSD2 (Seedance 2.0 — text/image/first-last/multi, 480p–1080p, audio),
BSD2F (Seedance 2.0 Fast — same modes, 720p max),
BSD25 (Seedance 2.5 — text/image/first-last/multi/video→video),
K30 (Kling 3.0 — text/image/first-last, std/pro/4K, 3–15s),
KV3TT (Kling V3 Turbo text-to-video, 720p/1080p, 3–15s),
KV3TI (Kling V3 Turbo image-to-video, 720p/1080p, 3–15s),
GT2V (Grok Imagine text-to-video, 6–30s, 480p/720p),
GI2V (Grok Imagine image-to-video, 6–30s, 480p/720p),
W27T (WAN 2.7 text-to-video, neg-prompt, 720p/1080p, 2–15s),
W27I (WAN 2.7 image+first-last-frame, 720p/1080p, 2–15s),
MH23P (MiniMax Hailuo 2.3 Pro image-to-video, 768P/1080P, 6s/10s),
MH23S (MiniMax Hailuo 2.3 Std image-to-video, 768P/1080P, 6s/10s),
GOMNI (Gemini Omni — text/image/multi/v2v, 4–10s, 720p/1080p/4K),
K25TP (Kling 2.5 Turbo Pro — text+image, 5s/10s),
K25TS (Kling 2.5 Turbo Std — image-to-video, 5s/10s),
MH02S (MiniMax Hailuo 02 Std — text/image/first-last, 6s/10s),
MH02P (MiniMax Hailuo 02 Pro — text/image/first-last, 6s),
MH02F (MiniMax Hailuo 02 Fast — image-to-video, 6s/10s),
W25P (WAN 2.5 Preview — text+image, 5s/10s, neg-prompt),
VEO31 (VEO 3.1 — text/image/multi/first-last, 4–8s, audio),
VEO31F (VEO 3.1 Fast — same capabilities, faster),
VEO31L (VEO 3.1 Lite — text/image/first-last).
Call getModelFields with type=video and no model_id to get
the live list with capabilities and parameter schemas.
Always a nested object — these keys never go at the top
level. Input slots use the same names on every model;
each takes an asset UUID from this workspace (from a
previous generateImage/createAsset/uploadAsset call), not
a raw URL. Engine-specific options (duration,
resolution, mode, negative_prompt, generate_audio)
are validated against the chosen model; a value it lacks
is rejected with the supported list.
Optional. Derived from the inputs when omitted: a source video → video_to_video; first/last frames → first_last_frame; several images → multiple_images; one image → reference_image; none → prompt. Set it only to force a mode; a mode the model lacks is rejected by name.
Storage mode for the rendered video — see operation description.
Response
Video generation completed (apikit blocks here until
job_status is terminal — the post-hook polls the
AIArtVariation internally). Final body is the AIArtVariation
with job_status="DONE"; on terminal failure the body is
{error: true, id, art_variation_id, status, message, payload}. Useful keys on success: output.file_url
(rendered video), output.thumbnail_cover_image, id
(variation id), payload.input (echo of the request).
