Build with Veevid API
Generate AI videos, images, speech and music, and run video tools like upscaling and lip-sync. One API key, 55 models and tools — the same ones you use on veevid.ai.
Get Your API Key →Quick Start
1. Get Your API Key
Create an API key at veevid.ai/settings/api-keys. Save it securely — it's only shown once.
mkdir -p ~/.config/veevid
echo "vv_sk_your_key_here" > ~/.config/veevid/api_key2. Get a Quote
Before generating, check the exact credit cost. A quote is free and never starts a generation:
curl -X POST https://veevid.ai/api/quote \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"mode": "veo3",
"generation_type": "text-to-video",
"video_quality": "standard"
}'Response:
{
"required_credits": 30,
"current_balance": 451,
"sufficient": true
}3. Generate a Video
curl -X POST https://veevid.ai/api/generate-video \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"prompt": "A golden retriever running through sunflowers, cinematic lighting",
"mode": "veo3",
"generation_type": "text-to-video",
"aspect_ratio": "16:9",
"video_quality": "standard"
}'Response:
{
"generation_id": "vg_abc123",
"status": "processing",
"video_url": null,
"image_url": null,
"provider_task_id": "...",
"video_quality": "standard"
}The same endpoint generates images and audio too — pick a mode and generation_type from Models.
4. Check Status
curl https://veevid.ai/api/video-generation/GENERATION_ID/status \
-H "Authorization: Bearer YOUR_API_KEY"When complete:
{
"id": "vg_abc123",
"status": "completed",
"media_kind": "video",
"video_url": "https://cdn.veevid.ai/generated-videos/...mp4",
"credits_used": 30,
...
}Typical generation time: 60–180 seconds for videos, a few seconds to a minute for images and audio.
Authentication
All API requests require a Bearer token:
Authorization: Bearer vv_sk_your_key_here- Keys start with
vv_sk_prefix - Create and manage keys at /settings/api-keys (key management is only available while signed in on the website, not via the API)
- Each key shares your account's credit balance
- Maximum 5 active keys per account
Security tips
- Never expose your API key in client-side code
- Use environment variables or secure config files
- Rotate keys if you suspect a leak
MCP Server
Veevid is also a remote MCP (Model Context Protocol) server, so AI assistants such as Claude Code, Cursor and Claude Desktop can list models, quote, generate and fetch results for you. The MCP tools run the same endpoints as the REST API on this page — same API key, same models, same credit balance and the same credit cost as on the website (see pricing).
| Setting | Value |
|---|---|
| Endpoint | https://veevid.ai/api/mcp |
| Transport | Streamable HTTP (POST, stateless) |
| Authentication | Header Authorization: Bearer vv_sk_your_key_here. API keys only — there is no OAuth sign-in, so use one of the setups below rather than a sign-in-based connector. |
1. Get an API Key
Sign in on the website and create a key at veevid.ai/settings/api-keys. Save it securely — it's only shown once. In the examples below, replace vv_sk_your_key_here with your key.
2. Connect Your Client
Claude Code
claude mcp add --transport http veevid https://veevid.ai/api/mcp \
--header "Authorization: Bearer vv_sk_your_key_here"The server is added to the current project. Add --scope user to use it in all your projects.
Cursor
Add this to ~/.cursor/mcp.json (all projects) or .cursor/mcp.json in a project:
{
"mcpServers": {
"veevid": {
"url": "https://veevid.ai/api/mcp",
"headers": {
"Authorization": "Bearer vv_sk_your_key_here"
}
}
}
}To keep the key out of the file, write "Bearer ${env:VEEVID_API_KEY}" and set VEEVID_API_KEY in your environment.
Claude Desktop
Claude Desktop's config file runs local servers, so connect through the mcp-remote bridge (requires Node.js). Open Settings → Developer → Edit Config (claude_desktop_config.json), add the server, then fully quit and restart Claude Desktop:
{
"mcpServers": {
"veevid": {
"command": "npx",
"args": [
"-y",
"mcp-remote",
"https://veevid.ai/api/mcp",
"--header",
"Authorization:${VEEVID_AUTH}"
],
"env": {
"VEEVID_AUTH": "Bearer vv_sk_your_key_here"
}
}
}
}Keep Authorization:${VEEVID_AUTH} without a space and put the key in env — spaces inside args break on Windows.
3. Tools
get_quote and generate take mode, generation_type and params, where params holds every other field of the request body (prompt, duration, aspect_ratio, video_quality, image, images, video, audio…) exactly as listed under Models. Input files are https URLs. Results are JSON. A failed call is marked as an error and carries an error message, plus http_status when the underlying API endpoint rejected the request.
| Tool | What it does |
|---|---|
list_models | Lists models (mode, generation_type, name, category, summary), optionally filtered by category: video, image, audio, video-tool or effect. Pass mode to get that model’s parameters, input files, upload folders, rules and an example request. |
get_credits | Returns your current credit balance. |
get_quote | Returns the exact credit cost of a request (required_credits, current_balance, sufficient) without running it. Free. |
generate | Starts a video, image or audio generation and returns its generation_id. Spends credits; they are refunded automatically if the job fails. |
get_generation | Returns the status and result of a generation by id — video_url holds the output file for video, image and audio. Set wait_seconds (up to 40) to wait for it to finish. |
upload_file | Uploads an image, video or audio file from source_url (https, up to 20MB) or data_base64 (up to 10MB) and returns a url to pass to generate. Use the upload folder list_models gives for the model. |
create_upload_url | Returns a presigned URL to PUT a large file straight to storage. Currently only for veed-clean-audio (folder remove-background-noise/input, up to 512MB). |
list_creations | Lists your past generations, newest first. Optional page, page_size (up to 50), generation_type and include_audio. |
4. Example: Quote, Generate, Get the Result
Ask your assistant something like "Make an 8-second video of a golden retriever running through sunflowers with Veo 3.1". It calls the tools in this order (numbers below are placeholders):
get_quote — arguments:
{
"mode": "veo3",
"generation_type": "text-to-video",
"params": {
"prompt": "A golden retriever running through sunflowers, cinematic lighting",
"video_quality": "standard",
"resolution": "720p",
"duration": "8",
"aspect_ratio": "16:9"
}
}Result:
{
"required_credits": <cost of this request>,
"current_balance": <your balance>,
"sufficient": true
}The server tells the assistant to confirm the cost with you before generating.
generate — the same arguments as get_quote. Result:
{
"generation_id": "vg_abc123",
"status": "processing",
"video_url": null,
...
}get_generation — arguments:
{
"id": "vg_abc123",
"wait_seconds": 40
}Result when completed:
{
"id": "vg_abc123",
"status": "completed",
"media_kind": "video",
"video_url": "https://cdn.veevid.ai/generated-videos/...mp4",
"error_message": null,
"credits_used": <credits charged>,
...
}Videos usually take 1–5 minutes. While status is still "processing", the assistant calls get_generation again. If generate times out, check list_creations before retrying — the job may already have started.
API Reference
Base URL: https://veevid.ai. All endpoints accept and return JSON unless noted.
POST /api/quote
Get the exact credit cost before generating. Send the same body you will send to /api/generate-video, including the prompt — GPT Image, Qwen Image, text-to-speech and music models check it. File URLs are only needed to quote speech-to-text and noise removal (the audio URL); for other models you can quote before uploading. The quote endpoint does not read video URLs and, for some models, does not count the images array, so describe your input files with these fields — each model's rules say which ones it needs:
| Quote field | Description |
|---|---|
input_image_count | Number of images (0–16), for the models whose rules ask for it |
has_video_input | Always true when you send reference videos, or a source video to edit or extend (not for character swap or the insert-shot source video) |
input_video_count | Number of reference videos (also read by MiniMax H3 Max and the insert-shot model) |
input_video_duration | Length in seconds of the input / reference video(s) |
reference_video_duration | Total length of the reference videos for an inserted shot |
| has_audio_input, input_audio_count, input_audio_duration | The same for audio inputs |
Request:
{
"mode": "kling-3",
"generation_type": "text-to-video",
"duration": "10",
"model_version": "kling-3-standard",
"generate_audio": true
}Response:
{
"required_credits": 250,
"current_balance": 451,
"sufficient": true
}POST /api/generate-video
Start a generation. Despite the name, this single endpoint handles every model: video, image, audio and video tools. Each model accepts its own parameters — see Models for the full list. Credits are deducted when the job is submitted and refunded automatically if it fails, unless the model's rules say otherwise.
Common Parameters:
| Parameter | Type | Description |
|---|---|---|
mode | string | Model selector (see Models). Required in practice — the server default is a legacy model. |
generation_type | string | What to do with the model, e.g. "text-to-video", "image-to-video", "reference-to-video", "text-to-image", "image-to-image", "text-to-speech", "video-upscale". Defaults to "text-to-video". |
prompt | string | Text prompt (max 20,000 characters; some models allow less). Prompts are sent to the model as-is — the automatic prompt translation on the website is not applied to API calls, so write prompts in English for best results. |
aspect_ratio | string | "16:9", "9:16", "1:1" etc. — allowed values depend on the model |
duration | string | Output length in seconds, as a string (model-specific). For tools it is often the length of your input video. |
video_quality | string | Tier or resolution, depending on the model (e.g. Veo 3.1 tier "lite" / "standard" / "pro"; "720p" / "1080p"; image models "1K" / "2K" / "4K") |
resolution | string | Output resolution for models that take it separately (e.g. Veo 3.1 "720p" / "1080p" / "4k") |
image | string | Single image URL (first frame, character photo…) |
images | string[] | Multiple image URLs (start/end frames, reference images, images to edit) |
video | string | Single video URL (source video for tools) |
videos | string[] | Reference video URLs |
audio | string | Audio URL (driving audio, audio to transcribe…). Some models take an audios array instead. |
model_version | string | Sub-variant (e.g. "kling-3-pro", "seedance-2.0-fast") |
Response:
{
"generation_id": "vg_abc123",
"status": "processing",
"video_url": null,
"image_url": null,
"provider_task_id": "...",
"video_quality": "standard"
}Most jobs return "processing" — poll the status endpoint. A few return "completed" right away with video_url (videos) or image_url (images).
GET /api/video-generation/{generation_id}/status
Poll for completion. Works for every kind of generation — videos, images and audio.
Response (completed):
{
"id": "vg_abc123",
"status": "completed",
"media_kind": "image",
"video_url": "https://cdn.veevid.ai/images/...jpg",
"error_message": null,
"credits_used": 5,
"created_at": "2026-10-06T08:00:00.000Z",
"completed_at": "2026-10-06T08:01:37.000Z",
"video_metadata": { "content_type": null, "file_name": null, "file_size": null },
"provider_task_id": "...",
"video_quality": "1K",
"result_images": ["https://cdn.veevid.ai/images/...jpg"]
}| Field | Description |
|---|---|
status | "processing" → "completed" or "failed" |
media_kind | "video", "image" or "audio" |
video_url | The result file for every media kind — video, image or audio (for images it is the first image) |
result_images | All output images, when a model returns more than one |
error_message | Why the generation failed (credits are normally refunded automatically) |
| lyrics, transcript, timestamps | Extra audio results: song lyrics (Lyria), the transcript (Scribe V2) and word timestamps (Eleven v4, when requested) |
seedance25_draft | Only for Seedance 2.5 / Veevid 2.5 Pro 480p drafts: { upgradable, upgrade_credits, expires_at, upgrade_generation_id } — see upgrade-draft below |
POST /api/video-generation/{generation_id}/upgrade-draft
Seedance 2.5 and Veevid 2.5 Pro videos generated at 480p are usually drafts (the status response includes seedance25_draft). A completed draft can be upgraded once, within 7 days, to a 1080p version with the same prompt, inputs and seed. No request body is needed; the upgrade is a new generation with its own generation_id to poll. Get its price first by sending only the draft's id to the quote endpoint:
curl -X POST https://veevid.ai/api/quote \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "draft_upgrade_id": "vg_abc123" }'
curl -X POST https://veevid.ai/api/video-generation/vg_abc123/upgrade-draft \
-H "Authorization: Bearer YOUR_API_KEY"Response:
{
"generation_id": "vg_def456",
"status": "processing",
"credits_used": 790,
"draft_id": "vg_abc123"
}Returns 409 if the draft is not completed yet or was already upgraded, 410 once it has expired, and 502 if the upgrade could not be submitted (the credits are refunded).
GET /api/user/creations
List your generations, newest first.
| Query | Description |
|---|---|
page | Page number, starting at 1 (default 1) |
pageSize | Items per page (default 12, max 50) |
generationType | Only return one generation_type, e.g. "text-to-image" (default: all) |
includeAudio | Audio generations are left out unless you pass includeAudio=1 |
curl "https://veevid.ai/api/user/creations?page=1&pageSize=20&includeAudio=1" \
-H "Authorization: Bearer YOUR_API_KEY"Response:
{
"creations": [
{
"id": "vg_abc123",
"mode": "veo3",
"generationType": "text-to-video",
"status": "completed",
"prompt": "A golden retriever running through sunflowers...",
"videoUrl": "https://cdn.veevid.ai/generated-videos/...mp4",
"creditsUsed": 30,
"createdAt": "2026-10-06T08:00:00.000Z",
"mediaKind": "video",
...
}
],
"totalPages": 3,
"currentPage": 1,
"totalCount": 41,
"hasMore": true
}Note that list items use camelCase field names (videoUrl, generationType), unlike the status endpoint.
GET | DELETE /api/user/creations/{id}
GET returns one generation with the same fields as a list item. DELETE deletes it, including its result files, and returns { "success": true }.
GET /api/credits
Check your current credit balance.
{
"success": true,
"credits": 451,
"balance": 451
}File Uploads
Image, video and audio inputs are passed as URLs. Most models accept any public https URL, but some tools only accept files uploaded to Veevid (noted in each model's rules). Upload a local file first, then pass the returned url.
POST /api/storage/upload
multipart/form-data with a file and a folder. Use the upload folder listed for the model's input — the folder decides which file types and sizes are accepted (images: JPEG/PNG/WebP, 10MB by default; video and audio folders allow larger files). Requests to this endpoint are capped at 100MB by our hosting platform, even where a folder or model allows more; for bigger files pass a public https URL if the model accepts one.
curl -X POST https://veevid.ai/api/storage/upload \
-H "Authorization: Bearer YOUR_API_KEY" \
-F "file=@./photo.jpg" \
-F "folder=image-generator"Response:
{
"url": "https://cdn.veevid.ai/image-generator/...jpg",
"key": "image-generator/...jpg"
}POST /api/storage/presign-upload
Direct upload for large files, currently only for the veed-clean-audio model (folder remove-background-noise/input, up to 512MB). Send { "folder", "filename", "size" } (size in bytes), then PUT the raw file to uploadUrl with exactly the returned contentType as the Content-Type header, and pass url to the model. Accepted extensions: .mp3 .wav .m4a .aac .ogg .opus .flac .mp4 .mov .m4v .webm. Returns 402 if your balance is below the minimum charge for one minute of audio.
{
"uploadUrl": "https://...r2.cloudflarestorage.com/...&X-Amz-Signature=...",
"url": "https://cdn.veevid.ai/remove-background-noise/input/...mp3",
"key": "remove-background-noise/input/...mp3",
"contentType": "audio/mpeg"
}Models
Every model below is called through POST /api/generate-video with its mode and generation_type. Open a model to see its parameters, input files, rules and an example request body. Parameters marked affects price change the credit cost — use /api/quote for the exact amount. Defaults are what the API uses when you leave a field out, which is not always what the website preselects, so send the values you care about explicitly.
Video Models
Text-to-video, image-to-video and reference-to-video generation.
| Model | mode | generation_type |
|---|---|---|
| Google Veo 3.1 | veo3 | text-to-video, image-to-video, reference-to-video |
| Runway Gen-4 Turbo | runway | text-to-video, image-to-video |
| Grok Imagine | grok-imagine | text-to-video, image-to-video |
| Kling 2.6 | kling-2-6 | text-to-video, image-to-video |
| Kling 3.0 | kling-3 | text-to-video, image-to-video |
| LTX 2.5 | ltx-2-5 | text-to-video, image-to-video |
| LTX 2.3 | ltx-2-3 | text-to-video, image-to-video |
| Wan Animate | wan-animate | image-to-video |
| Seedance 1.5 Pro | seedance-1.5-pro | text-to-video, image-to-video |
| Veevid 1.0 Pro | veevid-1.0-pro | text-to-video, image-to-video |
| Veevid 2.0 Pro | veevid-2.0-pro | text-to-video, image-to-video |
| Veevid 2.5 Pro | veevid-2.5-pro | text-to-video, image-to-video |
| Seedance 2.0 | seedance-2.0 | text-to-video, image-to-video, reference-to-video |
| Seedance 2.5 | seedance-2.5 | text-to-video, image-to-video, reference-to-video |
| Wan 2.6 | wan-2-6 | text-to-video, image-to-video, reference-to-video |
| Wan 2.7 | wan-2.7 | text-to-video, image-to-video, reference-to-video |
| Wan 3.0 | wan-3.0 | text-to-video, image-to-video, reference-to-video |
| MiniMax H3 | minimax-h3 | text-to-video, image-to-video, reference-to-video |
| MiniMax H3 Max | minimax-h3-max | text-to-video, image-to-video, reference-to-video |
| MiniMax H3 Max Camera Controls | minimax-h3-max-camera | camera-controls |
| MiniMax H3 Max Lip Sync | minimax-h3-max-lip-sync | lip-sync |
| Gemini Omni | gemini-omni | text-to-video, image-to-video, reference-to-video |
| Boreal | creatify-boreal | text-to-video, image-to-video |
Google Veo 3.1 mode: veo3
text-to-video
Generates a 4-8 second video with native audio from a text prompt using Google Veo 3.1 (Lite / Fast / Quality tiers).
| Parameter | Values & notes |
|---|---|
| promptstring · required | Scene description (required, non-empty, max 20000 characters). |
| video_qualitystringaffects pricedefault "standard" | "lite", "standard", "pro"Tier: lite = Lite, standard = Fast, pro = Quality. Always send video_quality. |
| resolutionstringaffects pricedefault "720p" | "720p", "1080p", "4k"Output resolution. |
| durationstringdefault "8" | "4", "6", "8"Seconds: "4", "6" or "8" (default 8). |
| aspect_ratiostringdefault "16:9" | "16:9", "9:16"16:9 (landscape) or 9:16 (portrait). |
| watermarkstring | Optional custom watermark text (max 200 characters). |
Rules
- resolution, when sent, must be 720p, 1080p or 4k (422 otherwise).
- duration, when sent, must be '4', '6' or '8' (400 otherwise); omitted duration becomes '8'.
- Send lite, standard or pro: any other accepted video_quality value (e.g. "1080p") is treated as the Quality (pro) tier — set the output size with
resolution. - Price depends only on the video_quality tier and resolution; duration and aspect_ratio do not change the price.
Example request body
{
"mode": "veo3",
"generation_type": "text-to-video",
"prompt": "A golden retriever running through sunflowers, cinematic lighting",
"video_quality": "standard",
"resolution": "720p",
"duration": "8",
"aspect_ratio": "16:9"
}image-to-video
Animates a first frame (optionally to a last frame) into a 4-8 second Veo 3.1 video with native audio.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Motion/scene description (required, non-empty, max 20000 characters). |
| imagesstring[] · required | [first_frame_url] or [first_frame_url, last_frame_url]. A single image string is also accepted. |
| video_qualitystringaffects pricedefault "standard" | "lite", "standard", "pro"Tier: lite = Lite, standard = Fast, pro = Quality. Always send video_quality. |
| resolutionstringaffects pricedefault "720p" | "720p", "1080p", "4k"Output resolution; same price table as text-to-video. |
| durationstringdefault "8" | "4", "6", "8"Seconds: "4", "6" or "8" (default 8). |
| aspect_ratiostringdefault "16:9" | "16:9", "9:16"16:9 (landscape) or 9:16 (portrait). |
| watermarkstring | Optional custom watermark text (max 200 characters). |
Input files
images: 1-2 image URLs: first frame, optional last frame (max 10MB per image)
Upload folders for /api/storage/upload
- image:
video-generator-veo3
Rules
- image or a non-empty images array is required (400 'Image is required for image-to-video generation').
- Send 1 or 2 images: the first frame and an optional last frame.
- resolution must be 720p/1080p/4k (422) and duration 4/6/8 (400) when sent.
- Send lite, standard or pro: any other accepted video_quality value (e.g. "1080p") is treated as the Quality (pro) tier — set the output size with
resolution. - Price depends only on the video_quality tier and resolution.
- To guide generation with reference images (rather than first/last frames), use generation_type reference-to-video.
Example request body
{
"mode": "veo3",
"generation_type": "image-to-video",
"prompt": "The woman slowly turns toward the camera and smiles",
"images": [
"https://example.com/first-frame.jpg"
],
"video_quality": "standard",
"resolution": "720p",
"duration": "8",
"aspect_ratio": "16:9"
}reference-to-video
Generates an 8-second Veo 3.1 video that keeps subjects consistent with 1-3 reference images.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Scene description (required, non-empty). |
| imagesstring[] · required | 1-3 reference image URLs. |
| video_qualitystringaffects pricedefault "standard" | "lite", "standard"Only the Lite (lite) and Fast (standard) tiers accept reference images. |
| resolutionstringaffects pricedefault "720p" | "720p", "1080p", "4k" |
| durationstringdefault "8" | "8"Locked to 8 seconds. |
| aspect_ratiostringdefault "16:9" | "16:9"Send 16:9; reference mode supports 16:9 only. |
| watermarkstring | Optional custom watermark text (max 200 characters). |
Input files
images: 1-3 reference image URLs (max 10MB per image)
Upload folders for /api/storage/upload
- image:
video-generator-veo3
Rules
- images must contain 1-3 URLs (400 'Reference-to-video requires 1-3 images').
- video_quality, when sent, must be lite or standard; the Quality tier is rejected with 422.
- duration, when sent, must be '8' (422).
- resolution must be 720p/1080p/4k (422) when sent.
Example request body
{
"mode": "veo3",
"generation_type": "reference-to-video",
"prompt": "The character from the reference walks through a neon-lit street at night",
"images": [
"https://example.com/character.png"
],
"video_quality": "standard",
"resolution": "720p",
"duration": "8",
"aspect_ratio": "16:9"
}Runway Gen-4 Turbo mode: runway
text-to-video
Generates a 5 or 10 second video from a text prompt with Runway Gen-4 Turbo.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Scene description (required, non-empty, max 20000 characters). |
| video_qualitystring · requiredaffects price | "720p", "1080p"Output resolution: 720p or 1080p. Required. |
| durationstringaffects pricedefault "5" | "5", "10"Seconds. |
| aspect_ratiostringdefault "16:9" | "16:9", "1:1", "9:16", "4:3", "3:4"Output aspect ratio; use one of the listed values. |
Rules
- duration must be 5 or 10 (400 'Runway only supports 5s or 10s durations').
- 1080p is only allowed with 5-second duration (400 '1080p is only supported for 5-second Runway videos').
- video_quality is required and must be 720p or 1080p; any other value (or omitting it) makes the request fail without charging.
Example request body
{
"mode": "runway",
"generation_type": "text-to-video",
"prompt": "A paper boat drifting down a rainy city gutter, macro shot",
"duration": "5",
"video_quality": "720p",
"aspect_ratio": "16:9"
}image-to-video
Animates a single image into a 5 or 10 second video with Runway Gen-4 Turbo.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Motion description (required, non-empty). |
| imagestring · required | Single image URL used as the start image. images[0] is also accepted. |
| video_qualitystring · requiredaffects price | "720p", "1080p"Output resolution: 720p or 1080p. Required (see constraints). |
| durationstringaffects pricedefault "5" | "5", "10"Seconds. |
| aspect_ratiostringdefault "16:9" | "16:9", "1:1", "9:16", "4:3", "3:4"Output aspect ratio; use one of the listed values. |
Input files
image: 1 image URL (start frame; max 10MB)
Upload folders for /api/storage/upload
- image:
video-generator-runway
Rules
- image or a non-empty images array is required (400 'Image is required for image-to-video generation').
- duration must be 5 or 10 (400).
- 1080p is only allowed with 5-second duration (400).
- video_quality is required and must be 720p or 1080p; omitting it makes the request fail.
Example request body
{
"mode": "runway",
"generation_type": "image-to-video",
"prompt": "The lighthouse beam sweeps across stormy waves",
"image": "https://example.com/lighthouse.jpg",
"duration": "5",
"video_quality": "720p",
"aspect_ratio": "16:9"
}Grok Imagine mode: grok-imagine
text-to-video
Generate a 6-30 second video with synchronized audio from a text prompt using xAI Grok Imagine.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt (required, non-empty). |
| durationstringaffects pricedefault "6" | 6–30Integer seconds sent as a string, 6-30. |
| video_qualitystringaffects pricedefault "720p" | "480p", "720p"Output resolution: 480p or 720p. Always send video_quality. |
| grokModestringdefault "normal" | "fun", "normal", "spicy"Grok style mode: fun, normal or spicy. |
| aspect_ratiostringdefault "16:9" | "2:3", "3:2", "1:1", "9:16", "16:9"Output aspect ratio; use one of the listed values. Always send aspect_ratio explicitly (16:9 is used when omitted). |
Rules
- duration must be an integer string between 6 and 30 (400 otherwise); omitted duration defaults to 6.
- video_quality, when sent, must be 480p or 720p (400 otherwise).
- grokMode must be fun, normal or spicy; aspect_ratio must be one of the listed values.
- prompt must be non-empty.
Example request body
{
"mode": "grok-imagine",
"generation_type": "text-to-video",
"prompt": "A cat DJ spinning records in a neon club, crowd cheering",
"duration": "6",
"video_quality": "480p",
"aspect_ratio": "1:1",
"grokMode": "normal"
}image-to-video
Animate one or more images into a 6-30 second video with audio using xAI Grok Imagine.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt (required, non-empty). |
| durationstringaffects pricedefault "6" | 6–30Integer seconds sent as a string, 6-30. |
| video_qualitystringaffects pricedefault "720p" | "480p", "720p"Output resolution: 480p or 720p. Always send video_quality. |
| grokModestringdefault "normal" | "fun", "normal", "spicy"Grok style mode: fun, normal or spicy. |
| aspect_ratiostring | "2:3", "3:2", "1:1", "9:16", "16:9"Output aspect ratio; use one of the listed values. When omitted, a single-image request follows the image and a multi-image request uses 16:9. |
| imagesarray | Optional array of all image URLs when more than one image is used (the first one also goes in image). |
| imagestring | First image URL. Required unless you send it as images[0]. |
Input files
image: First image URL.images: Optional array of up to 7 image URLs when more than one image is used (jpg/png/webp, max 10MB each).
Upload folders for /api/storage/upload
- image:
video-generator-grok-imagine
Rules
- image-to-video requires image or images (400).
- Send at most 7 images.
- aspect_ratio only applies when more than one image is used; with a single image the output follows the image.
- duration must be an integer string between 6 and 30 (400 otherwise); omitted duration defaults to 6.
- video_quality, when sent, must be 480p or 720p (400 otherwise).
- grokMode must be fun, normal or spicy; aspect_ratio must be one of the listed values.
- prompt must be non-empty.
Example request body
{
"mode": "grok-imagine",
"generation_type": "image-to-video",
"prompt": "She turns toward the camera and laughs as the wind blows her hair",
"image": "https://example.com/portrait.jpg",
"duration": "6",
"video_quality": "480p",
"aspect_ratio": "1:1",
"grokMode": "normal"
}Kling 2.6 mode: kling-2-6
text-to-video
Generates a cinematic video with optional synced native audio from a text prompt using Kling 2.6.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt (required, non-empty, max 20000 characters). |
| negative_promptstring | Optional things to avoid (max 2000 characters). |
| durationstringaffects pricedefault "5" | "5", "10"Seconds, sent as a string: "5" or "10". |
| aspect_ratiostringdefault "16:9" | "16:9", "9:16", "1:1"Output aspect ratio. |
| generate_audiobooleanaffects pricedefault true | Generate native audio. |
Rules
- duration must be '5' or '10'; aspect_ratio must be one of 16:9, 9:16, 1:1.
Example request body
{
"mode": "kling-2-6",
"generation_type": "text-to-video",
"prompt": "A red fox running through a snowy forest at sunrise, cinematic",
"duration": "5",
"aspect_ratio": "16:9",
"generate_audio": true
}image-to-video
Animates a start image into a video guided by a text prompt, with optional synced native audio, using Kling 2.6.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt describing the motion (required, non-empty). |
| imagestring · required | Public URL of the start frame image. |
| negative_promptstring | Optional things to avoid (max 2000 characters). |
| durationstringaffects pricedefault "5" | "5", "10"Seconds, sent as a string: "5" or "10". |
| aspect_ratiostringdefault "16:9" | "16:9"Send 16:9; the output follows the input image's aspect ratio. |
| generate_audiobooleanaffects pricedefault true | Generate native audio. |
Input files
image: 1 image URL used as the start frame (JPEG/PNG/WebP, max 10MB)
Upload folders for /api/storage/upload
- image:
video-generator-kling26
Rules
- image (or images[0]) is required for image-to-video, otherwise 400 'Image is required for image-to-video generation'.
- Only one image is used (start frame); there is no end-frame support for Kling 2.6.
- duration must be '5' or '10'.
Example request body
{
"mode": "kling-2-6",
"generation_type": "image-to-video",
"prompt": "The woman slowly turns her head and smiles",
"image": "https://cdn.veevid.ai/sample/motion-control-image-1.webp",
"duration": "5",
"aspect_ratio": "16:9",
"generate_audio": true
}Kling 3.0 mode: kling-3
text-to-video
Generates a 3-15 second video with optional native audio from a text prompt or a multi-shot prompt list using Kling 3.0 Standard/Pro.
| Parameter | Values & notes |
|---|---|
| promptstring | Single text prompt. Required unless multi_prompt is used; send either prompt or multi_prompt, not both (send prompt as an empty string when using multi_prompt). |
| multi_promptarray | Multi-shot prompts: array of {prompt: string (max 5000 characters), duration: string seconds 3-15}. Set duration to the sum of item durations. |
| model_versionstringaffects pricedefault "kling-3-standard" | "kling-3-standard", "kling-3-pro"Standard or Pro quality tier. |
| durationstringaffects pricedefault "5" | 3–15Total seconds as a string, integer step 1. |
| aspect_ratiostringdefault "16:9" | "16:9", "9:16", "1:1"Output aspect ratio. |
| generate_audiobooleanaffects pricedefault false | Native audio. When omitted the video is generated without audio (and priced as silent); the website sends true. |
| shot_typestringdefault "customize" | "customize", "intelligent"'customize' or 'intelligent'. |
| cfg_scalenumberdefault 0.5 | 0–1Prompt adherence, 0-1 (step 0.01). |
| negative_promptstring | Optional things to avoid (max 2000 characters). |
Rules
- duration must be an integer 3-15.
- aspect_ratio must be one of 16:9, 9:16, 1:1 when provided.
- model_version must be kling-3-standard or kling-3-pro, otherwise 400.
- prompt and multi_prompt must not both be non-empty; with multi_prompt, set
durationto the sum of item durations.
Example request body
{
"mode": "kling-3",
"generation_type": "text-to-video",
"prompt": "A drone shot flying over a neon-lit city at night, rain reflections",
"model_version": "kling-3-standard",
"duration": "5",
"aspect_ratio": "16:9",
"generate_audio": true,
"shot_type": "customize",
"cfg_scale": 0.5
}image-to-video
Animates a start frame (optionally to an end frame) into a 3-15 second video with optional native audio (Kling 3.0 Standard or Pro).
| Parameter | Values & notes |
|---|---|
| promptstring | Text prompt. Required unless multi_prompt is used; never send both non-empty. |
| imagestring | Public URL of the start frame. If images is also sent, images[0] takes precedence. |
| imagesarray | [start_frame_url] or [start_frame_url, end_frame_url]. images[0] is the start frame (takes precedence over image), images[1] the optional end frame. |
| multi_promptarray | Array of {prompt, duration} for multi-shot generation. Do not combine with an end frame. |
| model_versionstringaffects pricedefault "kling-3-standard" | "kling-3-standard", "kling-3-pro" |
| durationstringaffects pricedefault "5" | 3–15Total seconds as a string. |
| aspect_ratiostringdefault "16:9" | "16:9", "9:16", "1:1" |
| generate_audiobooleanaffects pricedefault false | Native audio. When omitted the video is generated without audio (and priced as silent); the website sends true. |
| shot_typestringdefault "customize" | "customize"Send 'customize'. |
| cfg_scalenumberdefault 0.5 | 0–1 |
| negative_promptstring | Optional negative prompt (max 2000 chars); omit when empty. |
Input files
image: 1 image URL (start frame); JPEG/PNG/WebP, max 10MBimages: 1-2 image URLs: [start_frame] or [start_frame, end_frame]; only the first two are used
Upload folders for /api/storage/upload
- image:
video-generator-kling3 - images:
video-generator-kling3
Rules
- image or images[0] is required, otherwise 400 'Image is required for image-to-video generation'.
- duration must be an integer 3-15; aspect_ratio one of 16:9, 9:16, 1:1; model_version kling-3-standard or kling-3-pro.
- images[1] is used as the end frame; an end frame cannot be combined with multi_prompt.
Example request body
{
"mode": "kling-3",
"generation_type": "image-to-video",
"prompt": "The character walks toward the camera as the wind blows her hair",
"image": "https://cdn.veevid.ai/sample/motion-control-image-1.webp",
"images": [
"https://cdn.veevid.ai/sample/motion-control-image-1.webp"
],
"model_version": "kling-3-pro",
"duration": "5",
"aspect_ratio": "16:9",
"generate_audio": true,
"shot_type": "customize",
"cfg_scale": 0.5
}LTX 2.5 mode: ltx-2-5
text-to-video
Generates a 720p-4K video with native audio from a text prompt using LTX 2.5 Fast or Pro; with a driving audio file it animates a scene to that audio (audio-to-video).
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt (required, non-empty). |
| model_versionstringaffects pricedefault "ltx-2.5-fast" | "ltx-2.5-fast", "ltx-2.5-pro"Fast or Pro tier. Fast: 720p-2160p, 6-20s. Pro: 720p/1080p, 6/8/10s. |
| durationstringaffects pricedefault "6" | "6", "8", "10", "12", "14", "16", "18", "20"Seconds, sent as a string. Pro supports only 6/8/10; 12-20 require ltx-2.5-fast. Ignored when audio is sent (the clip is as long as the audio). |
| video_qualitystringaffects pricedefault "1080p" | "720p", "1080p", "1440p", "2160p"Output resolution. Pro supports only 720p/1080p. Fast longer than 10s supports only 720p/1080p. Ignored when audio is sent (always 1080p). |
| aspect_ratiostringdefault "16:9" | "16:9", "9:16"16:9 or 9:16 ('auto' is also accepted when audio is sent). 16:9 when omitted. |
| target_fpsnumberdefault 25 | 24, 25, 48, 50Frames per second. Pro: 24, 25 or 50. Fast: 24, 25, 48 or 50; Fast longer than 10s only 24 or 25. Ignored when audio is sent. |
| generate_audiobooleandefault true | Generate a native audio track. Same price either way. Ignored when audio is sent. |
| camera_motionstring | "dolly_in", "dolly_out", "dolly_left", "dolly_right", "jib_up", "jib_down", "static", "focus_shift"Optional camera move. Omit to let the prompt decide. Not allowed together with audio (400). |
| audiostringaffects price | Optional driving audio (2-20s; Pro 2-10s). Switches to audio-to-video: the clip follows the audio and is as long as it, always at 1080p. The price then depends on the audio length. Must be a file uploaded with POST /api/storage/upload to folder 'ltx25/audio'; any other URL is rejected with 400 before charging. |
| input_audio_durationnumberaffects price | 2–20Length of the driving audio in seconds. The server measures the uploaded file and uses the measured length (your value is used only when it is within 0.15s of the measurement). Needed only to quote before uploading the audio. |
| guidance_scalenumber | 1–50Audio-to-video only (400 without audio). Omit to use the model default. |
Input files
audio: Optional: 1 driving audio file (2-20s, Pro 2-10s; max 15MB), uploaded to folder 'ltx25/audio'
Upload folders for /api/storage/upload
- audio:
ltx25/audio
Rules
- text-to-video does not accept images (400); use image-to-video.
- duration must be one of '6','8','10' for ltx-2.5-pro and '6'-'20' (even) for ltx-2.5-fast (400 otherwise).
- video_quality must be 720p or 1080p for ltx-2.5-pro; 720p, 1080p, 1440p or 2160p for ltx-2.5-fast (400 otherwise).
- ltx-2.5-fast longer than 10s supports only 720p/1080p and 24/25 fps (400 otherwise).
- Invalid combinations are rejected with 400 before charging; nothing is silently downgraded.
audiomust be uploaded first withPOST /api/storage/upload(folder 'ltx25/audio'); external URLs and files in other folders are rejected with 400 before charging.- If the server cannot read the length of the uploaded audio file, the request is rejected with 400; re-encode it as MP3 or WAV and upload again.
- Audio must be 2-20s long (Pro: 2-10s), measured on the server (400 otherwise).
- With
audio: duration, video_quality, target_fps and generate_audio are ignored; camera_motion and an end image are rejected (400). - To quote audio-to-video, send the uploaded
audioURL (the quote measures the file). Before uploading, sendhas_audio_input: trueandinput_audio_durationinstead. - Generation is async; poll by generation_id.
Example request body
{
"mode": "ltx-2-5",
"generation_type": "text-to-video",
"prompt": "A slow cinematic dolly shot through a neon-lit rainy street at night",
"model_version": "ltx-2.5-fast",
"duration": "6",
"video_quality": "720p",
"aspect_ratio": "16:9",
"target_fps": 25,
"generate_audio": true
}image-to-video
Animates a start image (optionally toward an end image) into a 720p-4K video with native audio using LTX 2.5 Fast or Pro; with a driving audio file it animates the image to that audio (audio-to-video).
| Parameter | Values & notes |
|---|---|
| promptstring | Text prompt describing the motion. Required unless audio is sent. |
| imagestring | Public URL of the start (first-frame) image. Required unless you send it as images[0]; when both are sent they are merged (image first, duplicates removed). |
| imagesstring[] | [start_url] or [start_url, end_url]. The second image is used as the end frame (not allowed with audio). More than 2 images is rejected (400). |
| model_versionstringaffects pricedefault "ltx-2.5-fast" | "ltx-2.5-fast", "ltx-2.5-pro"Fast or Pro tier. Fast: 720p-2160p, 6-20s. Pro: 720p/1080p, 6/8/10s. |
| durationstringaffects pricedefault "6" | "6", "8", "10", "12", "14", "16", "18", "20"Seconds, sent as a string. Pro supports only 6/8/10; 12-20 require ltx-2.5-fast. Ignored when audio is sent (the clip is as long as the audio). |
| video_qualitystringaffects pricedefault "1080p" | "720p", "1080p", "1440p", "2160p"Output resolution. Pro supports only 720p/1080p. Fast longer than 10s supports only 720p/1080p. Ignored when audio is sent (always 1080p). |
| aspect_ratiostringdefault "16:9" | "auto", "16:9", "9:16"'auto' follows the start image; otherwise 16:9 or 9:16. 16:9 when omitted (send 'auto' to keep the image's framing). |
| target_fpsnumberdefault 25 | 24, 25, 48, 50Frames per second. Pro: 24, 25 or 50. Fast: 24, 25, 48 or 50; Fast longer than 10s only 24 or 25. Ignored when audio is sent. |
| generate_audiobooleandefault true | Generate a native audio track. Same price either way. Ignored when audio is sent. |
| camera_motionstring | "dolly_in", "dolly_out", "dolly_left", "dolly_right", "jib_up", "jib_down", "static", "focus_shift"Optional camera move. Omit to let the prompt decide. Not allowed together with audio (400). |
| audiostringaffects price | Optional driving audio (2-20s; Pro 2-10s). Switches to audio-to-video: the clip follows the audio and is as long as it, always at 1080p. The price then depends on the audio length. Must be a file uploaded with POST /api/storage/upload to folder 'ltx25/audio'; any other URL is rejected with 400 before charging. |
| input_audio_durationnumberaffects price | 2–20Length of the driving audio in seconds. The server measures the uploaded file and uses the measured length (your value is used only when it is within 0.15s of the measurement). Needed only to quote before uploading the audio. |
| guidance_scalenumber | 1–50Audio-to-video only (400 without audio). Omit to use the model default. |
Input files
image: 1 image URL used as the first frameimages: Optional: [start, end] - 2 image URLs; the second is used as the last frameaudio: Optional: 1 driving audio file (2-20s, Pro 2-10s; max 15MB), uploaded to folder 'ltx25/audio'
Upload folders for /api/storage/upload
- image:
video-generator-ltx25 - end image:
video-generator-ltx25-end - audio:
ltx25/audio
Rules
- image-to-video requires a start image in image or images (400 'LTX 2.5 image-to-video requires a start image').
- Start frame = first image; end frame = second image if present. End image does not change the price.
- duration must be one of '6','8','10' for ltx-2.5-pro and '6'-'20' (even) for ltx-2.5-fast (400 otherwise).
- video_quality must be 720p or 1080p for ltx-2.5-pro; 720p, 1080p, 1440p or 2160p for ltx-2.5-fast (400 otherwise).
- ltx-2.5-fast longer than 10s supports only 720p/1080p and 24/25 fps (400 otherwise).
- Invalid combinations are rejected with 400 before charging; nothing is silently downgraded.
audiomust be uploaded first withPOST /api/storage/upload(folder 'ltx25/audio'); external URLs and files in other folders are rejected with 400 before charging.- If the server cannot read the length of the uploaded audio file, the request is rejected with 400; re-encode it as MP3 or WAV and upload again.
- Audio must be 2-20s long (Pro: 2-10s), measured on the server (400 otherwise).
- With
audio: duration, video_quality, target_fps and generate_audio are ignored; camera_motion and an end image are rejected (400). - To quote audio-to-video, send the uploaded
audioURL (the quote measures the file). Before uploading, sendhas_audio_input: trueandinput_audio_durationinstead. - Generation is async; poll by generation_id.
Example request body
{
"mode": "ltx-2-5",
"generation_type": "image-to-video",
"prompt": "The woman turns toward the camera and smiles as the wind moves her hair",
"image": "https://example.com/start.jpg",
"images": [
"https://example.com/start.jpg"
],
"model_version": "ltx-2.5-fast",
"duration": "6",
"video_quality": "1080p",
"aspect_ratio": "auto",
"target_fps": 25,
"generate_audio": true
}LTX 2.3 mode: ltx-2-3
text-to-video
Generates a high-resolution (1080p-2160p) video with optional native audio from a text prompt using LTX 2.3 or LTX 2.3 Fast.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt (required, non-empty, max 20000 chars). |
| model_versionstringaffects pricedefault "ltx-2.3" | "ltx-2.3", "ltx-2.3-fast"Standard or Fast variant. Use one of the listed values. |
| durationstring · requiredaffects price | "6", "8", "10", "12", "14", "16", "18", "20"Seconds, sent as a string. Standard ('ltx-2.3') only supports 6/8/10; 12-20 require 'ltx-2.3-fast'. Always send it explicitly. |
| video_qualitystringaffects pricedefault "1080p" | "1080p", "1440p", "2160p"Output resolution. Forced to 1080p when model_version is 'ltx-2.3-fast' and duration > 10. |
| aspect_ratiostringdefault "16:9" | "16:9", "9:16"16:9 or 9:16 for text-to-video. |
| target_fpsnumberdefault 24 | 24, 25, 48, 50Frames per second: 24, 25, 48 or 50. Send it explicitly (24 is used when omitted). Forced to 25 when model_version is 'ltx-2.3-fast' and duration > 10. |
| generate_audiobooleandefault true | Generate native audio track. |
Rules
- Always send
duration; it must be one of '6','8','10','12','14','16','18','20' (400 otherwise). - model_version 'ltx-2.3' supports only duration 6/8/10; 'ltx-2.3-fast' supports 6-20 (even). Other combinations make the generation fail.
- video_quality must be one of 1080p, 1440p, 2160p (400 otherwise); defaults to 1080p.
- target_fps must be 24, 25, 48 or 50; other values make the generation fail.
- With ltx-2.3-fast and duration > 10, output is always 1080p at 25 fps and priced at 1080p.
- aspect_ratio: use one of the listed values.
- Generation is async; poll by generation_id.
Example request body
{
"mode": "ltx-2-3",
"generation_type": "text-to-video",
"prompt": "A slow cinematic dolly shot through a neon-lit rainy street at night",
"model_version": "ltx-2.3",
"duration": "6",
"video_quality": "1080p",
"aspect_ratio": "16:9",
"target_fps": 25,
"generate_audio": true
}image-to-video
Animates a start image (optionally toward an end image) into a high-resolution video with optional native audio using LTX 2.3 or LTX 2.3 Fast.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt describing the motion (required, non-empty, max 20000 chars). |
| imagestring | Public URL of the start (first-frame) image. Required unless you send images; when both are sent, images[0] is used. |
| imagesstring[] | [start_url] or [start_url, end_url]. images[1], if present, is used as the end frame. |
| model_versionstringaffects pricedefault "ltx-2.3" | "ltx-2.3", "ltx-2.3-fast"Standard or Fast variant. Use one of the listed values. |
| durationstring · requiredaffects price | "6", "8", "10", "12", "14", "16", "18", "20"Seconds, sent as a string. Standard only supports 6/8/10; 12-20 require 'ltx-2.3-fast'. Always send it explicitly. |
| video_qualitystringaffects pricedefault "1080p" | "1080p", "1440p", "2160p"Output resolution. Forced to 1080p when Fast and duration > 10. |
| aspect_ratiostringdefault "16:9" | "auto", "16:9", "9:16"'auto' follows the input image; otherwise 16:9 or 9:16. |
| target_fpsnumberdefault 24 | 24, 25, 48, 50Frames per second: 24, 25, 48 or 50. Send it explicitly (24 is used when omitted). Forced to 25 when Fast and duration > 10. |
| generate_audiobooleandefault true | Generate native audio track. |
Input files
image: 1 image URL used as the first frame (JPEG/PNG/GIF, max 10MB)images: Optional: [start, end] - 2 image URLs; the second is used as the last frame
Upload folders for /api/storage/upload
- image:
video-generator-ltx23 - end image:
video-generator-ltx23-end
Rules
- image-to-video requires image or a non-empty images array (400 'Image is required for image-to-video generation').
- Start frame = image (or images[0] if image is absent); end frame = images[1] if present.
- Always send
duration; it must be one of '6','8','10','12','14','16','18','20' (400 otherwise). - model_version 'ltx-2.3' supports only duration 6/8/10; 'ltx-2.3-fast' supports 6-20 (even). Other combinations make the generation fail.
- video_quality must be one of 1080p, 1440p, 2160p (400 otherwise); defaults to 1080p.
- target_fps must be 24, 25, 48 or 50; other values make the generation fail.
- With ltx-2.3-fast and duration > 10, output is always 1080p at 25 fps and priced at 1080p.
- aspect_ratio: use one of the listed values.
- End image does not change the price.
Example request body
{
"mode": "ltx-2-3",
"generation_type": "image-to-video",
"prompt": "The woman turns toward the camera and smiles as the wind moves her hair",
"image": "https://example.com/start.jpg",
"images": [
"https://example.com/start.jpg"
],
"model_version": "ltx-2.3",
"duration": "6",
"video_quality": "1080p",
"aspect_ratio": "16:9",
"target_fps": 25,
"generate_audio": true
}Wan Animate mode: wan-animate
image-to-video
Animates a character image with the motion of a reference video (animate), or replaces the person in the reference video with the character (replace).
| Parameter | Values & notes |
|---|---|
| imagestring · required | Public URL of the character image. |
| videostring · required | Public URL of the reference (driving) video, up to 120 seconds. |
| animationModestringdefault "animate" | "animate", "replace"'animate' = character performs the reference video's motion; 'replace' = character replaces the person in the reference video. |
| video_qualitystring · requiredaffects price | "480p", "720p"Output resolution. Always send it ('480p' or '720p'). |
| durationstringaffects pricedefault "5" | 1–120Reference video length in whole seconds (rounded up), as a string. Send the real length of the reference video; it determines the price (5 when omitted). |
| promptstring | Optional text guidance, max 5000 characters; may be empty. |
| aspect_ratiostring | Optional; can be omitted (output follows the inputs). |
Input files
image: 1 character image URL (JPEG/PNG/GIF, max 10MB)video: 1 reference video URL, max 120 seconds (max 200MB)
Upload folders for /api/storage/upload
- image:
video-generator - video:
wan-animate/reference-videos
Rules
- image is required (400 'Character image is required for wan-animate generation').
- video is required (400 'Reference video is required for wan-animate generation').
- When sent, duration must be an integer > 0 and <= 120 (400).
- prompt, if sent, must be <= 5000 characters (400).
- generation_type must be 'image-to-video'.
- Always send
video_quality('480p' or '720p'). - animationMode must be 'animate' or 'replace'; defaults to 'animate'.
- Send
durationequal to the reference video length in whole seconds (rounded up). - Generation is async; poll by generation_id.
Example request body
{
"mode": "wan-animate",
"generation_type": "image-to-video",
"prompt": "",
"image": "https://example.com/character.png",
"video": "https://example.com/dance.mp4",
"duration": "8",
"video_quality": "480p",
"aspect_ratio": "",
"animationMode": "animate"
}Seedance 1.5 Pro mode: seedance-1.5-pro
text-to-video
Seedance 1.5 Pro text-to-video generation, 4-12s up to 1080p with optional native audio.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt (max 20000 chars). Required; an empty prompt makes the generation fail. |
| durationstringaffects pricedefault "5" | 4–12Output length in seconds as an integer string, 4-12; other values return 400. |
| video_qualitystringaffects pricedefault "720p" | "480p", "720p", "1080p"Output resolution; defaults to 720p. Use one of the listed values. |
| aspect_ratiostringdefault "16:9" | "auto", "21:9", "16:9", "4:3", "1:1", "3:4", "9:16"Output aspect ratio. 'auto' lets the model decide (for image-to-video it follows the input image). Send it explicitly; when omitted, 16:9 is used. |
| generate_audiobooleanaffects pricedefault true | Generate synchronized audio. Silent output costs half. |
| camera_fixedbooleandefault false | Lock the camera (no camera movement). Does not change price. |
| seedintegerdefault -1 | Random seed; -1 means random. |
Rules
- duration must be an integer string between 4 and 12; otherwise 400.
- camera_fixed and seed do not affect price.
- Only text-to-video and image-to-video are supported.
- prompt must be non-empty.
Example request body
{
"mode": "seedance-1.5-pro",
"generation_type": "text-to-video",
"prompt": "A red fox running through fresh snow at sunrise, cinematic tracking shot",
"duration": "5",
"aspect_ratio": "16:9",
"video_quality": "720p",
"generate_audio": true,
"camera_fixed": false,
"seed": -1
}image-to-video
Seedance 1.5 Pro image-to-video (first frame, optional last frame) generation, 4-12s up to 1080p with optional native audio.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt (max 20000 chars). Required; an empty prompt makes the generation fail. |
| imagestring | Public https URL of the first-frame image. Required unless you send it as images[0]. |
| imagesarray<string> | Optional [first_frame_url, last_frame_url]; when 2+ entries are sent, the last one is used as the last frame. Send image as well. |
| durationstringaffects pricedefault "5" | 4–12Output length in seconds as an integer string, 4-12; other values return 400. |
| video_qualitystringaffects pricedefault "720p" | "480p", "720p", "1080p"Output resolution; defaults to 720p. Use one of the listed values. |
| aspect_ratiostringdefault "16:9" | "auto", "21:9", "16:9", "4:3", "1:1", "3:4", "9:16"Output aspect ratio. 'auto' lets the model decide (for image-to-video it follows the input image). Send it explicitly; when omitted, 16:9 is used. |
| generate_audiobooleanaffects pricedefault true | Generate synchronized audio. Silent output costs half. |
| camera_fixedbooleandefault false | Lock the camera (no camera movement). Does not change price. |
| seedintegerdefault -1 | Random seed; -1 means random. |
Input files
image: 1 image URL used as the first frame (JPEG/PNG/WebP, max 10MB)images: optional 2nd image URL as last frame: send images=[first,last]
Upload folders for /api/storage/upload
- image:
seedance15pro
Rules
- duration must be an integer string between 4 and 12; otherwise 400.
- camera_fixed and seed do not affect price.
- Only text-to-video and image-to-video are supported.
- prompt must be non-empty.
Example request body
{
"mode": "seedance-1.5-pro",
"generation_type": "image-to-video",
"prompt": "A red fox running through fresh snow at sunrise, cinematic tracking shot",
"duration": "5",
"aspect_ratio": "auto",
"video_quality": "720p",
"generate_audio": true,
"camera_fixed": false,
"seed": -1,
"image": "https://example.com/first-frame.jpg",
"images": [
"https://example.com/first-frame.jpg"
]
}Veevid 1.0 Pro mode: veevid-1.0-pro
text-to-video
Veevid 1.0 Pro text-to-video generation, 4-12s up to 1080p with optional native audio.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt (max 20000 chars). Required; an empty prompt makes the generation fail. |
| durationstringaffects pricedefault "5" | 4–12Output length in seconds as an integer string, 4-12; other values return 400. |
| video_qualitystringaffects pricedefault "720p" | "480p", "720p", "1080p"Output resolution; defaults to 720p. Use one of the listed values. |
| aspect_ratiostringdefault "16:9" | "auto", "21:9", "16:9", "4:3", "1:1", "3:4", "9:16"Output aspect ratio. 'auto' lets the model decide (for image-to-video it follows the input image). Send it explicitly; when omitted, 16:9 is used. |
| generate_audiobooleanaffects pricedefault true | Generate synchronized audio. Silent output costs half. |
| camera_fixedbooleandefault false | Lock the camera (no camera movement). Does not change price. |
| seedintegerdefault -1 | Random seed; -1 means random. |
Rules
- duration must be an integer string between 4 and 12; otherwise 400.
- camera_fixed and seed do not affect price.
- Only text-to-video and image-to-video are supported.
- prompt must be non-empty.
Example request body
{
"mode": "veevid-1.0-pro",
"generation_type": "text-to-video",
"prompt": "A red fox running through fresh snow at sunrise, cinematic tracking shot",
"duration": "5",
"aspect_ratio": "16:9",
"video_quality": "720p",
"generate_audio": true,
"camera_fixed": false,
"seed": -1
}image-to-video
Veevid 1.0 Pro image-to-video (first frame, optional last frame) generation, 4-12s up to 1080p with optional native audio.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt (max 20000 chars). Required; an empty prompt makes the generation fail. |
| imagestring | Public https URL of the first-frame image. Required unless you send it as images[0]. |
| imagesarray<string> | Optional [first_frame_url, last_frame_url]; when 2+ entries are sent, the last one is used as the last frame. Send image as well. |
| durationstringaffects pricedefault "5" | 4–12Output length in seconds as an integer string, 4-12; other values return 400. |
| video_qualitystringaffects pricedefault "720p" | "480p", "720p", "1080p"Output resolution; defaults to 720p. Use one of the listed values. |
| aspect_ratiostringdefault "16:9" | "auto", "21:9", "16:9", "4:3", "1:1", "3:4", "9:16"Output aspect ratio. 'auto' lets the model decide (for image-to-video it follows the input image). Send it explicitly; when omitted, 16:9 is used. |
| generate_audiobooleanaffects pricedefault true | Generate synchronized audio. Silent output costs half. |
| camera_fixedbooleandefault false | Lock the camera (no camera movement). Does not change price. |
| seedintegerdefault -1 | Random seed; -1 means random. |
Input files
image: 1 image URL used as the first frame (JPEG/PNG/WebP, max 10MB)images: optional 2nd image URL as last frame: send images=[first,last]
Upload folders for /api/storage/upload
- image:
veevid-1-0-pro
Rules
- duration must be an integer string between 4 and 12; otherwise 400.
- camera_fixed and seed do not affect price.
- Only text-to-video and image-to-video are supported.
- prompt must be non-empty.
Example request body
{
"mode": "veevid-1.0-pro",
"generation_type": "image-to-video",
"prompt": "A red fox running through fresh snow at sunrise, cinematic tracking shot",
"duration": "5",
"aspect_ratio": "auto",
"video_quality": "720p",
"generate_audio": true,
"camera_fixed": false,
"seed": -1,
"image": "https://example.com/first-frame.jpg",
"images": [
"https://example.com/first-frame.jpg"
]
}Veevid 2.0 Pro mode: veevid-2.0-pro
text-to-video
Veevid 2.0 Pro text-to-video, 4-15s, Mini/Fast/Standard tiers up to 4K with native audio.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt (max 20000 chars). Required for text-to-video. |
| model_versionstringaffects pricedefault "seedance-2.0" | "seedance-2.0-mini", "seedance-2.0-fast", "seedance-2.0"Quality tier: Mini, Fast or Standard ('seedance-2.0'). Always send model_version; when omitted, Standard is used. Mini/Fast support only 480p/720p. |
| durationstringaffects pricedefault "5" | 4–15Output length in seconds as an integer string; values outside 4-15 are clamped to 4-15. |
| video_qualitystringaffects pricedefault "480p" | "480p", "720p", "1080p", "4k"Output resolution. 1080p/4k are available only with Standard ('seedance-2.0'); with Mini/Fast they are reduced to 720p and priced as 720p. Defaults to 480p; use one of the listed values. |
| aspect_ratiostringdefault "16:9" | "auto", "21:9", "16:9", "4:3", "1:1", "3:4", "9:16"Output aspect ratio; 'auto' adapts to the content. Use one of the listed values; when omitted, 16:9 is used. |
| generate_audiobooleandefault true | Generate native audio. Does not change price. |
| seedintegerdefault -1 | Random seed; -1 means random. |
Rules
- duration is clamped to 4-15 seconds (for both price and output); no range error is returned.
- Always send
model_version; values other than the Mini/Fast tiers are treated as Standard ('seedance-2.0'). - generate_audio and seed do not affect price.
Example request body
{
"mode": "veevid-2.0-pro",
"generation_type": "text-to-video",
"prompt": "A chef flips a pancake in a sunlit kitchen, slow motion, warm tones",
"model_version": "seedance-2.0-fast",
"duration": "5",
"aspect_ratio": "16:9",
"video_quality": "480p",
"generate_audio": true,
"seed": -1
}image-to-video
Veevid 2.0 Pro image-to-video with first/optional last frame, 4-15s, Mini/Fast/Standard tiers up to 4K with native audio.
| Parameter | Values & notes |
|---|---|
| promptstring | Optional text prompt (max 20000 chars). |
| imagestring | Public https URL of the first-frame image. Required unless you send it as images[0]. |
| imagesarray<string> | Optional [first_frame_url, last_frame_url]; when 2+ entries are sent, the last one is used as the last frame. |
| model_versionstringaffects pricedefault "seedance-2.0" | "seedance-2.0-mini", "seedance-2.0-fast", "seedance-2.0"Quality tier: Mini, Fast or Standard ('seedance-2.0'). Always send model_version; when omitted, Standard is used. Mini/Fast support only 480p/720p. |
| durationstringaffects pricedefault "5" | 4–15Output length in seconds as an integer string; values outside 4-15 are clamped to 4-15. |
| video_qualitystringaffects pricedefault "480p" | "480p", "720p", "1080p", "4k"Output resolution. 1080p/4k are available only with Standard ('seedance-2.0'); with Mini/Fast they are reduced to 720p and priced as 720p. Defaults to 480p; use one of the listed values. |
| aspect_ratiostringdefault "16:9" | "auto", "21:9", "16:9", "4:3", "1:1", "3:4", "9:16"Output aspect ratio; 'auto' adapts to the content. Use one of the listed values; when omitted, 16:9 is used. |
| generate_audiobooleandefault true | Generate native audio. Does not change price. |
| seedintegerdefault -1 | Random seed; -1 means random. |
Input files
image: 1 image URL used as first frame (JPEG/PNG/WebP, max 30MB)images: optional [first,last] to add a last frame
Upload folders for /api/storage/upload
- image:
seedance20
Rules
- duration is clamped to 4-15 seconds (for both price and output); no range error is returned.
- Always send
model_version; values other than the Mini/Fast tiers are treated as Standard ('seedance-2.0'). - generate_audio and seed do not affect price.
Example request body
{
"mode": "veevid-2.0-pro",
"generation_type": "image-to-video",
"prompt": "A chef flips a pancake in a sunlit kitchen, slow motion, warm tones",
"model_version": "seedance-2.0-fast",
"duration": "5",
"aspect_ratio": "auto",
"video_quality": "480p",
"generate_audio": true,
"seed": -1,
"image": "https://example.com/first-frame.jpg",
"images": [
"https://example.com/first-frame.jpg"
]
}Veevid 2.5 Pro mode: veevid-2.5-pro
text-to-video
Veevid 2.5 Pro text-to-video, one-take videos 4-30s up to 1080p with native audio.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt (max 20000 chars). A prompt or at least one image is required. |
| durationstringaffects pricedefault "5" | 4–30Output length: integer string 4-30, or '-1' to let the model choose the length. Anything else returns 400. |
| video_qualitystringaffects pricedefault "480p" | "480p", "720p", "1080p"Output resolution (1080p is 10-bit HEVC). Defaults to 480p; use one of the listed values. |
| aspect_ratiostringdefault "16:9" | "auto", "21:9", "16:9", "4:3", "1:1", "3:4", "9:16"Output aspect ratio; 'auto' adapts to the content. Use one of the listed values; when omitted, 16:9 is used. |
| output_formatstringdefault "mp4" | "mp4", "mov"Container. mov = H.264 yuv444p + PCM audio (poor browser playback). Only mp4/mov are valid for this mode. |
| generate_audiobooleandefault true | Generate native audio. Does not change price. |
| seedintegerdefault -1 | Random seed; -1 means random. |
Rules
- duration must be a pure integer string 4-30 or '-1'; otherwise 400.
- output_format must be mp4 or mov for this mode; image formats (PNG/JPEG) return 400.
- output_format, generate_audio and seed do not affect price.
- 480p results are usually drafts: the status response then includes
seedance25_draft, and the draft can be upgraded once to a 1080p version within 7 days withPOST /api/video-generation/{id}/upgrade-draft.
Example request body
{
"mode": "veevid-2.5-pro",
"generation_type": "text-to-video",
"prompt": "A lighthouse keeper climbs the spiral stairs at dusk, one continuous take",
"duration": "5",
"aspect_ratio": "16:9",
"video_quality": "480p",
"output_format": "mp4",
"generate_audio": true,
"seed": -1
}image-to-video
Veevid 2.5 Pro image-to-video with first/optional last frame, one-take videos 4-30s up to 1080p with native audio.
| Parameter | Values & notes |
|---|---|
| promptstring | Text prompt (max 20000 chars). A prompt or at least one image is required. |
| imagestring | Public https URL of the first-frame image. Required unless you send it as images[0]. |
| imagesarray<string> | Optional [first_frame_url, last_frame_url]; images[0] is used as first frame and the last entry (when 2+) as last frame. |
| durationstringaffects pricedefault "5" | 4–30Output length: integer string 4-30, or '-1' to let the model choose the length. Anything else returns 400. |
| video_qualitystringaffects pricedefault "480p" | "480p", "720p", "1080p"Output resolution (1080p is 10-bit HEVC). Defaults to 480p; use one of the listed values. |
| aspect_ratiostringdefault "auto" | "auto", "21:9", "16:9", "4:3", "1:1", "3:4", "9:16"Not applied for image-to-video: the output follows the first frame's aspect ratio. Send 'auto'. |
| output_formatstringdefault "mp4" | "mp4", "mov"Container. mov = H.264 yuv444p + PCM audio (poor browser playback). Only mp4/mov are valid for this mode. |
| generate_audiobooleandefault true | Generate native audio. Does not change price. |
| seedintegerdefault -1 | Random seed; -1 means random. |
Input files
image: 1 image URL used as first frame (JPEG/PNG/WebP, max 30MB)images: optional [first,last] to add a last frame
Upload folders for /api/storage/upload
- image:
seedance25
Rules
- image (or images) is required for image-to-video, else 400.
- For image-to-video the aspect ratio always follows the first frame, regardless of aspect_ratio.
- duration must be a pure integer string 4-30 or '-1'; otherwise 400.
- output_format must be mp4 or mov for this mode; image formats (PNG/JPEG) return 400.
- output_format, generate_audio and seed do not affect price.
- 480p results are usually drafts: the status response then includes
seedance25_draft, and the draft can be upgraded once to a 1080p version within 7 days withPOST /api/video-generation/{id}/upgrade-draft.
Example request body
{
"mode": "veevid-2.5-pro",
"generation_type": "image-to-video",
"prompt": "A lighthouse keeper climbs the spiral stairs at dusk, one continuous take",
"duration": "5",
"aspect_ratio": "auto",
"video_quality": "480p",
"output_format": "mp4",
"generate_audio": true,
"seed": -1,
"image": "https://example.com/first-frame.jpg",
"images": [
"https://example.com/first-frame.jpg"
]
}Seedance 2.0 mode: seedance-2.0
text-to-video
Generates a 4-15s video with optional native audio from a text prompt using ByteDance Seedance 2.0 (Mini/Fast/Standard tiers).
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt (max 20000 chars). Required for text-to-video. |
| model_versionstringaffects pricedefault "seedance-2.0" | "seedance-2.0-mini", "seedance-2.0-fast", "seedance-2.0"Quality tier: 'seedance-2.0-mini', 'seedance-2.0-fast' or 'seedance-2.0' (Standard). Always send model_version; when omitted, Standard is used. Mini and Fast support only 480p/720p. |
| video_qualitystringaffects pricedefault "480p" | "480p", "720p", "1080p", "4k"Output resolution ('4K' is also accepted as 4k). Defaults to 480p; use one of the listed values. 1080p and 4k are only available with 'seedance-2.0'; with Mini/Fast they fall back to 720p. |
| durationstringaffects pricedefault "5" | 4–15Output length in seconds as an integer string (e.g. "5"), 4-15. Out-of-range values are clamped to 4-15; non-numeric values fall back to 5. |
| aspect_ratiostringdefault "16:9" | "auto", "16:9", "9:16", "1:1", "4:3", "3:4", "21:9"Output aspect ratio; 'auto' (or 'adaptive') adapts to the content. Use one of the listed values; when omitted, 16:9 is used. |
| generate_audiobooleandefault true | Generate native audio track. |
| seedintegerdefault -1 | Random seed; -1 = random. |
Rules
- duration must be a JSON string (e.g. "5"); a number returns 400.
- duration is clamped to 4-15 seconds (for both price and output) instead of being rejected.
- Always send
model_version; values other than seedance-2.0 / seedance-2.0-fast / seedance-2.0-mini are treated as seedance-2.0 (Standard). - seedance-2.0-fast and seedance-2.0-mini support only 480p/720p; requested 1080p/4k is reduced to 720p and priced as 720p.
- Credits are deducted at submission (402 if insufficient); unused credits are refunded automatically if the generation fails.
Example request body
{
"mode": "seedance-2.0",
"generation_type": "text-to-video",
"prompt": "A red fox running through fresh snow at sunrise, cinematic tracking shot",
"model_version": "seedance-2.0-fast",
"video_quality": "480p",
"duration": "5",
"aspect_ratio": "16:9",
"generate_audio": true,
"seed": -1
}image-to-video
Animates a first-frame image (optionally with a last-frame image) into a 4-15s video with Seedance 2.0.
| Parameter | Values & notes |
|---|---|
| promptstring | Optional text prompt (max 20000 chars) when image or reference inputs are given. |
| model_versionstringaffects pricedefault "seedance-2.0" | "seedance-2.0-mini", "seedance-2.0-fast", "seedance-2.0"Quality tier: 'seedance-2.0-mini', 'seedance-2.0-fast' or 'seedance-2.0' (Standard). Always send model_version; when omitted, Standard is used. Mini and Fast support only 480p/720p. |
| video_qualitystringaffects pricedefault "480p" | "480p", "720p", "1080p", "4k"Output resolution ('4K' is also accepted as 4k). Defaults to 480p; use one of the listed values. 1080p and 4k are only available with 'seedance-2.0'; with Mini/Fast they fall back to 720p. |
| durationstringaffects pricedefault "5" | 4–15Output length in seconds as an integer string (e.g. "5"), 4-15. Out-of-range values are clamped to 4-15; non-numeric values fall back to 5. |
| aspect_ratiostringdefault "16:9" | "auto", "16:9", "9:16", "1:1", "4:3", "3:4", "21:9"Output aspect ratio; 'auto' (or 'adaptive') adapts to the content. Use one of the listed values; when omitted, 16:9 is used. |
| generate_audiobooleandefault true | Generate native audio track. |
| seedintegerdefault -1 | Random seed; -1 = random. |
| imagestring | First-frame image URL. Required unless you send it as images[0]. |
| imagesarray<string> | [first] or [first, last]. images[0] is used as first frame (overrides image), and the last element becomes the last frame when length > 1. |
Input files
image: 1 public image URL used as first frame (JPEG/PNG/WebP, max 30MB)images: optional [first_frame_url, last_frame_url] for first+last-frame mode
Upload folders for /api/storage/upload
- image:
seedance20
Rules
- duration must be a JSON string (e.g. "5"); a number returns 400.
- duration is clamped to 4-15 seconds (for both price and output) instead of being rejected.
- Always send
model_version; values other than seedance-2.0 / seedance-2.0-fast / seedance-2.0-mini are treated as seedance-2.0 (Standard). - seedance-2.0-fast and seedance-2.0-mini support only 480p/720p; requested 1080p/4k is reduced to 720p and priced as 720p.
- Credits are deducted at submission (402 if insufficient); unused credits are refunded automatically if the generation fails.
- aspect_ratio is applied as sent for image-to-video.
- Price is the same as text-to-video (images do not add cost).
Example request body
{
"mode": "seedance-2.0",
"generation_type": "image-to-video",
"prompt": "The woman turns toward the camera and smiles",
"image": "https://example.com/first.jpg",
"images": [
"https://example.com/first.jpg"
],
"model_version": "seedance-2.0-fast",
"video_quality": "720p",
"duration": "5",
"aspect_ratio": "auto",
"generate_audio": true,
"seed": -1
}reference-to-video
Generates a 4-15s video guided by multimodal references (up to 9 images, 3 videos, 3 audio clips) with Seedance 2.0 omni reference.
| Parameter | Values & notes |
|---|---|
| promptstring | Optional text prompt (max 20000 chars) when image or reference inputs are given. |
| model_versionstringaffects pricedefault "seedance-2.0" | "seedance-2.0-mini", "seedance-2.0-fast", "seedance-2.0"Quality tier: 'seedance-2.0-mini', 'seedance-2.0-fast' or 'seedance-2.0' (Standard). Always send model_version; when omitted, Standard is used. Mini and Fast support only 480p/720p. |
| video_qualitystringaffects pricedefault "480p" | "480p", "720p", "1080p", "4k"Output resolution ('4K' is also accepted as 4k). Defaults to 480p; use one of the listed values. 1080p and 4k are only available with 'seedance-2.0'; with Mini/Fast they fall back to 720p. |
| durationstringaffects pricedefault "5" | 4–15Output length in seconds as an integer string (e.g. "5"), 4-15. Out-of-range values are clamped to 4-15; non-numeric values fall back to 5. |
| aspect_ratiostringdefault "16:9" | "auto", "16:9", "9:16", "1:1", "4:3", "3:4", "21:9"Output aspect ratio; 'auto' (or 'adaptive') adapts to the content. Use one of the listed values; when omitted, 16:9 is used. |
| generate_audiobooleandefault true | Generate native audio track. |
| seedintegerdefault -1 | Random seed; -1 = random. |
| imagesarray<string> | Reference image URLs (role reference_image). Max 9; extras are ignored. |
| videosarray<string>affects price | Reference video URLs (role reference_video), mp4/mov. Max 3; extras are ignored. |
| input_video_durationnumberaffects price | 0–15Total seconds of all reference videos (rounded up). Required when videos are sent; must equal the total reference video length. >15 is rejected with 400; values below 2 are priced as 2. |
| audiostring | Optional; if sent, set it to the first reference audio URL (audios[0]). |
| audiosarray<string> | Reference audio URLs (role reference_audio). Max 3, .wav/.mp3, total <= 15s; extras are ignored. Audio cannot be the only reference: at least one image or video is required. |
Input files
images: 0-9 reference image URLs (JPEG/PNG/WebP; max 30MB each)videos: 0-3 reference video URLs (mp4/mov, max 50MB each, total <= 15s)audios: 0-3 reference audio URLs (wav/mp3, max 15MB each, total <= 15s); requires at least one image or video
Upload folders for /api/storage/upload
- image:
seedance20 - video:
seedance20/reference-videos - audio:
seedance20/audio
Rules
- duration must be a JSON string (e.g. "5"); a number returns 400.
- duration is clamped to 4-15 seconds (for both price and output) instead of being rejected.
- Always send
model_version; values other than seedance-2.0 / seedance-2.0-fast / seedance-2.0-mini are treated as seedance-2.0 (Standard). - seedance-2.0-fast and seedance-2.0-mini support only 480p/720p; requested 1080p/4k is reduced to 720p and priced as 720p.
- Credits are deducted at submission (402 if insufficient); unused credits are refunded automatically if the generation fails.
- At least one reference image or video is required (audio-only is not allowed); otherwise the generation fails.
- When sending videos, send
input_video_durationequal to the total reference video length in seconds (rounded up); > 15 is rejected with 400. - Send at most 9 images, 3 videos and 3 audios; extras are ignored.
- Without reference videos (images/audio only), price equals the text-to-video price.
- When quoting with
/api/quote, addhas_video_input: trueandinput_video_durationif you send reference videos (the quote endpoint does not read video URLs, so the price would come out too low).
Example request body
{
"mode": "seedance-2.0",
"generation_type": "reference-to-video",
"prompt": "The character from image 1 dances in the style of video 1",
"images": [
"https://example.com/character.png"
],
"videos": [
"https://example.com/dance.mp4"
],
"input_video_duration": 6,
"model_version": "seedance-2.0-fast",
"video_quality": "480p",
"duration": "5",
"aspect_ratio": "auto",
"generate_audio": true,
"seed": -1
}Seedance 2.5 mode: seedance-2.5
text-to-video
Generates a 4-30s one-take video with optional native audio from a text prompt using ByteDance Seedance 2.5.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt (max 20000 chars). Required for text-to-video. |
| video_qualitystringaffects pricedefault "480p" | "480p", "720p", "1080p"Output resolution; always send video_quality explicitly. 1080p is 10-bit HEVC (may not play in some browsers). '4k'/'4K'/'2160p' are treated as 1080p. |
| durationstringaffects pricedefault "5" | 4–30Output seconds as a string containing a plain integer from 4 to 30, or "-1" to let the model choose the length. Anything else (e.g. "3", "31", "5.0", "0x10") is rejected with 400. |
| aspect_ratiostringdefault "16:9" | "auto", "16:9", "9:16", "1:1", "4:3", "3:4", "21:9"Output aspect ratio; 'auto' lets the model choose. Send aspect_ratio explicitly and use one of the listed values. |
| output_formatstringdefault "mp4" | "mp4", "mov"Output container. mov = H.264 yuv444p with PCM audio (often not playable in browsers). Any other value is rejected with 400. |
| generate_audiobooleandefault true | Generate a native audio track. Default true. |
| seedintegerdefault -1 | Random seed; -1 = random. |
Rules
- duration must be a JSON string matching /^-?\d+$/ and equal -1 or 4-30, otherwise 400
- output_format must be mp4 or mov (400 otherwise)
- Always send
video_qualityas 480p, 720p or 1080p ('4k' is treated as 1080p) - No model_version parameter (single tier)
- Price can be previewed with POST /api/quote; /api/generate-video returns 402 if credits are insufficient. Credits for failed generations are refunded automatically.
- 480p results are usually drafts: the status response then includes
seedance25_draft, and the draft can be upgraded once to a 1080p version within 7 days withPOST /api/video-generation/{id}/upgrade-draft.
Example request body
{
"mode": "seedance-2.5",
"generation_type": "text-to-video",
"prompt": "A lighthouse keeper climbs the spiral stairs at dusk, one continuous shot",
"video_quality": "480p",
"duration": "5",
"aspect_ratio": "16:9",
"output_format": "mp4",
"generate_audio": true,
"seed": -1
}image-to-video
Animates a first-frame image (optionally to a last-frame image) into a 4-30s video with Seedance 2.5; aspect ratio follows the input image.
| Parameter | Values & notes |
|---|---|
| promptstring | Text prompt (max 20000 chars). Optional when a first-frame image is given. |
| video_qualitystringaffects pricedefault "480p" | "480p", "720p", "1080p"Output resolution; always send video_quality explicitly. 1080p is 10-bit HEVC (may not play in some browsers). '4k'/'4K'/'2160p' are treated as 1080p. |
| durationstringaffects pricedefault "5" | 4–30Output seconds as a string containing a plain integer from 4 to 30, or "-1" to let the model choose the length. Anything else (e.g. "3", "31", "5.0", "0x10") is rejected with 400. |
| aspect_ratiostringdefault "auto" | "auto", "16:9", "9:16", "1:1", "4:3", "3:4", "21:9"Ignored for image-to-video: the output aspect ratio follows the first-frame image. |
| output_formatstringdefault "mp4" | "mp4", "mov"Output container. mov = H.264 yuv444p with PCM audio (often not playable in browsers). Any other value is rejected with 400. |
| generate_audiobooleandefault true | Generate a native audio track. Default true. |
| seedintegerdefault -1 | Random seed; -1 = random. |
| imagestring | First-frame image URL (or provide images[0]). |
| imagesarray<string> | [first_frame_url] or [first_frame_url, last_frame_url]. images[0] is the first frame (overrides image); when there is more than one entry, the last element is used as the last frame. |
Input files
image: 1 public image URL used as the first frame (JPEG/PNG/WEBP; max 30MB)images: optional [first_frame_url, last_frame_url]
Upload folders for /api/storage/upload
- image:
seedance25
Rules
- duration must be a JSON string matching /^-?\d+$/ and equal -1 or 4-30, otherwise 400
- output_format must be mp4 or mov (400 otherwise)
- Always send
video_qualityas 480p, 720p or 1080p ('4k' is treated as 1080p) - No model_version parameter (single tier)
- Price can be previewed with POST /api/quote; /api/generate-video returns 402 if credits are insufficient. Credits for failed generations are refunded automatically.
- image-to-video requires
imageorimages(400 'Image is required for image-to-video generation') - aspect_ratio has no effect for image-to-video; the output follows the first-frame image
- Price is the same as text-to-video (images do not add cost)
- 480p results are usually drafts: the status response then includes
seedance25_draft, and the draft can be upgraded once to a 1080p version within 7 days withPOST /api/video-generation/{id}/upgrade-draft.
Example request body
{
"mode": "seedance-2.5",
"generation_type": "image-to-video",
"prompt": "The camera slowly pushes in as leaves drift past",
"image": "https://example.com/first.jpg",
"images": [
"https://example.com/first.jpg",
"https://example.com/last.jpg"
],
"video_quality": "720p",
"duration": "8",
"aspect_ratio": "auto",
"output_format": "mp4",
"generate_audio": true,
"seed": -1
}reference-to-video
Generates a 4-30s video guided by up to 30 reference images, 10 reference videos and 10 reference audio clips (audio-only allowed) with Seedance 2.5.
| Parameter | Values & notes |
|---|---|
| promptstring | Text prompt (max 20000 chars). Optional when reference assets are given; can refer to them (e.g. 'image 1', 'video 1'). |
| video_qualitystringaffects pricedefault "480p" | "480p", "720p", "1080p"Output resolution; always send video_quality explicitly. 1080p is 10-bit HEVC (may not play in some browsers). '4k'/'4K'/'2160p' are treated as 1080p. |
| durationstringaffects pricedefault "5" | 4–30Output seconds as a string containing a plain integer from 4 to 30, or "-1" to let the model choose the length. Anything else (e.g. "3", "31", "5.0", "0x10") is rejected with 400. |
| aspect_ratiostringdefault "16:9" | "auto", "16:9", "9:16", "1:1", "4:3", "3:4", "21:9"Output aspect ratio; 'auto' lets the model choose. Send aspect_ratio explicitly and use one of the listed values. |
| output_formatstringdefault "mp4" | "mp4", "mov"Output container. mov = H.264 yuv444p with PCM audio (often not playable in browsers). Any other value is rejected with 400. |
| generate_audiobooleandefault true | Generate a native audio track. Default true. |
| seedintegerdefault -1 | Random seed; -1 = random. |
| imagesarray<string> | Reference image URLs. Max 30. |
| videosarray<string>affects price | Reference video URLs (mp4/mov). Max 10, total length <= 30s. Including reference videos switches to reference-video pricing. |
| input_video_durationnumberaffects price | 0–30Total length in seconds of all reference videos (round the sum up to a whole second). Max 30; larger values are rejected with 400. Send it equal to the total reference video length whenever videos is sent. Minimum billed input length is 2s. |
| audiostring | Optional single reference audio URL; may be the same as audios[0]. |
| audiosarray<string> | Reference audio URLs (wav/mp3). Max 10, total length <= 30s. Audio-only reference is allowed. |
Input files
images: 0-30 reference image URLs (JPEG/PNG/WEBP; max 30MB each)videos: 0-10 reference video URLs (mp4/mov, max 50MB each, total <= 30s)audios: 0-10 reference audio URLs (wav/mp3, max 15MB each, total <= 30s)
Upload folders for /api/storage/upload
- image:
seedance25 - video:
seedance25/reference-videos - audio:
seedance25/audio
Rules
- duration must be a JSON string matching /^-?\d+$/ and equal -1 or 4-30, otherwise 400
- output_format must be mp4 or mov (400 otherwise)
- Always send
video_qualityas 480p, 720p or 1080p ('4k' is treated as 1080p) - No model_version parameter (single tier)
- Price can be previewed with POST /api/quote; /api/generate-video returns 402 if credits are insufficient. Credits for failed generations are refunded automatically.
- At least one reference image, video or audio is required; the generation fails otherwise
- Send
input_video_durationequal to the total reference video length (seconds) whenevervideosis sent; values > 30 are rejected with 400 - aspect_ratio is honored for reference-to-video
- Without reference videos (images/audio only), price equals text-to-video price
- 480p results are usually drafts: the status response then includes
seedance25_draft, and the draft can be upgraded once to a 1080p version within 7 days withPOST /api/video-generation/{id}/upgrade-draft. - When quoting with
/api/quote, addhas_video_input: trueandinput_video_durationif you send reference videos (the quote endpoint does not read video URLs, so the price would come out too low).
Example request body
{
"mode": "seedance-2.5",
"generation_type": "reference-to-video",
"prompt": "The product from image 1 rotates on a marble table, lit like video 1",
"images": [
"https://example.com/product.png"
],
"videos": [
"https://example.com/lighting.mp4"
],
"input_video_duration": 8,
"video_quality": "480p",
"duration": "10",
"aspect_ratio": "9:16",
"output_format": "mp4",
"generate_audio": true,
"seed": -1
}Wan 2.6 mode: wan-2-6
text-to-video
Generates a 5-15 second video with synchronized audio from a text prompt using Wan 2.6.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt (max 20000 chars). Always send a non-empty prompt. |
| negative_promptstring | Optional negative prompt (max 2000 chars). Omit when empty. |
| durationstringaffects pricedefault "5" | "5", "10", "15"Output seconds, sent as a string. |
| video_qualitystringaffects pricedefault "720p" | "720p", "1080p"Output resolution. Use one of the listed values. |
| seedintegerdefault -1 | Random seed; -1 = random, otherwise 0-2147483647. |
| enable_prompt_expansionbooleandefault true | Let the model rewrite/expand the prompt. |
| shot_typestringdefault "single" | "single", "multi"multi = multi-shot video. |
| aspect_ratiostringdefault "16:9" | "16:9", "9:16", "1:1", "4:3", "3:4"Output aspect ratio. Use one of the listed values. |
| audiostring | Optional driving/background audio URL (mp3/ogg/wav/m4a/aac, max 15MB) for audio-visual sync. |
Input files
audio: optional 1 audio URL (mp3/ogg/wav/m4a/aac)
Upload folders for /api/storage/upload
- audio:
wan26/audio
Rules
- duration must be one of 5, 10, 15 (default 5)
- video_quality and aspect_ratio are not validated; always send one of the listed values
Example request body
{
"mode": "wan-2-6",
"generation_type": "text-to-video",
"prompt": "A red fox running through fresh snow at sunrise, cinematic",
"duration": "5",
"video_quality": "720p",
"aspect_ratio": "16:9"
}image-to-video
Animates a single image into a 5-15 second video with Wan 2.6.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt (max 20000 chars). Always send a non-empty prompt. |
| negative_promptstring | Optional negative prompt (max 2000 chars). Omit when empty. |
| imagestring · required | Public https URL of the source image (used as the first frame). |
| durationstringaffects pricedefault "5" | "5", "10", "15"Output seconds, sent as a string. |
| video_qualitystringaffects pricedefault "720p" | "720p", "1080p"Output resolution. Use one of the listed values. |
| seedintegerdefault -1 | Random seed; -1 = random, otherwise 0-2147483647. |
| enable_prompt_expansionbooleandefault true | Let the model rewrite/expand the prompt. |
| shot_typestringdefault "single" | "single", "multi"multi = multi-shot video. |
| audiostring | Optional driving/background audio URL (mp3/ogg/wav/m4a/aac, max 15MB) for audio-visual sync. |
Input files
image: 1 image URL (first frame); jpg/png/webp, max 10MBaudio: optional 1 audio URL (mp3/ogg/wav/m4a/aac)
Upload folders for /api/storage/upload
- image:
video-generator - audio:
wan26/audio
Rules
- image (or images[0]) is required
- aspect_ratio is not used for image-to-video; the output follows the image
- duration must be one of 5, 10, 15
Example request body
{
"mode": "wan-2-6",
"generation_type": "image-to-video",
"prompt": "The character turns and smiles at the camera",
"image": "https://example.com/input.jpg",
"duration": "5",
"video_quality": "720p"
}reference-to-video
Generates a new video guided by 1-3 reference videos (subject/motion reference) with Wan 2.6.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt (max 20000 chars). Always send a non-empty prompt. |
| negative_promptstring | Optional negative prompt (max 2000 chars). Omit when empty. |
| videosarray<string> · required | 1-3 reference video URLs. |
| videostring | Optional single reference video URL; either video or videos satisfies the requirement. |
| durationstringaffects pricedefault "5" | "5", "10"Output seconds: 5 or 10 (15 is rejected for reference-to-video). |
| video_qualitystringaffects pricedefault "720p" | "720p", "1080p"Output resolution. Use one of the listed values. |
| seedintegerdefault -1 | Random seed; -1 = random, otherwise 0-2147483647. |
| enable_prompt_expansionbooleandefault true | Let the model rewrite/expand the prompt. |
| shot_typestringdefault "single" | "single", "multi"multi = multi-shot video. |
| aspect_ratiostringdefault "16:9" | "16:9", "9:16", "1:1", "4:3", "3:4"Output aspect ratio. Use one of the listed values. |
Input files
videos: 1-3 reference video URLs (mp4/mov/mkv; max 10MB each)
Upload folders for /api/storage/upload
- video:
wan26/reference-videos
Rules
- At least one reference video (video or videos) is required
- At most 3 reference videos
- duration 15 is rejected (only 5 or 10)
- Audio input is not supported for reference-to-video; do not send
audio
Example request body
{
"mode": "wan-2-6",
"generation_type": "reference-to-video",
"prompt": "The character from the reference video dances in a neon-lit street",
"videos": [
"https://example.com/ref.mp4"
],
"duration": "5",
"video_quality": "720p",
"aspect_ratio": "16:9"
}Wan 2.7 mode: wan-2.7
text-to-video
Generates a 2-15 second video from a text prompt with Wan 2.7 (smoother motion, better scene fidelity).
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt (max 20000 chars). Always send a non-empty prompt. |
| negative_promptstring | Optional negative prompt (max 2000 chars). Omit when empty. |
| durationstringaffects pricedefault "5" | 2–15Output seconds as an integer string. |
| video_qualitystringaffects pricedefault "720p" | "720p", "1080p"Output resolution. |
| aspect_ratiostringdefault "16:9" | "16:9", "9:16", "1:1", "4:3", "3:4"Output aspect ratio. |
| enable_prompt_expansionbooleandefault true | Let the model expand the prompt. |
| audiostring | Optional audio URL to drive the video (mp3/ogg/wav/m4a/aac, max 15MB). |
| seedintegerdefault -1 | Random seed; -1 = random. |
Input files
audio: optional 1 audio URL (mp3/ogg/wav/m4a/aac)
Upload folders for /api/storage/upload
- audio:
wan27/audio
Rules
- duration must be an integer 2-15 (default 5)
- video_quality, if sent, must be 720p or 1080p
- aspect_ratio, if sent, must be one of 16:9, 9:16, 1:1, 4:3, 3:4
Example request body
{
"mode": "wan-2.7",
"generation_type": "text-to-video",
"prompt": "A paper boat drifting down a rain-soaked city gutter, macro shot",
"duration": "5",
"video_quality": "720p",
"aspect_ratio": "16:9"
}image-to-video
Animates a start image (optionally to an end frame) into a 2-15 second video with Wan 2.7.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt (max 20000 chars). Always send a non-empty prompt. |
| negative_promptstring | Optional negative prompt (max 2000 chars). Omit when empty. |
| imagestring | Start-frame image URL. Required unless you send it as images[0]. |
| imagesarray<string> | Only for first/last-frame mode: [start_image_url, end_image_url]; images[1] becomes the end frame. |
| durationstringaffects pricedefault "5" | 2–15Output seconds as an integer string. |
| video_qualitystringaffects pricedefault "720p" | "720p", "1080p"Output resolution. |
| enable_prompt_expansionbooleandefault true | Let the model expand the prompt. |
| audiostring | Optional audio URL to drive the video (mp3/ogg/wav/m4a/aac, max 15MB). |
| seedintegerdefault -1 | Random seed; -1 = random. |
Input files
image: 1 start-frame image URL (jpg/png/webp, max 20MB)images: optional [start, end] pair for first/last-frame interpolationaudio: optional 1 audio URL
Upload folders for /api/storage/upload
- image:
video-generator-wan27 - audio:
wan27/audio
Rules
- image (or images[0]) is required
- aspect_ratio is not needed for image-to-video (the output follows the image); if sent it must be one of the listed values
- duration must be an integer 2-15
Example request body
{
"mode": "wan-2.7",
"generation_type": "image-to-video",
"prompt": "Gentle wind moves the hair, camera slowly pushes in",
"image": "https://example.com/start.jpg",
"duration": "5",
"video_quality": "720p"
}reference-to-video
Generates a video guided by up to 3 reference images and/or up to 3 reference videos with Wan 2.7.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt (max 20000 chars). Always send a non-empty prompt. |
| negative_promptstring | Optional negative prompt (max 2000 chars). Omit when empty. |
| imagesarray<string> | 0-3 reference image URLs. |
| videosarray<string> | 0-3 reference video URLs. |
| videostring | Optional single reference video URL; may be the same as videos[0]. |
| durationstringaffects pricedefault "5" | 2–15Output seconds as an integer string. |
| video_qualitystringaffects pricedefault "720p" | "720p", "1080p"Output resolution; 720p when omitted. |
| aspect_ratiostringdefault "16:9" | "16:9", "9:16", "1:1", "4:3", "3:4"Output aspect ratio. |
| shot_typestringdefault "single" | "single", "multi"single or multi-shot output. |
| seedintegerdefault -1 | Random seed; -1 = random (values >= 0 fix the seed). |
Input files
images: 0-3 reference image URLs (jpg/png/webp, max 20MB)videos: 0-3 reference video URLs (mp4/mov, max 100MB each)
Upload folders for /api/storage/upload
- image:
video-generator-wan27 - video:
wan27/reference-videos
Rules
- At least one reference image or video is required
- At most 3 reference images and at most 3 reference videos
- duration must be an integer 2-15
- audio and enable_prompt_expansion are not supported for reference-to-video
Example request body
{
"mode": "wan-2.7",
"generation_type": "reference-to-video",
"prompt": "The woman from the reference image walks through a sunflower field",
"images": [
"https://example.com/ref.jpg"
],
"duration": "5",
"video_quality": "720p",
"aspect_ratio": "16:9",
"shot_type": "single"
}Wan 3.0 mode: wan-3.0
text-to-video
Generates a 2-30 second video with optional native audio from a text prompt using Wan 3.0.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt. Required for text-to-video. |
| durationstringaffects pricedefault "5" | 2–30Output seconds as an integer string (2-30); -1 (smart duration) is rejected. |
| video_qualitystringaffects pricedefault "720p" | "480p", "720p", "1080p"Resolution. Use one of the listed values ('4k' is treated as 1080p). |
| aspect_ratiostringdefault "16:9" | "adaptive", "16:9", "9:16", "1:1", "4:3", "3:4"Output aspect ratio; 'adaptive' lets the model choose. Send aspect_ratio explicitly and use one of the listed values. |
| generate_audiobooleandefault true | Generate a synchronized soundtrack (default true; send false to turn it off). Does not change the price. |
| seedinteger | Accepted but currently not passed to the model. |
Rules
- duration must be an integer 2-30; -1 (smart duration) is rejected
- generate_audio does not change the price
Example request body
{
"mode": "wan-3.0",
"generation_type": "text-to-video",
"prompt": "A drone shot over a misty mountain lake at dawn, birds taking off",
"duration": "5",
"video_quality": "720p",
"aspect_ratio": "adaptive",
"generate_audio": true
}image-to-video
Animates a first frame (optionally to a last frame) into a 2-30 second video with Wan 3.0.
| Parameter | Values & notes |
|---|---|
| promptstring | Motion/scene prompt (optional for image-to-video). |
| imagestring | First-frame image URL. Required unless you send it as images[0]. |
| imagesarray<string> | [first_frame] or [first_frame, last_frame]; images[0] is the first frame and the last element (when 2 or more) is the last frame. |
| durationstringaffects pricedefault "5" | 2–30Output seconds as an integer string (2-30); -1 (smart duration) is rejected. |
| video_qualitystringaffects pricedefault "720p" | "480p", "720p", "1080p"Resolution. Use one of the listed values ('4k' is treated as 1080p). |
| aspect_ratiostringdefault "16:9" | "adaptive", "16:9", "9:16", "1:1", "4:3", "3:4"Output aspect ratio; 'adaptive' lets the model choose. Send aspect_ratio explicitly and use one of the listed values. |
| generate_audiobooleandefault true | Generate a synchronized soundtrack (default true; send false to turn it off). Does not change the price. |
| seedinteger | Accepted but currently not passed to the model. |
Input files
image: 1 first-frame image URL (jpg/png/webp/bmp/tiff, max 20MB)images: optional [first, last] frame pair
Upload folders for /api/storage/upload
- image:
wan30
Rules
- image or images is required
- files (document) and links (webpage) cannot be combined with image-to-video
- duration must be an integer 2-30; -1 rejected
Example request body
{
"mode": "wan-3.0",
"generation_type": "image-to-video",
"prompt": "The cat slowly opens its eyes and yawns",
"image": "https://example.com/first.jpg",
"images": [
"https://example.com/first.jpg"
],
"duration": "5",
"video_quality": "720p",
"aspect_ratio": "adaptive",
"generate_audio": true
}reference-to-video
Generates a 2-30 second video from any mix of reference images, videos, audio, one document (document-to-video) or one webpage (webpage-to-video) with Wan 3.0.
| Parameter | Values & notes |
|---|---|
| promptstring | Instruction describing how to use the references (e.g. refer to 'Image 1', 'Video 1'). |
| imagesarray<string> | 0-10 reference image URLs. |
| videosarray<string> | 0-5 reference video URLs, total length <= 15s. Always pass reference videos in videos (plural), together with input_video_duration. |
| input_video_durationnumberaffects price | 0–15Total seconds of all reference videos (round the sum up to a whole second; max 15). Send it equal to the total reference video length whenever videos is sent. |
| audiosarray<string> | 0-5 reference audio URLs (wav/mp3), total <= 15s. |
| audiostring | Single reference audio URL; used only when audios is absent. |
| filesarray<string> | Document-to-video: at most 1 document URL (.doc .docx .xls .xlsx .ppt .pptx .pdf .txt .md .key .pages .numbers; <=100MB, <=50 pages). Mutually exclusive with links. |
| linksarray<string> | Webpage-to-video: at most 1 public webpage URL. Mutually exclusive with files. |
| durationstringaffects pricedefault "5" | 2–30Output seconds as an integer string (2-30); -1 (smart duration) is rejected. |
| video_qualitystringaffects pricedefault "720p" | "480p", "720p", "1080p"Resolution. Use one of the listed values ('4k' is treated as 1080p). |
| aspect_ratiostringdefault "16:9" | "adaptive", "16:9", "9:16", "1:1", "4:3", "3:4"Output aspect ratio; 'adaptive' lets the model choose. Send aspect_ratio explicitly and use one of the listed values. |
| generate_audiobooleandefault true | Generate a synchronized soundtrack (default true; send false to turn it off). Does not change the price. |
| seedinteger | Accepted but currently not passed to the model. |
Input files
images: 0-10 image URLs (jpg/png/webp/bmp/tiff, max 20MB each)videos: 0-5 video URLs (mp4/mov, max 100MB each), combined <= 15saudio: 0-5 audio URLs via audios (wav/mp3, max 15MB each), combined <= 15sfiles: 0-1 document URL (100MB, 50 pages max)links: 0-1 webpage URL
Upload folders for /api/storage/upload
- image:
wan30 - video:
wan30/reference-videos - audio:
wan30/audio - document:
wan30/documents
Rules
- At most 10 images, 5 videos, 5 audios, 1 file, 1 link (exceeding returns 400, never truncated)
- files and links cannot be sent together
- Provide at least one reference image, video, audio, document or webpage
- input_video_duration must be <= 15, and input_video_duration + duration must be <= 30 (400 otherwise)
- Always pass reference videos in
videos(plural) and sendinput_video_durationequal to their total length (seconds) - duration must be an integer 2-30; -1 rejected
- When quoting with
/api/quote, addhas_video_input: trueandinput_video_durationif you send reference videos (the quote endpoint does not read video URLs, so the price would come out too low).
Example request body
{
"mode": "wan-3.0",
"generation_type": "reference-to-video",
"prompt": "Turn this product page into a 10 second promo video",
"links": [
"https://example.com/product"
],
"duration": "10",
"video_quality": "720p",
"aspect_ratio": "16:9",
"generate_audio": true
}MiniMax H3 mode: minimax-h3
text-to-video
Generates a 4-15 second video from a text prompt with MiniMax H3 (Hailuo 03) at 768P or 2K.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt (max 20000 chars). Always send a non-empty prompt. |
| durationstringaffects pricedefault "5" | 4–15Output seconds as an integer string. |
| video_qualitystringaffects pricedefault "768p" | "768p", "2k"Resolution. '2K', '1440p', '2160p' are treated as 2k. Use one of the listed values. |
| aspect_ratiostringdefault "16:9" | "21:9", "16:9", "4:3", "1:1", "3:4", "9:16"Output aspect ratio. 'adaptive' is not supported for text-to-video; use one of the listed values. |
Rules
- duration must be an integer between 4 and 15 (400 otherwise); sent as a string
- Seed, negative prompt and audio options are not supported for this model
Example request body
{
"mode": "minimax-h3",
"generation_type": "text-to-video",
"prompt": "A red fox running through fresh snow at sunrise, cinematic tracking shot",
"duration": "5",
"aspect_ratio": "16:9",
"video_quality": "768p"
}image-to-video
Animates a first-frame image (optionally toward a last frame) into a 4-15 second MiniMax H3 video at 768P or 2K.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt describing the motion. Always send a non-empty prompt. |
| durationstringaffects pricedefault "5" | 4–15Output seconds as an integer string. |
| video_qualitystringaffects pricedefault "768p" | "768p", "2k" |
| imagestring | First frame image URL. Either image or images[0] must be present. |
| imagesstring[]affects price | [first_frame] or [first_frame, last_frame]. images[0] is the first frame; with 2 or more entries the LAST element is used as the last frame. Max 9. |
| aspect_ratiostringdefault "adaptive" | "adaptive", "21:9", "16:9", "4:3", "1:1", "3:4", "9:16"Ignored for image-to-video; the output follows the first frame. |
Input files
image: 1 public image URL used as the first frame (jpg/png/webp, max 20MB)images: Optional array [first_frame_url, last_frame_url]; only the first and last entries are used
Upload folders for /api/storage/upload
- image:
minimax-h3
Rules
- image-to-video requires image or a non-empty images array (400 'Image is required for image-to-video generation')
- duration must be an integer 4-15
- At most 9 images
- The last frame is taken from the last element of images only when images has 2 or more entries
- Only the first and last images are used, but every image after the 5th is charged: send at most 2. When quoting with
/api/quote, sendinput_image_count.
Example request body
{
"mode": "minimax-h3",
"generation_type": "image-to-video",
"prompt": "The woman turns toward the camera and smiles as wind moves her hair",
"duration": "5",
"aspect_ratio": "adaptive",
"video_quality": "768p",
"image": "https://example.com/first-frame.jpg",
"images": [
"https://example.com/first-frame.jpg"
]
}reference-to-video
Generates a MiniMax H3 video guided by up to 9 reference images, 3 reference videos and 3 reference audio clips.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt; can refer to the reference assets. Always send a non-empty prompt. |
| durationstringaffects pricedefault "5" | 4–15Output seconds as an integer string. |
| video_qualitystringaffects pricedefault "768p" | "768p", "2k" |
| aspect_ratiostringdefault "16:9" | "adaptive", "21:9", "16:9", "4:3", "1:1", "3:4", "9:16"'adaptive' lets the model choose from the references. Use one of the listed values. |
| imagesstring[]affects price | Reference image URLs, max 9. |
| videosstring[]affects price | Reference video URLs, max 3, total duration <= 15s. Singular 'video' is accepted as a fallback when 'videos' is absent. |
| input_video_durationnumberaffects price | 0–15Total seconds of the reference videos (max 15). Send it equal to the total reference video length whenever videos are sent. |
| audiosstring[] | Reference audio URLs, max 3 (free). |
| audiostring | Singular reference audio; used only if 'audios' is absent. |
Input files
images: 0-9 reference image URLs (jpg/png/webp/bmp/tiff, 20MB each)video: 0-3 reference video URLs via 'videos' (mp4/mov, 100MB each, total <= 15s)audio: 0-3 reference audio URLs via 'audios' (wav/mp3, 15MB each); cannot be the only reference
Upload folders for /api/storage/upload
- image:
minimax-h3 - video:
minimax-h3/reference-videos - audio:
minimax-h3/audio
Rules
- At most 9 reference images, 3 reference videos, 3 reference audio files (400 if exceeded, no silent truncation)
- At least one reference image or video is required; reference audio cannot be used alone
- input_video_duration must be <= 15 seconds in total
- Send
input_video_durationequal to the total reference video length (seconds) whenever videos are sent - duration must be an integer 4-15
- When quoting with
/api/quote, addhas_video_input: trueandinput_video_durationif you send reference videos (the quote endpoint does not read video URLs, so the price would come out too low). Also sendinput_image_count(number of reference images).
Example request body
{
"mode": "minimax-h3",
"generation_type": "reference-to-video",
"prompt": "The character from image 1 walks into the room from video 1 and waves",
"duration": "5",
"aspect_ratio": "adaptive",
"video_quality": "768p",
"images": [
"https://example.com/character.png"
],
"videos": [
"https://example.com/room.mp4"
],
"input_video_duration": 6
}MiniMax H3 Max mode: minimax-h3-max
text-to-video
Fast (turbo) MiniMax H3 Max text-to-video, 5-15 seconds at 480P/768P/1080P.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt. Required; always send a non-empty prompt. |
| durationstringaffects pricedefault "5" | 5–15Output seconds as an integer string. |
| video_qualitystringaffects pricedefault "768p" | "480p", "768p", "1080p"Use one of the listed values (2k is not supported). |
| aspect_ratiostringdefault "16:9" | "21:9", "16:9", "4:3", "1:1", "3:4", "9:16"'adaptive' is not supported for text-to-video. Use one of the listed values. |
| prompt_expansion_modestringdefault "balanced" | "disabled", "fast", "balanced", "quality"Prompt rewriting strength. Values outside the list are rejected with 400. |
| seedinteger | Optional; a value >= 0 fixes the random seed. |
Rules
- generation_type must be text-to-video, image-to-video or reference-to-video for this mode (400 otherwise)
- duration must be an integer between 5 and 15
Example request body
{
"mode": "minimax-h3-max",
"generation_type": "text-to-video",
"prompt": "A paper boat drifting down a rainy city gutter, macro shot",
"duration": "5",
"aspect_ratio": "16:9",
"video_quality": "768p",
"prompt_expansion_mode": "balanced"
}image-to-video
Fast MiniMax H3 Max image-to-video from a first frame (optional end frame), 5-15 seconds.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Motion prompt. Required; always send a non-empty prompt. |
| durationstringaffects pricedefault "5" | 5–15Output seconds as an integer string. |
| video_qualitystringaffects pricedefault "768p" | "480p", "768p", "1080p" |
| imagestring | First frame URL (image or images[0] required). |
| imagesstring[] | [first_frame] or [first_frame, end_frame]. With 2 or more entries, the last one is used as the end frame. |
| aspect_ratiostringdefault "16:9" | "21:9", "16:9", "4:3", "1:1", "3:4", "9:16"Ignored for image-to-video; the output follows the first frame. |
| prompt_expansion_modestringdefault "balanced" | "disabled", "fast", "balanced", "quality"Prompt rewriting strength. |
| seedinteger | Optional; a value >= 0 fixes the random seed. |
Input files
image: 1 public image URL used as the first frameimages: Optional [first_frame_url, end_frame_url]
Upload folders for /api/storage/upload
- image:
minimax-h3-max
Rules
- image-to-video requires a first frame (image or images[0]); 400 otherwise
- duration must be an integer 5-15
Example request body
{
"mode": "minimax-h3-max",
"generation_type": "image-to-video",
"prompt": "The camera slowly pushes in as the dog tilts its head",
"duration": "5",
"aspect_ratio": "16:9",
"video_quality": "768p",
"prompt_expansion_mode": "balanced",
"image": "https://example.com/dog.jpg",
"images": [
"https://example.com/dog.jpg"
]
}reference-to-video
Fast MiniMax H3 Max video guided by reference images, videos and audio (up to 12 files total).
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt. Required; always send a non-empty prompt. |
| durationstringaffects pricedefault "5" | 5–15Output seconds as an integer string. |
| video_qualitystringaffects pricedefault "768p" | "480p", "768p", "1080p" |
| aspect_ratiostringdefault "16:9" | "adaptive", "21:9", "16:9", "4:3", "1:1", "3:4", "9:16"'adaptive' is available only for reference-to-video. Use one of the listed values. |
| prompt_expansion_modestringdefault "balanced" | "disabled", "fast", "balanced", "quality"Prompt rewriting strength. |
| imagesstring[]affects price | Reference image URLs, max 9. Resize images so the long edge is <= 1024 px. |
| videosstring[]affects price | Reference video URLs, max 3, each 2-15s, total <= 15s. Singular 'video' accepted as fallback. |
| audiosstring[]affects price | Reference audio URLs, max 3, each 2-15s, total <= 15s. Singular 'audio' accepted as fallback. |
| input_video_durationnumberaffects price | 0–15Total reference video seconds (max 15). Send it equal to the total reference video length whenever videos are sent. |
| input_audio_durationnumberaffects price | 0–15Total reference audio seconds (max 15). Send it equal to the total reference audio length whenever audios are sent. |
| seedinteger | Optional; a value >= 0 fixes the random seed. |
Input files
images: 0-9 reference image URLs (jpg/png/webp, 20MB, long edge <= 1024 px)video: 0-3 reference video URLs via 'videos' (mp4/mov, 100MB each, 2-15s each, total <= 15s)audio: 0-3 reference audio URLs via 'audios' (wav/mp3, 15MB each, 2-15s each, total <= 15s); cannot be the only reference
Upload folders for /api/storage/upload
- image:
minimax-h3-max - video:
minimax-h3-max/reference-videos - audio:
minimax-h3-max/audio
Rules
- At least one reference image, video or audio is required
- Reference audio cannot be used alone; add at least one image or video
- Max 9 images, 3 videos, 3 audio files, and max 12 files in total
- Reference videos and reference audio each total <= 15s (two separate limits); each clip must be at least 2s
- Send
input_video_duration/input_audio_durationequal to the total reference video / audio length (seconds) - duration must be an integer 5-15
- Resize reference images so the long edge is <= 1024 px
- When quoting with
/api/quote, describe the references withinput_image_count,has_video_input+input_video_count+input_video_duration, andhas_audio_input+input_audio_count+input_audio_duration(the quote endpoint does not read video URLs, so the price would come out too low).
Example request body
{
"mode": "minimax-h3-max",
"generation_type": "reference-to-video",
"prompt": "The product from image 1 rotates on a marble table, soft studio light",
"duration": "5",
"aspect_ratio": "adaptive",
"video_quality": "768p",
"prompt_expansion_mode": "balanced",
"images": [
"https://example.com/product.jpg"
]
}MiniMax H3 Max Camera Controls mode: minimax-h3-max-camera
camera-controls
Turns a single image into a 5-15 second video with a scripted camera move (orbit, dolly, crane, etc.) using MiniMax H3 Max.
| Parameter | Values & notes |
|---|---|
| imagestring | Source image URL (image or images[0]). Output aspect ratio follows the image. |
| imagesstring[] | Optional; if sent, must contain at most 1 entry (the source image). No end frame. |
| camera_presetstringdefault "orbit-360" | "orbit-360", "birds-eye", "whip-spin", "hero-reveal", "dolly-in", "dolly-out", "arc-left", "arc-right", "worms-eye"Named trajectory. Unknown names are rejected (400). Ignored when camera_trajectory is provided. Defaults to orbit-360 when both are omitted. |
| camera_trajectoryarray | Explicit keyframes. Array of 2-12 objects {time: number 0..1 strictly increasing, azimuth: degrees (signed, multi-turn allowed), elevation: number -90..90, distance: number > 0}. Total |azimuth| travel <= 11520 degrees (32 turns). Takes precedence over camera_preset. |
| promptstring | Optional. If empty, the default instruction 'The entire scene is frozen. Only the camera moves.' is used. |
| durationstringaffects pricedefault "5" | 5–15Output seconds as an integer string. |
| video_qualitystringaffects pricedefault "768p" | "480p", "768p", "1080p"Use one of the listed values. |
| prompt_expansion_modestringdefault "balanced" | "disabled", "fast", "balanced", "quality"Prompt rewriting strength. |
| seedinteger | Optional non-negative integer for reproducible results; omit for random. |
Input files
image: Exactly 1 public image URL (source image / first frame); no end frame, no reference videos or audio
Upload folders for /api/storage/upload
- image:
minimax-h3-max
Rules
- generation_type must be 'camera-controls' (image-to-video is rejected with 400)
- A source image is required; more than 1 image, any video, or any audio is rejected with 400
- duration must be an integer 5-15
- camera_preset must be one of the 9 preset ids unless camera_trajectory is given
- camera_trajectory: 2-12 keyframes, time in [0,1] strictly increasing, elevation in [-90,90], distance > 0, total azimuth travel <= 11520 degrees
- aspect_ratio has no effect (output follows the source image)
Example request body
{
"mode": "minimax-h3-max-camera",
"generation_type": "camera-controls",
"prompt": "",
"duration": "5",
"video_quality": "768p",
"prompt_expansion_mode": "balanced",
"image": "https://example.com/statue.jpg",
"images": [
"https://example.com/statue.jpg"
],
"camera_preset": "orbit-360"
}MiniMax H3 Max Lip Sync mode: minimax-h3-max-lip-sync
lip-sync
Makes a portrait photo speak in sync with an audio track (5-14.8s) using MiniMax H3 Max; output length follows the audio.
| Parameter | Values & notes |
|---|---|
| imagestring | Portrait image URL (image or images[0]). Image aspect ratio (width/height) must be between 0.4 and 2.5. |
| imagesstring[] | Optional; if sent, must contain at most 1 entry (the portrait image). |
| audiosstring[] · required | Exactly 1 driving audio URL. Singular audio is accepted as a fallback if audios is absent. |
| input_audio_durationnumberaffects price | 5–3600Audio length in seconds. Always send it, equal to the audio length. Billed seconds = ceil(min(input_audio_duration, 14.8)), min 5. Values < 5 are rejected. |
| video_qualitystringaffects pricedefault "768p" | "480p", "768p", "1080p", "2k"Use one of the listed values. |
| enable_transcriptionbooleandefault true | Transcribe the audio to guide mouth shapes. Defaults to true. |
| seedinteger | Optional non-negative integer for reproducible results; omit for random. |
Input files
image: 1 public portrait image URL (aspect ratio 0.4-2.5)audio: 1 public audio URL via 'audios' (wav/mp3/m4a/aac, up to 15MB); at least 5s; audio over 14.8s is clipped to the first 14.8s
Upload folders for /api/storage/upload
- image:
minimax-h3-max - audio:
minimax-h3-max/audio
Rules
- generation_type must be 'lip-sync' (400 otherwise)
- Requires one portrait image and one audio track; more than 1 image, more than 1 audio, or any video input is rejected (400)
- Portrait image aspect ratio (width/height) must be between 0.4 and 2.5
- Send input_audio_duration equal to the audio length; it must be at least 5s (400). Audio longer than 14.8s is allowed and clipped
- prompt, duration, aspect_ratio and prompt_expansion_mode have no effect for this mode
- Cannot re-sync an existing video; image-to-video lip sync only
- If the audio turns out longer than the declared input_audio_duration, the difference is charged after generation.
Example request body
{
"mode": "minimax-h3-max-lip-sync",
"generation_type": "lip-sync",
"image": "https://example.com/portrait.jpg",
"images": [
"https://example.com/portrait.jpg"
],
"audios": [
"https://example.com/speech.mp3"
],
"input_audio_duration": 8.4,
"video_quality": "768p",
"enable_transcription": true
}Gemini Omni mode: gemini-omni
text-to-video
Google Gemini Omni text-to-video with native audio, 4-10s, optionally guided by one reference video.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt (required, non-empty), max 20000 characters. |
| durationstringaffects pricedefault "8" | "4", "6", "8", "10"Seconds. Ignored (and not billed) when a video is supplied; output length then follows the input video clip (max 10s). |
| video_qualitystringaffects pricedefault "1080p" | "720p", "1080p", "4k"Output resolution. The resolution field with the same values is also accepted and takes precedence. |
| aspect_ratiostringdefault "16:9" | "16:9", "9:16"Only landscape or portrait. |
| videosstring[]affects price | Optional, at most 1 video URL used as motion/style reference. |
| input_video_durationnumber | 0–30Length of the supplied input video in seconds. Send it whenever a video is supplied; the output covers the first min(input_video_duration, 10) seconds. |
| seedinteger | 0–2147483647Optional, 0-2147483647; omit for random. |
Input files
videos: optional, max 1 video URL (mp4/mov, up to 100MB; input video must be <= 30s)
Upload folders for /api/storage/upload
- video:
gemini-omni/reference-videos
Rules
- duration must be one of 4, 6, 8, 10 (omitted defaults to 8)
- video_quality (if sent) must be 720p, 1080p or 4k; aspect_ratio (if sent) must be 16:9 or 9:16
- at most 1 video; asset quota images + 2*videos <= 7
- input_video_duration must be <= 30s
- prompt is required, max 20000 characters
- When quoting with
/api/quote, addhas_video_input: trueandinput_video_durationif you send a video (the quote endpoint does not read video URLs, so the price would come out too low).
Example request body
{
"mode": "gemini-omni",
"generation_type": "text-to-video",
"prompt": "A golden retriever surfing a wave at sunset, cinematic, with ocean sounds",
"duration": "8",
"video_quality": "1080p",
"aspect_ratio": "16:9"
}image-to-video
Google Gemini Omni image-to-video: animates up to 7 input images into a 4-10s video with native audio.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt, max 20000 characters. |
| imagesstring[] · required | 1-7 image URLs. A single image field is also accepted. |
| durationstringaffects pricedefault "8" | "4", "6", "8", "10"Seconds. Ignored and not billed when a video is supplied. |
| video_qualitystringaffects pricedefault "1080p" | "720p", "1080p", "4k"Output resolution; resolution is also accepted and takes precedence. |
| aspect_ratiostringdefault "16:9" | "16:9", "9:16" |
| videosstring[]affects price | Optional, at most 1 video URL. |
| input_video_durationnumber | 0–30Length of the input video in seconds. Send it whenever a video is supplied; the output covers the first min(input_video_duration, 10) seconds. |
| seedinteger | 0–2147483647 |
Input files
images: 1-7 image URLs (up to 20MB each)videos: optional, max 1 video URL (mp4/mov, <= 30s, up to 100MB)
Upload folders for /api/storage/upload
- image:
gemini-omni - video:
gemini-omni/reference-videos
Rules
- image-to-video requires
imageor a non-emptyimages - max 7 images and max 1 video; images + 2*videos <= 7 (so with a video, at most 5 images)
- video_quality must be 720p, 1080p or 4k; aspect_ratio must be 16:9 or 9:16
- input_video_duration must be <= 30s
- same price table as text-to-video
- duration must be 4, 6, 8 or 10, also when a video is supplied (it is then ignored).
- When quoting with
/api/quote, addhas_video_input: trueandinput_video_durationif you send a video (the quote endpoint does not read video URLs, so the price would come out too low).
Example request body
{
"mode": "gemini-omni",
"generation_type": "image-to-video",
"prompt": "The woman in the photo turns toward the camera and smiles as wind moves her hair",
"images": [
"https://example.com/portrait.jpg"
],
"duration": "6",
"video_quality": "1080p",
"aspect_ratio": "9:16"
}reference-to-video
Google Gemini Omni reference-to-video: generates a video guided by up to 7 reference images and/or 1 reference video.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt, max 20000 characters. |
| imagesstring[] | 0-7 reference image URLs. At least one image or one video is required. |
| videosstring[]affects price | 0-1 reference video URL. |
| input_video_durationnumber | 0–30Length of the reference video in seconds; output length = min(this, 10). Send it whenever a video is supplied. |
| durationstringaffects pricedefault "8" | "4", "6", "8", "10"Only used/billed when no video is supplied. |
| video_qualitystringaffects pricedefault "1080p" | "720p", "1080p", "4k" |
| aspect_ratiostringdefault "16:9" | "16:9", "9:16" |
| seedinteger | 0–2147483647 |
Input files
images: 0-7 reference image URLs (up to 20MB each)videos: 0-1 reference video URL (mp4/mov, <= 30s; only the first 10s is used)
Upload folders for /api/storage/upload
- image:
gemini-omni - video:
gemini-omni/reference-videos
Rules
- at least one reference image or video is required
- reference quota: images <= 7, videos <= 1, images + 2*videos <= 7
- input_video_duration must be <= 30s
- same price table as text-to-video
- duration must be 4, 6, 8 or 10, also when a video is supplied (it is then ignored).
- When quoting with
/api/quote, addhas_video_input: trueandinput_video_durationif you send a reference video (the quote endpoint does not read video URLs, so the price would come out too low).
Example request body
{
"mode": "gemini-omni",
"generation_type": "reference-to-video",
"prompt": "The character from the reference images walks through a neon-lit Tokyo street at night",
"images": [
"https://example.com/character-front.png",
"https://example.com/character-side.png"
],
"duration": "8",
"video_quality": "1080p",
"aspect_ratio": "16:9"
}Boreal mode: creatify-boreal
text-to-video
Creatify Boreal text-to-video for ads, UGC and presenter clips with native synced speech, 1-20s, optionally driven by an uploaded audio track.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Prompt; may optionally be structured with section headers [VISUAL], [SPEECH], [SOUNDS], [TEXT] (dialogue in quotes). |
| durationstringaffects pricedefault "10" | 1–20Integer seconds. |
| video_qualitystringaffects pricedefault "720p" | "720p", "1080p", "2k"Output resolution ('2K' uppercase also accepted). |
| aspect_ratiostringdefault "auto" | "auto", "16:9", "9:16", "1:1", "4:3", "3:4" |
| audiostring | Optional driving audio URL; if given, the audio track is kept and the model lip-syncs to it instead of generating speech. |
Input files
audio: optional, 1 audio URL (.mp3/.ogg/.wav/.m4a/.aac)
Upload folders for /api/storage/upload
- audio:
creatify-boreal/audio
Rules
- generation_type must be text-to-video or image-to-video
- duration must be an integer between 1 and 20 (omitted defaults to 10)
- video_quality must be 720p, 1080p, 2k or 2K; aspect_ratio must be auto, 16:9, 9:16, 1:1, 4:3 or 3:4
- prompt is required
Example request body
{
"mode": "creatify-boreal",
"generation_type": "text-to-video",
"prompt": "[VISUAL]\nA young woman holds a can of sparkling water in a sunny kitchen and talks to the camera.\n\n[SPEECH]\n\"This is the only drink I grab after my workout.\"",
"duration": "10",
"video_quality": "720p",
"aspect_ratio": "9:16"
}image-to-video
Creatify Boreal image-to-video: turns a product or presenter image into an ad/UGC clip with native synced speech, optionally driven by audio.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Prompt; optional [VISUAL]/[SPEECH]/[SOUNDS]/[TEXT] sections. |
| imagestring · required | 1 image URL used as the starting image (images[0] is also accepted). |
| durationstringaffects pricedefault "10" | 1–20Integer seconds. |
| video_qualitystringaffects pricedefault "720p" | "720p", "1080p", "2k" |
| aspect_ratiostringdefault "auto" | "auto", "16:9", "9:16", "1:1", "4:3", "3:4"'auto' follows the input image. |
| audiostring | Optional driving audio URL; the model lip-syncs to it instead of generating speech. |
Input files
image: 1 image URLaudio: optional, 1 audio URL (.mp3/.ogg/.wav/.m4a/.aac)
Upload folders for /api/storage/upload
- image:
creatify-boreal - audio:
creatify-boreal/audio
Rules
- image-to-video requires
imageor non-emptyimages - duration integer 1-20; video_quality 720p/1080p/2k/2K; aspect_ratio auto/16:9/9:16/1:1/4:3/3:4
Example request body
{
"mode": "creatify-boreal",
"generation_type": "image-to-video",
"prompt": "[VISUAL]\nThe presenter picks up the product and shows it to the camera.\n\n[SPEECH]\n\"Meet the lightest headphones we've ever made.\"",
"image": "https://example.com/presenter.jpg",
"duration": "8",
"video_quality": "1080p",
"aspect_ratio": "auto"
}Image Models
Text-to-image and image editing (image-to-image).
| Model | mode | generation_type |
|---|---|---|
| Google Nano Banana | nano-banana | text-to-image |
| Google Nano Banana (Edit) | nano-banana-edit | image-to-image |
| Google Nano Banana 2.1 | nano-banana-2-1 | text-to-image, image-to-image |
| Google Nano Banana 2 | nano-banana-2 | text-to-image, image-to-image |
| Flux Kontext | flux-kontext | text-to-image, image-to-image |
| FLUX 3 Image | flux-3-image | text-to-image, image-to-image |
| Seedream 5.0 | seedream-5.0-lite | text-to-image, image-to-image |
| Seedream 5.0 Pro | seedream-5.0-pro | text-to-image, image-to-image |
| Seedream 5.0 Flash | seedream-5.0-flash | text-to-image, image-to-image |
| Ideogram 4.5 | ideogram-4-5 | text-to-image, image-to-image |
| GPT Image 2 | gpt-image-2 | text-to-image, image-to-image |
| GPT Image 2.5 | gpt-image-2-5 | text-to-image, image-to-image |
| Qwen Image 2.1 | qwen-image-2-1 | text-to-image, image-to-image |
| Virtual Try-On | virtual-try-on | image-to-image |
| Product Holding | product-holding | image-to-image |
Google Nano Banana mode: nano-banana
text-to-image
Generates an image from text with Google Nano Banana (or Nano Banana Pro via modelVersion).
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt, up to 20000 characters. |
| modelVersionstringaffects pricedefault "nano-banana" | "nano-banana", "nano-banana-pro"nano-banana = base model; nano-banana-pro = Nano Banana Pro. |
| video_qualitystringaffects pricedefault "standard" | "standard", "1K", "2K", "4K"Send 'standard' with nano-banana; with nano-banana-pro, send the output resolution 1K, 2K or 4K. |
| aspect_ratiostringdefault "16:9" | "auto", "1:1", "9:16", "16:9", "3:4", "4:3", "3:2", "2:3", "5:4", "4:5", "21:9"Use one of the listed values. nano-banana-pro does not support auto (it produces 1:1). |
| output_formatstringdefault "PNG" | "PNG", "JPEG"Output image format. |
| durationstringdefault "0" | Send "0" (not used for images). |
Rules
- aspect_ratio: use one of the listed values; with nano-banana-pro, auto produces 1:1.
- Text-to-image does not use images; send images: [] or omit it.
- Asynchronous: the response is {generation_id, status:'processing', image_url:null, video_url:null, video_quality}. Poll GET /api/video-generation/{generation_id}/status until status='completed'; the first image URL is then in video_url and all output images are in result_images (array). The status response has no image_url field.
Example request body
{
"mode": "nano-banana",
"generation_type": "text-to-image",
"prompt": "A beautiful sunset over mountains with golden light",
"duration": "0",
"aspect_ratio": "auto",
"video_quality": "standard",
"modelVersion": "nano-banana",
"output_format": "PNG",
"images": []
}Google Nano Banana (Edit) mode: nano-banana-edit
image-to-image
Edits one or more uploaded images from a text instruction with Google Nano Banana (or Nano Banana Pro).
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt, up to 20000 characters. |
| imagesstring[] · required | Input image URLs to edit (1-10). |
| modelVersionstringaffects pricedefault "nano-banana" | "nano-banana", "nano-banana-pro"nano-banana = base model; nano-banana-pro = Nano Banana Pro. |
| video_qualitystringaffects pricedefault "standard" | "standard", "1K", "2K", "4K"Send 'standard' with nano-banana; with nano-banana-pro, send the output resolution 1K, 2K or 4K. |
| aspect_ratiostringdefault "16:9" | "auto", "1:1", "9:16", "16:9", "3:4", "4:3", "3:2", "2:3", "5:4", "4:5", "21:9"Use one of the listed values. nano-banana-pro does not support auto (it produces 1:1). |
| output_formatstringdefault "PNG" | "PNG", "JPEG"Output image format. |
| durationstringdefault "0" | Send "0" (not used for images). |
Input files
images: 1-10 image URLs; public https URLs or files uploaded with POST /api/storage/upload
Upload folders for /api/storage/upload
- image:
image-generator
Rules
- aspect_ratio: use one of the listed values; with nano-banana-pro, auto produces 1:1.
- image-to-image requires a non-empty images array (or image) (400 otherwise).
- Send 1-10 input images.
- Asynchronous: the response is {generation_id, status:'processing', image_url:null, video_url:null, video_quality}. Poll GET /api/video-generation/{generation_id}/status until status='completed'; the first image URL is then in video_url and all output images are in result_images (array). The status response has no image_url field.
Example request body
{
"mode": "nano-banana-edit",
"generation_type": "image-to-image",
"prompt": "Change the background to a blue sky",
"duration": "0",
"aspect_ratio": "auto",
"video_quality": "standard",
"modelVersion": "nano-banana",
"output_format": "PNG",
"images": [
"https://example.com/input.jpg"
]
}Google Nano Banana 2.1 mode: nano-banana-2-1
text-to-image
Generates a 1K/2K/4K image from text with Google Nano Banana 2.1.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt, up to 5000 characters. |
| resolutionstringaffects pricedefault "1K" | "1K", "2K", "4K"Output tier. |
| aspect_ratiostringdefault "16:9" | "auto", "1:1", "16:9", "9:16", "4:3", "3:4", "3:2", "2:3", "5:4", "4:5", "21:9", "4:1", "1:4", "8:1", "1:8"auto lets the model pick from the prompt. Always send it (the API default is 16:9). |
| output_formatstringdefault "JPEG" | "JPEG", "PNG"Output format: JPEG or PNG (default JPEG). |
| durationstringdefault "0" | Send "0" (not used for images). |
Rules
- Price by resolution only: 1K = 3 credits, 2K = 5, 4K = 8. Aspect ratio and the number of reference images do not change the price.
- generation_type must be text-to-image or image-to-image (422).
- aspect_ratio must be one of the listed values (400).
- resolution must be 1K, 2K or 4K (422).
- prompt max 5000 characters (400).
- This is a different model from nano-banana-2 (separate mode and price).
- Asynchronous: the response is {generation_id, status:'processing', image_url:null, video_url:null, video_quality}. Poll GET /api/video-generation/{generation_id}/status until status='completed'; the first image URL is then in video_url and all output images are in result_images (array). The status response has no image_url field.
Example request body
{
"mode": "nano-banana-2-1",
"generation_type": "text-to-image",
"prompt": "A minimalist coffee shop poster with the headline \"SLOW MORNINGS\" in bold serif type",
"duration": "0",
"aspect_ratio": "1:1",
"resolution": "1K",
"output_format": "JPEG",
"images": []
}image-to-image
Edits/combines up to 10 reference images from a prompt with Google Nano Banana 2.1 at 1K-4K.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt, up to 5000 characters. |
| imagesstring[] · required | 1-10 reference image URLs. |
| resolutionstringaffects pricedefault "1K" | "1K", "2K", "4K"Output tier. |
| aspect_ratiostringdefault "16:9" | "auto", "1:1", "16:9", "9:16", "4:3", "3:4", "3:2", "2:3", "5:4", "4:5", "21:9", "4:1", "1:4", "8:1", "1:8"auto keeps the aspect ratio of the first reference image. Always send it (the API default is 16:9). |
| output_formatstringdefault "JPEG" | "JPEG", "PNG"Output format: JPEG or PNG (default JPEG). |
| durationstringdefault "0" | Send "0" (not used for images). |
Input files
images: 1-10 reference image URLs
Upload folders for /api/storage/upload
- image:
image-generator
Rules
- Price by resolution only: 1K = 3 credits, 2K = 5, 4K = 8. Aspect ratio and the number of reference images do not change the price.
- generation_type must be text-to-image or image-to-image (422).
- aspect_ratio must be one of the listed values (400).
- resolution must be 1K, 2K or 4K (422).
- prompt max 5000 characters (400).
- image-to-image requires a non-empty images array (or image) (400 otherwise).
- Max 10 reference images (400).
- This is a different model from nano-banana-2 (separate mode and price).
- Asynchronous: the response is {generation_id, status:'processing', image_url:null, video_url:null, video_quality}. Poll GET /api/video-generation/{generation_id}/status until status='completed'; the first image URL is then in video_url and all output images are in result_images (array). The status response has no image_url field.
Example request body
{
"mode": "nano-banana-2-1",
"generation_type": "image-to-image",
"prompt": "A minimalist coffee shop poster with the headline \"SLOW MORNINGS\" in bold serif type",
"duration": "0",
"aspect_ratio": "auto",
"resolution": "1K",
"output_format": "JPEG",
"images": [
"https://example.com/ref.jpg"
]
}Google Nano Banana 2 mode: nano-banana-2
text-to-image
Generates an image from text with Google Nano Banana 2 (1K-4K).
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt, up to 20000 characters. |
| video_qualitystringaffects pricedefault "1K" | "1K", "2K", "4K" |
| aspect_ratiostringdefault "16:9" | "1:1", "2:3", "3:2", "3:4", "4:3", "4:5", "5:4", "9:16", "16:9", "21:9", "1:4", "4:1", "1:8", "8:1"Output aspect ratio. Use one of the listed values. |
| output_formatstringdefault "JPG" | "JPG", "PNG"Output format: JPG or PNG (default JPG). |
| google_searchbooleandefault false | Enable Google Search grounding. |
| image_searchbooleandefault false | Optional; has no effect for this model and can be omitted. |
| durationstringdefault "0" | Send "0" (not used for images). |
Rules
- Use one of the listed values.
- aspect_ratio: use one of the listed values.
- Asynchronous: the response is {generation_id, status:'processing', image_url:null, video_url:null, video_quality}. Poll GET /api/video-generation/{generation_id}/status until status='completed'; the first image URL is then in video_url and all output images are in result_images (array). The status response has no image_url field.
Example request body
{
"mode": "nano-banana-2",
"generation_type": "text-to-image",
"prompt": "A beautiful portrait with stunning details",
"duration": "0",
"aspect_ratio": "1:1",
"video_quality": "1K",
"output_format": "JPG",
"google_search": false,
"image_search": false,
"images": []
}image-to-image
Edits/combines reference images from a text prompt with Google Nano Banana 2.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt, up to 20000 characters. |
| imagesstring[] · required | Reference image URLs (1-14). |
| video_qualitystringaffects pricedefault "1K" | "1K", "2K", "4K" |
| aspect_ratiostringdefault "16:9" | "match_input_image", "1:1", "2:3", "3:2", "3:4", "4:3", "4:5", "5:4", "9:16", "16:9", "21:9", "1:4", "4:1", "1:8", "8:1"Output aspect ratio. 'match_input_image' keeps the input image's ratio. Use one of the listed values. |
| output_formatstringdefault "JPG" | "JPG", "PNG"Output format: JPG or PNG (default JPG). |
| google_searchbooleandefault false | Enable Google Search grounding. |
| image_searchbooleandefault false | Optional; has no effect for this model and can be omitted. |
| durationstringdefault "0" | Send "0" (not used for images). |
Input files
images: 1-14 reference image URLs
Upload folders for /api/storage/upload
- image:
image-generator
Rules
- Use one of the listed values.
- aspect_ratio: use one of the listed values.
- image-to-image requires a non-empty images array (or image) (400 otherwise).
- Send 1-14 reference images.
- Asynchronous: the response is {generation_id, status:'processing', image_url:null, video_url:null, video_quality}. Poll GET /api/video-generation/{generation_id}/status until status='completed'; the first image URL is then in video_url and all output images are in result_images (array). The status response has no image_url field.
Example request body
{
"mode": "nano-banana-2",
"generation_type": "image-to-image",
"prompt": "A beautiful portrait with stunning details",
"duration": "0",
"aspect_ratio": "match_input_image",
"video_quality": "1K",
"output_format": "JPG",
"google_search": false,
"image_search": false,
"images": [
"https://example.com/ref.jpg"
]
}Flux Kontext mode: flux-kontext
text-to-image
Generates an image from text with Flux Kontext (Flux 1 or Flux 2, Pro or Max).
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt. |
| fluxModelVersionstringaffects pricedefault "flux1" | "flux1", "flux2"Flux Kontext 1 (flux1) or Flux 2 (flux2). |
| fluxModelstringaffects pricedefault "flux-kontext-pro" | "flux-kontext-pro", "flux-kontext-max"Pro or Max (with flux2, Max is Flux 2 Flex). |
| fluxResolutionstringaffects pricedefault "1K" | "1K", "2K"Output resolution; affects price and output only with flux2. |
| video_qualitystringdefault "1K" | "1K", "2K"Send the same value as fluxResolution. |
| aspect_ratiostringdefault "16:9" | "16:9", "21:9", "4:3", "1:1", "3:4", "9:16", "16:21", "3:2", "2:3", "auto"With flux1: 16:9, 21:9, 4:3, 1:1, 3:4, 9:16, 16:21. With flux2: 1:1, 4:3, 3:4, 16:9, 9:16, 3:2, 2:3, auto. |
| output_formatstring | "PNG", "JPEG"Output format, PNG or JPEG (flux1 only). When omitted the model picks the format. |
| promptUpsamplingbooleandefault false | Prompt upsampling (flux1 only). |
| enableTranslationbooleandefault false | Prompt translation (flux1 only). |
| durationstringdefault "0" | Send "0" (not used for images). |
Rules
- aspect_ratio: use a value supported by the selected fluxModelVersion.
- Asynchronous: the response is {generation_id, status:'processing', image_url:null, video_url:null, video_quality}. Poll GET /api/video-generation/{generation_id}/status until status='completed'; the first image URL is then in video_url and all output images are in result_images (array). The status response has no image_url field.
Example request body
{
"mode": "flux-kontext",
"generation_type": "text-to-image",
"prompt": "Serene mountain landscape at sunset with lake reflection",
"duration": "0",
"aspect_ratio": "16:9",
"video_quality": "1K",
"fluxModel": "flux-kontext-pro",
"fluxModelVersion": "flux1",
"fluxResolution": "1K",
"promptUpsampling": false,
"enableTranslation": true,
"output_format": "PNG"
}image-to-image
Edits an image (or up to 8 references with Flux 2) from a text instruction with Flux Kontext.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt. |
| imagestring | Input image URL (first image); flux1 uses only this. Required unless you send it as images[0]. |
| imagesstring[] | flux2 only: all reference image URLs, up to 8. |
| fluxModelVersionstringaffects pricedefault "flux1" | "flux1", "flux2"Flux Kontext 1 (flux1) or Flux 2 (flux2). |
| fluxModelstringaffects pricedefault "flux-kontext-pro" | "flux-kontext-pro", "flux-kontext-max"Pro or Max (with flux2, Max is Flux 2 Flex). |
| fluxResolutionstringaffects pricedefault "1K" | "1K", "2K"Output resolution; affects price and output only with flux2. |
| video_qualitystringdefault "1K" | "1K", "2K"Send the same value as fluxResolution. |
| aspect_ratiostringdefault "16:9" | "16:9", "21:9", "4:3", "1:1", "3:4", "9:16", "16:21", "3:2", "2:3", "auto"With flux1: 16:9, 21:9, 4:3, 1:1, 3:4, 9:16, 16:21. With flux2: 1:1, 4:3, 3:4, 16:9, 9:16, 3:2, 2:3, auto. |
| output_formatstring | "PNG", "JPEG"Output format, PNG or JPEG (flux1 only). When omitted the model picks the format. |
| promptUpsamplingbooleandefault false | Prompt upsampling (flux1 only). |
| enableTranslationbooleandefault false | Prompt translation (flux1 only). |
| durationstringdefault "0" | Send "0" (not used for images). |
Input files
image: 1 image URL (flux1)images: 1-8 image URLs (flux2 only)
Upload folders for /api/storage/upload
- image:
flux-kontext
Rules
- aspect_ratio: use a value supported by the selected fluxModelVersion.
- image-to-image requires a non-empty images array (or image) (400 otherwise).
- flux1 uses only the single input image in image; flux2 accepts up to 8 reference images in images.
- Asynchronous: the response is {generation_id, status:'processing', image_url:null, video_url:null, video_quality}. Poll GET /api/video-generation/{generation_id}/status until status='completed'; the first image URL is then in video_url and all output images are in result_images (array). The status response has no image_url field.
Example request body
{
"mode": "flux-kontext",
"generation_type": "image-to-image",
"prompt": "Serene mountain landscape at sunset with lake reflection",
"duration": "0",
"aspect_ratio": "16:9",
"video_quality": "1K",
"fluxModel": "flux-kontext-pro",
"fluxModelVersion": "flux1",
"fluxResolution": "1K",
"promptUpsampling": false,
"enableTranslation": true,
"output_format": "PNG",
"image": "https://example.com/input.jpg"
}FLUX 3 Image mode: flux-3-image
text-to-image
Generates a native 1K/2K/4K image from text with FLUX 3.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt, up to 5000 characters. |
| resolutionstringaffects pricedefault "1K" | "1K", "2K", "4K"Output tier. |
| aspect_ratiostringaffects pricedefault "16:9" | "21:9", "2:1", "16:9", "3:2", "7:5", "4:3", "5:4", "1:1", "4:5", "3:4", "5:7", "2:3", "9:16", "1:2"Output aspect ratio (no auto). Always send it. |
| output_formatstringdefault "JPEG" | "JPEG", "PNG"Output format: JPEG or PNG (default JPEG). |
| durationstringdefault "0" | Send "0" (not used for images). |
Rules
- generation_type must be text-to-image or image-to-image (422).
- aspect_ratio must be one of the 14 listed values (400). Always send aspect_ratio.
- resolution must be 1K, 2K or 4K (422).
- prompt max 5000 characters (400).
- Asynchronous: the response is {generation_id, status:'processing', image_url:null, video_url:null, video_quality}. Poll GET /api/video-generation/{generation_id}/status until status='completed'; the first image URL is then in video_url and all output images are in result_images (array). The status response has no image_url field.
Example request body
{
"mode": "flux-3-image",
"generation_type": "text-to-image",
"prompt": "A bakery storefront at golden hour with a hand-painted sign that reads \"Morning Loaf\"",
"duration": "0",
"aspect_ratio": "1:1",
"resolution": "1K",
"output_format": "JPEG",
"images": []
}image-to-image
Edits/combines up to 10 reference images from a prompt with FLUX 3 at 1K-4K.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt, up to 5000 characters. |
| imagesstring[] · required | 1-10 reference image URLs. |
| resolutionstringaffects pricedefault "1K" | "1K", "2K", "4K"Output tier. |
| aspect_ratiostringaffects pricedefault "16:9" | "21:9", "2:1", "16:9", "3:2", "7:5", "4:3", "5:4", "1:1", "4:5", "3:4", "5:7", "2:3", "9:16", "1:2"Output aspect ratio (no auto). Always send it. |
| output_formatstringdefault "JPEG" | "JPEG", "PNG"Output format: JPEG or PNG (default JPEG). |
| durationstringdefault "0" | Send "0" (not used for images). |
Input files
images: 1-10 reference image URLs
Upload folders for /api/storage/upload
- image:
image-generator
Rules
- generation_type must be text-to-image or image-to-image (422).
- aspect_ratio must be one of the 14 listed values (400). Always send aspect_ratio.
- resolution must be 1K, 2K or 4K (422).
- prompt max 5000 characters (400).
- image-to-image requires a non-empty images array (or image) (400 otherwise).
- Max 10 reference images (400); references do not change the price.
- Asynchronous: the response is {generation_id, status:'processing', image_url:null, video_url:null, video_quality}. Poll GET /api/video-generation/{generation_id}/status until status='completed'; the first image URL is then in video_url and all output images are in result_images (array). The status response has no image_url field.
Example request body
{
"mode": "flux-3-image",
"generation_type": "image-to-image",
"prompt": "A bakery storefront at golden hour with a hand-painted sign that reads \"Morning Loaf\"",
"duration": "0",
"aspect_ratio": "1:1",
"resolution": "1K",
"output_format": "JPEG",
"images": [
"https://example.com/ref.jpg"
]
}Seedream 5.0 mode: seedream-5.0-lite
text-to-image
Generates HD images (optionally a batch) from text with Seedream 5.0 Lite.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt. |
| video_qualitystringdefault "2K" | "2K", "3K"Output size: 2K or 3K. |
| aspect_ratiostringdefault "16:9" | "1:1", "2:3", "3:2", "3:4", "4:3", "9:16", "16:9", "21:9"Output aspect ratio. Use one of the listed values and always send it. |
| output_formatstringdefault "PNG" | "PNG", "JPEG"Output format: PNG or JPEG. |
| sequential_image_generationstringdefault "disabled" | "disabled", "auto"'auto' enables batch (multi-image) output. |
| max_imagesnumberaffects pricedefault 1 | 1–15Number of images in a batch. Set it above 1 only together with sequential_image_generation='auto'. |
| durationstringdefault "0" | Send "0" (not used for images). |
Rules
- Set max_images > 1 only with sequential_image_generation='auto'.
- If a batch returns fewer images than requested, unused credits are refunded automatically.
- Always send aspect_ratio, using one of the listed values.
- Asynchronous: the response is {generation_id, status:'processing', image_url:null, video_url:null, video_quality}. Poll GET /api/video-generation/{generation_id}/status until status='completed'; the first image URL is then in video_url and all output images are in result_images (array). The status response has no image_url field. Batch results appear in result_images.
Example request body
{
"mode": "seedream-5.0-lite",
"generation_type": "text-to-image",
"prompt": "A cozy reading nook with warm light",
"duration": "0",
"aspect_ratio": "1:1",
"video_quality": "2K",
"output_format": "PNG",
"sequential_image_generation": "disabled",
"max_images": 1,
"images": []
}image-to-image
Edits/blends reference images into new HD images (optionally a batch) with Seedream 5.0 Lite.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt. |
| imagesstring[] · required | Reference image URLs (1-14). |
| video_qualitystringdefault "2K" | "2K", "3K"Output size: 2K or 3K. |
| aspect_ratiostringdefault "16:9" | "match_input_image", "1:1", "2:3", "3:2", "3:4", "4:3", "9:16", "16:9", "21:9"Output aspect ratio; 'match_input_image' keeps the input image's ratio. Use one of the listed values and always send it. |
| output_formatstringdefault "PNG" | "PNG", "JPEG"Output format: PNG or JPEG. |
| sequential_image_generationstringdefault "disabled" | "disabled", "auto"'auto' enables batch (multi-image) output. |
| max_imagesnumberaffects pricedefault 1 | 1–15Number of images in a batch. Set it above 1 only together with sequential_image_generation='auto'. |
| durationstringdefault "0" | Send "0" (not used for images). |
Input files
images: 1-14 reference image URLs
Upload folders for /api/storage/upload
- image:
image-generator
Rules
- Set max_images > 1 only with sequential_image_generation='auto'.
- If a batch returns fewer images than requested, unused credits are refunded automatically.
- Always send aspect_ratio, using one of the listed values.
- image-to-image requires a non-empty images array (or image) (400 otherwise).
- Send 1-14 reference images.
- Asynchronous: the response is {generation_id, status:'processing', image_url:null, video_url:null, video_quality}. Poll GET /api/video-generation/{generation_id}/status until status='completed'; the first image URL is then in video_url and all output images are in result_images (array). The status response has no image_url field. Batch results appear in result_images.
Example request body
{
"mode": "seedream-5.0-lite",
"generation_type": "image-to-image",
"prompt": "A cozy reading nook with warm light",
"duration": "0",
"aspect_ratio": "match_input_image",
"video_quality": "2K",
"output_format": "PNG",
"sequential_image_generation": "disabled",
"max_images": 1,
"images": [
"https://example.com/ref.jpg"
]
}Seedream 5.0 Pro mode: seedream-5.0-pro
text-to-image
Generates images with dense text and infographics from a prompt with Seedream 5.0 Pro.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt, up to 5000 characters. |
| resolutionstringaffects pricedefault "1K" | "1K", "2K" |
| video_qualitystringdefault "1K" | "1K", "2K"Send the same value as resolution; resolution determines price and output. |
| aspect_ratiostringdefault "16:9" | "1:1", "4:3", "3:4", "16:9", "9:16", "2:3", "3:2", "21:9"Output aspect ratio; no auto (also for image-to-image). Always send it. |
| output_formatstring | "PNG", "JPEG"Output format: PNG or JPEG. When omitted the model picks the format. |
| durationstringdefault "0" | Send "0" (not used for images). |
Rules
- aspect_ratio must be one of 1:1, 4:3, 3:4, 16:9, 9:16, 2:3, 3:2, 21:9 (400). Always send aspect_ratio.
- resolution must be 1K or 2K (422).
- prompt max 5000 characters (400).
- Asynchronous: the response is {generation_id, status:'processing', image_url:null, video_url:null, video_quality}. Poll GET /api/video-generation/{generation_id}/status until status='completed'; the first image URL is then in video_url and all output images are in result_images (array). The status response has no image_url field.
Example request body
{
"mode": "seedream-5.0-pro",
"generation_type": "text-to-image",
"prompt": "An infographic poster explaining the water cycle with labeled arrows",
"duration": "0",
"aspect_ratio": "1:1",
"video_quality": "1K",
"resolution": "1K",
"output_format": "PNG",
"images": []
}image-to-image
Precision-edits up to 10 reference images from a prompt with Seedream 5.0 Pro.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt, up to 5000 characters. |
| imagesstring[] · requiredaffects price | 1-10 reference image URLs. |
| resolutionstringaffects pricedefault "1K" | "1K", "2K" |
| video_qualitystringdefault "1K" | "1K", "2K"Send the same value as resolution; resolution determines price and output. |
| aspect_ratiostringdefault "16:9" | "1:1", "4:3", "3:4", "16:9", "9:16", "2:3", "3:2", "21:9"Output aspect ratio; no auto (also for image-to-image). Always send it. |
| output_formatstring | "PNG", "JPEG"Output format: PNG or JPEG. When omitted the model picks the format. |
| durationstringdefault "0" | Send "0" (not used for images). |
Input files
images: 1-10 reference image URLs
Upload folders for /api/storage/upload
- image:
image-generator
Rules
- aspect_ratio must be one of 1:1, 4:3, 3:4, 16:9, 9:16, 2:3, 3:2, 21:9 (400). Always send aspect_ratio.
- resolution must be 1K or 2K (422).
- prompt max 5000 characters (400).
- image-to-image requires a non-empty images array (or image) (400 otherwise).
- Max 10 reference images (400).
- Asynchronous: the response is {generation_id, status:'processing', image_url:null, video_url:null, video_quality}. Poll GET /api/video-generation/{generation_id}/status until status='completed'; the first image URL is then in video_url and all output images are in result_images (array). The status response has no image_url field.
- When quoting with
/api/quote, sendinput_image_count(the number of images) — the quote does not count theimagesarray for this model, so it would come out too low.
Example request body
{
"mode": "seedream-5.0-pro",
"generation_type": "image-to-image",
"prompt": "An infographic poster explaining the water cycle with labeled arrows",
"duration": "0",
"aspect_ratio": "1:1",
"video_quality": "1K",
"resolution": "1K",
"output_format": "PNG",
"images": [
"https://example.com/ref.jpg"
]
}Seedream 5.0 Flash mode: seedream-5.0-flash
text-to-image
Fast, low-cost text-to-image up to 2K with Seedream 5.0 Flash.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt, max 5000 characters. |
| resolutionstringdefault "1K" | "1K", "1.5K", "2K"Output size (1K, 1.5K or 2K). Set it with resolution; video_quality is not used. |
| aspect_ratiostringdefault "16:9" | "1:1", "4:3", "3:4", "16:9", "9:16", "2:3", "3:2", "21:9"Output aspect ratio; no auto option. Always send it explicitly. |
| output_formatstringdefault "PNG" | "PNG", "JPEG"Output format in uppercase: PNG or JPEG. |
| durationstringdefault "0" | Send "0". |
Rules
- generation_type must be text-to-image or image-to-image (422).
- aspect_ratio must be one of 1:1, 4:3, 3:4, 16:9, 9:16, 2:3, 3:2, 21:9 (400). Always send
aspect_ratioexplicitly. - resolution must be 1K, 1.5K or 2K (422).
- prompt max 5000 characters (400).
- Asynchronous: POST returns {generation_id, status:'processing'}. Poll GET /api/video-generation/{generation_id}/status until status='completed'; the first image URL is in video_url and all output images are in result_images (array). The status response has no image_url field.
Example request body
{
"mode": "seedream-5.0-flash",
"generation_type": "text-to-image",
"prompt": "A cozy coffee shop interior at golden hour, warm light through large windows",
"duration": "0",
"aspect_ratio": "1:1",
"resolution": "1K",
"output_format": "PNG",
"images": []
}image-to-image
Fast, low-cost edits of up to 10 reference images up to 2K with Seedream 5.0 Flash.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Text prompt, max 5000 characters. |
| imagesstring[] · required | 1-10 reference image URLs. |
| resolutionstringdefault "1K" | "1K", "1.5K", "2K"Output size (1K, 1.5K or 2K). Set it with resolution; video_quality is not used. |
| aspect_ratiostringdefault "16:9" | "1:1", "4:3", "3:4", "16:9", "9:16", "2:3", "3:2", "21:9"Output aspect ratio; no auto option. Always send it explicitly. |
| output_formatstringdefault "PNG" | "PNG", "JPEG"Output format in uppercase: PNG or JPEG. |
| durationstringdefault "0" | Send "0". |
Input files
images: 1-10 reference image URLs.
Upload folders for /api/storage/upload
- image:
image-generator
Rules
- generation_type must be text-to-image or image-to-image (422).
- aspect_ratio must be one of 1:1, 4:3, 3:4, 16:9, 9:16, 2:3, 3:2, 21:9 (400). Always send
aspect_ratioexplicitly. - resolution must be 1K, 1.5K or 2K (422).
- prompt max 5000 characters (400).
- image-to-image requires a non-empty images array (or image) (400).
- At most 10 reference images (400).
- Asynchronous: POST returns {generation_id, status:'processing'}. Poll GET /api/video-generation/{generation_id}/status until status='completed'; the first image URL is in video_url and all output images are in result_images (array). The status response has no image_url field.
Example request body
{
"mode": "seedream-5.0-flash",
"generation_type": "image-to-image",
"prompt": "A cozy coffee shop interior at golden hour, warm light through large windows",
"duration": "0",
"aspect_ratio": "1:1",
"resolution": "1K",
"output_format": "PNG",
"images": [
"https://example.com/ref.jpg"
]
}Ideogram 4.5 mode: ideogram-4-5
text-to-image
Generates 1-8 images with accurate in-image text from a prompt with Ideogram 4.5.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Required, non-empty, max 10000 characters. |
| qualitystringaffects pricedefault "low" | "low", "medium", "high" |
| num_imagesnumberaffects pricedefault 1 | 1–8Images per request (1-8); the price is multiplied by num_images. |
| enable_prompt_expansionbooleandefault true | Text-to-image only. |
| aspect_ratiostringdefault "1:1" | "1:1", "16:9", "9:16", "4:3", "3:4", "3:2", "2:3", "5:4", "4:5", "16:10", "10:16", "2:1", "1:2", "12:5", "5:12", "22:9", "9:22", "23:9", "9:23", "8:3", "3:8", "3:1", "1:3"Output aspect ratio. |
| resolutionstring | "1K", "2K"Resolution tier for the chosen aspect_ratio; defaults to the lowest available. 12:5, 5:12, 22:9, 9:22, 23:9, 9:23, 8:3, 3:8, 3:1, 1:3 are 2K only. |
| seednumber | Optional integer seed. |
| durationstringdefault "0" | Send "0". |
Rules
- Credits for images that are not delivered are refunded automatically.
- generation_type must be text-to-image or image-to-image.
- prompt required, max 10000 characters (422).
- num_images must be an integer 1-8 (422).
- resolution must be valid for the chosen aspect_ratio (422).
- text-to-image must not include images (422); edit_precision and mask_url are rejected (422).
- aspect_ratio auto is not allowed for text-to-image.
- Asynchronous: POST returns {generation_id, status:'processing'}. Poll GET /api/video-generation/{generation_id}/status until status='completed'; the first image URL is in video_url and all output images are in result_images (array). The status response has no image_url field. With num_images > 1, all images are in result_images.
Example request body
{
"mode": "ideogram-4-5",
"generation_type": "text-to-image",
"prompt": "A vintage travel poster with the bold title \"VISIT LISBON\" and a yellow tram",
"duration": "0",
"quality": "low",
"num_images": 1,
"aspect_ratio": "1:1",
"resolution": "1K",
"enable_prompt_expansion": true
}image-to-image
Edits an image (optionally with a mask and reference images) with accurate text rendering via Ideogram 4.5.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Required, non-empty, max 10000 characters. |
| imagesstring[] · required | images[0] = image to edit; images[1..] = up to 4 reference images (3 with a mask). http(s) URLs only. |
| qualitystringaffects pricedefault "very_low" | "very_low", "low", "medium", "high" |
| num_imagesnumberaffects pricedefault 1 | 1–8Images per request (1-8); the price is multiplied by num_images. |
| edit_precisionstringdefault "regular" | "regular", "high"Edit precision. |
| mask_urlstring | Black/white mask URL, same size as the source image; black = area to change. |
| aspect_ratiostringdefault "auto" | "auto", "1:1", "16:9", "9:16", "4:3", "3:4", "3:2", "2:3", "5:4", "4:5", "16:10", "10:16", "2:1", "1:2", "12:5", "5:12", "22:9", "9:22", "23:9", "9:23", "8:3", "3:8", "3:1", "1:3"auto keeps the source image size. |
| resolutionstring | "1K", "2K"Resolution tier for the chosen aspect_ratio; defaults to the lowest available. 12:5, 5:12, 22:9, 9:22, 23:9, 9:23, 8:3, 3:8, 3:1, 1:3 are 2K only. Not allowed with aspect_ratio auto. |
| seednumber | Optional integer seed. |
| durationstringdefault "0" | Send "0". |
Input files
images: 1 image to edit + 0-4 reference image URLs (0-3 with a mask).mask: Optional mask_url (PNG, black = edit area, same size as the source).
Upload folders for /api/storage/upload
- image:
image-generator - mask:
image-generator
Rules
- Credits for images that are not delivered are refunded automatically.
- generation_type must be text-to-image or image-to-image.
- prompt required, max 10000 characters (422).
- num_images must be an integer 1-8 (422).
- resolution must be valid for the chosen aspect_ratio (422).
- very_low quality is available for image-to-image only.
- Needs at least 1 image; at most 4 reference images after the first (3 when mask_url is set). Image URLs must be http(s).
- enable_prompt_expansion is rejected for image-to-image (422).
- A non-auto aspect_ratio/resolution is only allowed with edit_precision=regular and no mask (422).
- mask_url is only accepted for qwen-image-2-1 and ideogram-4-5 and must be an http(s) URL.
- Asynchronous: POST returns {generation_id, status:'processing'}. Poll GET /api/video-generation/{generation_id}/status until status='completed'; the first image URL is in video_url and all output images are in result_images (array). The status response has no image_url field. With num_images > 1, all images are in result_images.
Example request body
{
"mode": "ideogram-4-5",
"generation_type": "image-to-image",
"prompt": "Replace the sign text with \"OPEN\"",
"duration": "0",
"quality": "very_low",
"num_images": 1,
"aspect_ratio": "auto",
"images": [
"https://example.com/source.jpg"
],
"edit_precision": "regular"
}GPT Image 2 mode: gpt-image-2
text-to-image
Generates photoreal images with sharp text rendering from a prompt with GPT Image 2.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Required, non-empty, max 20000 characters. |
| resolutionstringaffects pricedefault "1K" | "1K", "2K", "4K" |
| aspect_ratiostringdefault "auto" | "auto", "1:1", "3:2", "2:3", "4:3", "3:4", "5:4", "4:5", "16:9", "9:16", "2:1", "1:2", "3:1", "1:3", "21:9", "9:21"Omitted = auto. |
| backgroundstringdefault "auto" | "auto", "opaque", "transparent"Values other than auto are only allowed at 1K. |
| durationstringdefault "0" | Send "0". |
Rules
- generation_type must be text-to-image or image-to-image; prompt required, max 20000 characters (422).
- aspect_ratio auto (or omitted) only supports 1K (422).
- 1:1 does not support 4K (422).
- 5:4, 4:5, 3:1, 1:3 and 9:21 are not supported at 2K; 3:1, 1:3 and 9:21 are not supported at 4K (422).
- background other than auto is only allowed at 1K (422).
- Asynchronous: POST returns {generation_id, status:'processing'}. Poll GET /api/video-generation/{generation_id}/status until status='completed'; the first image URL is in video_url and all output images are in result_images (array). The status response has no image_url field.
Example request body
{
"mode": "gpt-image-2",
"generation_type": "text-to-image",
"prompt": "A modern storefront poster with the headline \"Summer Sale\" in bold serif type",
"duration": "0",
"aspect_ratio": "auto",
"resolution": "1K",
"images": []
}image-to-image
Edits/combines up to 16 reference images from a prompt with GPT Image 2.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Required, non-empty, max 20000 characters. |
| imagesstring[] · required | 1-16 reference image URLs. |
| resolutionstringaffects pricedefault "1K" | "1K", "2K", "4K" |
| aspect_ratiostringdefault "auto" | "auto", "1:1", "3:2", "2:3", "4:3", "3:4", "5:4", "4:5", "16:9", "9:16", "2:1", "1:2", "3:1", "1:3", "21:9", "9:21"Omitted = auto. |
| backgroundstringdefault "auto" | "auto", "opaque", "transparent"Values other than auto are only allowed at 1K. |
| durationstringdefault "0" | Send "0". |
Input files
images: 1-16 reference image URLs.
Upload folders for /api/storage/upload
- image:
image-generator
Rules
- generation_type must be text-to-image or image-to-image; prompt required, max 20000 characters (422).
- aspect_ratio auto (or omitted) only supports 1K (422).
- 1:1 does not support 4K (422).
- 5:4, 4:5, 3:1, 1:3 and 9:21 are not supported at 2K; 3:1, 1:3 and 9:21 are not supported at 4K (422).
- background other than auto is only allowed at 1K (422).
- image-to-image requires a non-empty images array (or image) (400).
- At most 16 reference images (400).
- Asynchronous: POST returns {generation_id, status:'processing'}. Poll GET /api/video-generation/{generation_id}/status until status='completed'; the first image URL is in video_url and all output images are in result_images (array). The status response has no image_url field.
Example request body
{
"mode": "gpt-image-2",
"generation_type": "image-to-image",
"prompt": "A modern storefront poster with the headline \"Summer Sale\" in bold serif type",
"duration": "0",
"aspect_ratio": "auto",
"resolution": "1K",
"images": [
"https://example.com/ref.jpg"
]
}GPT Image 2.5 mode: gpt-image-2-5
text-to-image
Generates sharper, detailed images from a prompt with GPT Image 2.5 (Flare/Sunburst).
| Parameter | Values & notes |
|---|---|
| promptstring · required | Required, non-empty, max 20000 characters. |
| gptImage25Versionstringdefault "flare" | "flare", "sunburst"Quality tier; flare and sunburst currently cost the same. |
| resolutionstringaffects pricedefault "1K" | "1K", "2K", "4K" |
| aspect_ratiostringdefault "auto" | "auto", "1:1", "3:2", "2:3", "4:3", "3:4", "16:9", "9:16", "21:9", "27:16", "16:27", "9:8", "8:9"Omitted = auto. 27:16, 16:27, 9:8, 8:9 are 1K only. |
| backgroundstringdefault "auto" | "auto", "opaque", "transparent"Available at all resolutions. |
| durationstringdefault "0" | Send "0". |
Rules
- generation_type must be text-to-image or image-to-image; prompt required, max 20000 characters (422).
- 27:16, 16:27, 9:8, 8:9 are only available at 1K (422).
- Asynchronous: POST returns {generation_id, status:'processing'}. Poll GET /api/video-generation/{generation_id}/status until status='completed'; the first image URL is in video_url and all output images are in result_images (array). The status response has no image_url field.
Example request body
{
"mode": "gpt-image-2-5",
"generation_type": "text-to-image",
"prompt": "A cinematic night city poster with neon reflections on a rainy street",
"duration": "0",
"aspect_ratio": "auto",
"resolution": "1K",
"gptImage25Version": "flare",
"background": "auto",
"images": []
}image-to-image
Edits up to 16 reference images with stronger reference fidelity using GPT Image 2.5.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Required, non-empty, max 20000 characters. |
| imagesstring[] · required | 1-16 reference image URLs. |
| gptImage25Versionstringdefault "flare" | "flare", "sunburst"Quality tier; flare and sunburst currently cost the same. |
| resolutionstringaffects pricedefault "1K" | "1K", "2K", "4K" |
| aspect_ratiostringdefault "auto" | "auto", "1:1", "3:2", "2:3", "4:3", "3:4", "16:9", "9:16", "21:9", "27:16", "16:27", "9:8", "8:9"Omitted = auto. 27:16, 16:27, 9:8, 8:9 are 1K only. |
| backgroundstringdefault "auto" | "auto", "opaque", "transparent"Available at all resolutions. |
| durationstringdefault "0" | Send "0". |
Input files
images: 1-16 reference image URLs.
Upload folders for /api/storage/upload
- image:
image-generator
Rules
- generation_type must be text-to-image or image-to-image; prompt required, max 20000 characters (422).
- 27:16, 16:27, 9:8, 8:9 are only available at 1K (422).
- image-to-image requires a non-empty images array (or image) (400).
- At most 16 reference images (400).
- Asynchronous: POST returns {generation_id, status:'processing'}. Poll GET /api/video-generation/{generation_id}/status until status='completed'; the first image URL is in video_url and all output images are in result_images (array). The status response has no image_url field.
Example request body
{
"mode": "gpt-image-2-5",
"generation_type": "image-to-image",
"prompt": "A cinematic night city poster with neon reflections on a rainy street",
"duration": "0",
"aspect_ratio": "auto",
"resolution": "1K",
"gptImage25Version": "flare",
"background": "auto",
"images": [
"https://example.com/ref.jpg"
]
}Qwen Image 2.1 mode: qwen-image-2-1
text-to-image
Generates images (including native transparent PNGs) from a prompt with Qwen Image 2.1.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Required, non-empty, max 5000 characters. |
| resolutionstringaffects pricedefault "1K" | "1K", "2K" |
| aspect_ratiostringdefault "1:1" | "1:1", "4:3", "3:4", "3:2", "2:3", "16:9", "9:16"Text-to-image default 1:1; auto is not available. |
| backgroundstringdefault "opaque" | "opaque", "transparent"transparent returns a native transparent PNG; 'auto' is rejected. |
| durationstringdefault "0" | Send "0". |
Rules
- generation_type must be text-to-image or image-to-image; prompt required, max 5000 characters (422).
- resolution must be 1K or 2K (422).
- background must be opaque or transparent (422).
- Text-to-image aspect ratios: 1:1, 4:3, 3:4, 3:2, 2:3, 16:9, 9:16 (auto is rejected, 422).
- mask_url is rejected for text-to-image (422).
- Asynchronous: POST returns {generation_id, status:'processing'}. Poll GET /api/video-generation/{generation_id}/status until status='completed'; the first image URL is in video_url and all output images are in result_images (array). The status response has no image_url field.
Example request body
{
"mode": "qwen-image-2-1",
"generation_type": "text-to-image",
"prompt": "A glossy cartoon sticker of a corgi astronaut, transparent background",
"duration": "0",
"aspect_ratio": "1:1",
"resolution": "1K",
"background": "transparent",
"images": []
}image-to-image
Edits up to 10 reference images, or inpaints one image with a mask, using Qwen Image 2.1.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Required, non-empty, max 5000 characters. |
| imagesstring[] · required | 1-10 reference image URLs (exactly 1 with mask_url). |
| resolutionstringaffects pricedefault "1K" | "1K", "2K" |
| aspect_ratiostringdefault "auto" | "auto", "1:1", "4:3", "3:4", "3:2", "2:3", "16:9", "9:16"Image-to-image default auto. When mask_url is set, auto is always used. |
| backgroundstringdefault "opaque" | "opaque", "transparent"transparent returns a native transparent PNG; 'auto' is rejected. |
| mask_urlstring | Black/white mask URL for inpainting; white = area to change. |
| durationstringdefault "0" | Send "0". |
Input files
images: 1-10 reference image URLs (exactly 1 when using a mask).mask: Optional mask_url (black/white PNG, white = edit area).
Upload folders for /api/storage/upload
- image:
image-generator - mask:
image-generator
Rules
- generation_type must be text-to-image or image-to-image; prompt required, max 5000 characters (422).
- resolution must be 1K or 2K (422).
- background must be opaque or transparent (422).
- image-to-image requires a non-empty images array (or image) (400).
- At most 10 reference images (422).
- mask_url requires exactly 1 image and cannot be combined with background=transparent (422); with a mask, aspect_ratio is always auto.
- Image-to-image aspect ratios: auto, 1:1, 4:3, 3:4, 3:2, 2:3, 16:9, 9:16 (422).
- Asynchronous: POST returns {generation_id, status:'processing'}. Poll GET /api/video-generation/{generation_id}/status until status='completed'; the first image URL is in video_url and all output images are in result_images (array). The status response has no image_url field.
Example request body
{
"mode": "qwen-image-2-1",
"generation_type": "image-to-image",
"prompt": "A glossy cartoon sticker of a corgi astronaut, transparent background",
"duration": "0",
"aspect_ratio": "auto",
"resolution": "1K",
"background": "opaque",
"images": [
"https://example.com/ref.jpg"
]
}Virtual Try-On mode: virtual-try-on
image-to-image
Dresses the person in a photo in 1-3 garment images.
| Parameter | Values & notes |
|---|---|
| imagesstring[] · required | [person photo, 1-3 garment images]; images[0] is always the person. |
| promptstring | Optional extra instruction appended to the built-in try-on instruction, max 1000 characters. |
| aspect_ratiostringdefault "auto" | "auto", "1:1", "2:3", "3:2", "3:4", "4:3", "4:5", "5:4", "9:16", "16:9"Omit or send auto to follow the person photo's ratio. |
| durationstringdefault "0" | Send "0". |
Input files
images: images[0] = person photo URL, images[1..3] = 1-3 garment image URLs.
Upload folders for /api/storage/upload
- image:
virtual-try-on
Rules
- generation_type must be image-to-image (422).
- images must be [person, 1-3 garment images] (422).
- prompt max 1000 characters (422).
- aspect_ratio must be auto or one of 1:1, 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9 (422).
- Asynchronous: POST returns {generation_id, status:'processing'}. Poll GET /api/video-generation/{generation_id}/status until status='completed'; the first image URL is in video_url and all output images are in result_images (array). The status response has no image_url field.
Example request body
{
"mode": "virtual-try-on",
"generation_type": "image-to-image",
"prompt": "",
"duration": "0",
"images": [
"https://example.com/person.jpg",
"https://example.com/garment.jpg"
]
}Product Holding mode: product-holding
image-to-image
Places 1-3 product images into the hands of the person in a photo.
| Parameter | Values & notes |
|---|---|
| imagesstring[] · required | [person photo, 1-3 product images]; images[0] is always the person. |
| promptstring | Optional extra instruction appended to the built-in product-holding instruction, max 1000 characters. |
| aspect_ratiostringdefault "auto" | "auto", "1:1", "2:3", "3:2", "3:4", "4:3", "4:5", "5:4", "9:16", "16:9"Omit or send auto to follow the person photo's ratio. |
| durationstringdefault "0" | Send "0". |
Input files
images: images[0] = person photo URL, images[1..3] = 1-3 product image URLs.
Upload folders for /api/storage/upload
- image:
virtual-try-on
Rules
- generation_type must be image-to-image (422).
- images must be [person, 1-3 product images] (422).
- prompt max 1000 characters (422).
- aspect_ratio must be auto or one of 1:1, 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9 (422).
- Asynchronous: POST returns {generation_id, status:'processing'}. Poll GET /api/video-generation/{generation_id}/status until status='completed'; the first image URL is in video_url and all output images are in result_images (array). The status response has no image_url field.
Example request body
{
"mode": "product-holding",
"generation_type": "image-to-image",
"prompt": "",
"duration": "0",
"images": [
"https://example.com/person.jpg",
"https://example.com/product.jpg"
]
}Audio Models
Text-to-speech, music, speech-to-text and audio cleanup. Audio results are returned in video_url with media_kind "audio".
| Model | mode | generation_type |
|---|---|---|
| Grok TTS | grok-tts | text-to-speech |
| Eleven v4 | eleven-v4 | text-to-speech |
| Eleven v4 Turbo | eleven-v4-turbo | text-to-speech |
| Lyria 3.5 (Text to Music) | lyria-music | text-to-music |
| Lyria 3.5 (Image to Music) | lyria-music | text-to-music |
| Scribe V2 | scribe-v2 | speech-to-text |
| VEED Clean Audio | veed-clean-audio | remove-background-noise |
Grok TTS mode: grok-tts
text-to-speech
Converts text to speech with xAI Grok TTS, with selectable voice, language and output codec.
| Parameter | Values & notes |
|---|---|
| promptstring · requiredaffects price | 1–15000Text to speak (sent as-is, never auto-translated). Length is counted in characters. |
| durationstringdefault "0" | "0"Send "0" (audio has no duration input). |
| voicestringdefault "eve" | "carina", "zagan", "helix", "orion", "luna", "iris", "altair", "zenith", "perseus", "helios", "lux", "kepler", "rigel", "cosmo", "celeste", "ursa", "sirius", "lumen", "castor", "naksh", "atlas", "aurora", "liora", "ara", "eve", "leo", "rex", "sal"Voice name. |
| languagestringdefault "auto" | "auto", "en", "zh", "ja", "ko", "fr", "de", "it", "es-ES", "es-MX", "pt-BR", "pt-PT", "ru", "tr", "vi", "id", "hi", "bn", "ar-EG", "ar-SA", "ar-AE"BCP-47 language; 'auto' = auto-detect. |
| audio_codecstringdefault "mp3" | "mp3", "wav", "pcm", "mulaw", "alaw"Output format. pcm/mulaw/alaw are headerless raw data (download only, not playable in a browser). |
| sample_ratenumberdefault 24000 | 8000, 16000, 22050, 24000, 44100, 48000Output sample rate in Hz. |
| bit_ratenumberdefault 128000 | 32000, 64000, 96000, 128000, 192000Only applies when audio_codec is mp3. The maximum depends on sample_rate: up to 8000 Hz → 64000, up to 24000 Hz → 160000 (so 128000 is the highest allowed value), above → 192000. |
| seednumber | Accepted for backward compatibility; can be omitted. |
Rules
- prompt is required and must be non-empty after trim; max 15000 characters.
- generation_type must be 'text-to-speech'; text-to-speech only accepts modes eleven-v4-turbo, eleven-v4, grok-tts.
- voice, language, audio_codec, sample_rate and bit_rate must be one of the listed values when provided, else 400.
- bit_rate above the max for the sample_rate (sample_rate defaults to 24000 when omitted) is rejected with 400.
- Eleven-only fields (stability, similarity_boost, audio_output_format, text_normalization, timestamps=true) are rejected with 400.
Result
Async: POST returns {generation_id, status:'processing'}. Poll GET /api/video-generation/{id}/status; when status='completed', video_url is the generated audio file, media_kind is 'audio', and video_metadata gives content_type/file_name/file_size. No lyrics/transcript/timestamps.
Example request body
{
"mode": "grok-tts",
"generation_type": "text-to-speech",
"prompt": "Welcome to Veevid. Let's make something amazing today.",
"duration": "0",
"voice": "eve",
"language": "auto",
"audio_codec": "mp3",
"sample_rate": 24000,
"bit_rate": 128000
}Eleven v4 mode: eleven-v4
text-to-speech
Converts text to speech with ElevenLabs Eleven v4, the higher-quality tier.
| Parameter | Values & notes |
|---|---|
| promptstring · requiredaffects price | 1–5000Text to speak (not auto-translated). Max 5000 characters. |
| durationstringdefault "0" | "0"Send "0". |
| voicestringdefault "Rachel" | "Rachel", "Aria", "Roger", "Sarah", "Laura", "Charlie", "George", "Callum", "River", "Liam", "Charlotte", "Alice", "Matilda", "Will", "Jessica", "Eric", "Chris", "Brian", "Daniel", "Lily", "Bill"Preset voice name (case-sensitive). |
| languagestringdefault "auto" | "auto", "en", "zh", "yue", "es", "pt", "fr", "de", "it", "ja", "ko", "hi", "ar", "ru", "af", "as", "ast", "az", "be", "bg", "bn", "bs", "ca", "ceb", "cs", "cy", "da", "el", "et", "fa", "fi", "fil", "gl", "gu", "ha", "he", "hr", "hu", "hy", "id", "is", "jv", "ka", "kk", "kn", "ky", "lb", "ln", "lt", "lv", "mi", "mk", "ml", "mn", "mr", "ms", "mt", "my", "ne", "nl", "no", "oc", "or", "pa", "pl", "ps", "ro", "sd", "sk", "sl", "so", "sr", "sv", "sw", "ta", "te", "tg", "th", "tr", "uk", "ur", "uz", "vi"ISO 639-1 code (a few 3-letter codes: yue, ast, ceb, fil); 'auto' = let the model detect. |
| stabilitynumberdefault 0.5 | 0–1Lower = more expressive, higher = more stable. |
| similarity_boostnumberdefault 0.75 | 0–1Higher = closer to the original voice. |
| seednumber | 0–4294967295Integer seed; omit for a random seed. |
| audio_output_formatstringdefault "mp3_44100_128" | "mp3_22050_32", "mp3_44100_32", "mp3_44100_64", "mp3_44100_96", "mp3_44100_128", "mp3_44100_192", "opus_48000_32", "opus_48000_64", "opus_48000_96", "opus_48000_128", "opus_48000_192", "pcm_8000", "pcm_16000", "pcm_22050", "pcm_24000", "pcm_44100", "pcm_48000", "ulaw_8000", "alaw_8000"codec_samplerate_bitrate. opus -> .ogg; pcm/ulaw/alaw are headerless raw files (.pcm/.ulaw/.alaw). |
| text_normalizationstringdefault "auto" | "auto", "on", "off"Whether numbers/dates/abbreviations are expanded before synthesis. |
| timestampsbooleandefault false | Return per-character timestamps in the status response. |
Rules
- prompt is required and non-empty after trim; max 5000 characters.
- generation_type must be 'text-to-speech'.
- Grok-only fields audio_codec / sample_rate / bit_rate are rejected with 400 (use audio_output_format).
- voice, language, audio_output_format and text_normalization must be one of the listed values when provided, else 400.
- stability and similarity_boost must be finite numbers in [0,1]; seed must be an integer in [0, 4294967295].
Result
Async: poll GET /api/video-generation/{id}/status. On completed, video_url is the generated audio file and media_kind is 'audio'. If timestamps:true was sent, the response also has 'timestamps': an array of word groups like {characters:[...], character_start_times_seconds:[...], character_end_times_seconds:[...]} (only present when non-empty).
Example request body
{
"mode": "eleven-v4",
"generation_type": "text-to-speech",
"prompt": "Welcome to Veevid. Let's make something amazing today.",
"duration": "0",
"voice": "Rachel",
"language": "auto",
"stability": 0.5,
"similarity_boost": 0.75,
"audio_output_format": "mp3_44100_128",
"text_normalization": "auto",
"timestamps": false
}Eleven v4 Turbo mode: eleven-v4-turbo
text-to-speech
Converts text to speech with ElevenLabs Eleven v4 Turbo: half the price of Eleven v4, with lower latency.
| Parameter | Values & notes |
|---|---|
| promptstring · requiredaffects price | 1–5000Text to speak (not auto-translated). Max 5000 characters. |
| durationstringdefault "0" | "0"Send "0". |
| voicestringdefault "Rachel" | "Rachel", "Aria", "Roger", "Sarah", "Laura", "Charlie", "George", "Callum", "River", "Liam", "Charlotte", "Alice", "Matilda", "Will", "Jessica", "Eric", "Chris", "Brian", "Daniel", "Lily", "Bill"Preset voice name (case-sensitive). |
| languagestringdefault "auto" | "auto", "en", "zh", "yue", "es", "pt", "fr", "de", "it", "ja", "ko", "hi", "ar", "ru", "af", "as", "ast", "az", "be", "bg", "bn", "bs", "ca", "ceb", "cs", "cy", "da", "el", "et", "fa", "fi", "fil", "gl", "gu", "ha", "he", "hr", "hu", "hy", "id", "is", "jv", "ka", "kk", "kn", "ky", "lb", "ln", "lt", "lv", "mi", "mk", "ml", "mn", "mr", "ms", "mt", "my", "ne", "nl", "no", "oc", "or", "pa", "pl", "ps", "ro", "sd", "sk", "sl", "so", "sr", "sv", "sw", "ta", "te", "tg", "th", "tr", "uk", "ur", "uz", "vi"Same language list as eleven-v4. |
| stabilitynumberdefault 0.5 | 0–1 |
| similarity_boostnumberdefault 0.75 | 0–1 |
| seednumber | 0–4294967295Integer seed; omit for a random seed. |
| audio_output_formatstringdefault "mp3_44100_128" | "mp3_22050_32", "mp3_44100_32", "mp3_44100_64", "mp3_44100_96", "mp3_44100_128", "mp3_44100_192", "opus_48000_32", "opus_48000_64", "opus_48000_96", "opus_48000_128", "opus_48000_192", "pcm_8000", "pcm_16000", "pcm_22050", "pcm_24000", "pcm_44100", "pcm_48000", "ulaw_8000", "alaw_8000" |
| text_normalizationstringdefault "auto" | "auto", "on", "off" |
| timestampsbooleandefault false | Return per-character timestamps in the status response. |
Rules
- Same validation rules as eleven-v4; only the price differs.
Result
Same as eleven-v4: video_url is the generated audio file; optional 'timestamps' array when timestamps:true.
Example request body
{
"mode": "eleven-v4-turbo",
"generation_type": "text-to-speech",
"prompt": "Welcome to Veevid. Let's make something amazing today.",
"duration": "0",
"voice": "Rachel",
"language": "auto",
"stability": 0.5,
"similarity_boost": 0.75,
"audio_output_format": "mp3_44100_128",
"text_normalization": "auto",
"timestamps": false
}Lyria 3.5 (Text to Music) mode: lyria-music
text-to-music
Generates a music track (about 1-2 minutes, vocals or instrumental) from a text prompt with Google Lyria 3.5.
| Parameter | Values & notes |
|---|---|
| promptstring · required | 1–5000Music description (style, mood, instruments, tempo, structure). There is no duration parameter; length can only be described in the prompt (observed minimum ~60s). |
| durationstringdefault "0" | "0"Send "0". |
| imagestring | Optional single reference image https URL. |
Input files
image: Optional, 1 image https URL (jpeg/png/webp, max 10MB; upload it withPOST /api/storage/upload). Only one image is used.
Upload folders for /api/storage/upload
- image:
lyria-music
Rules
- prompt is required and non-empty after trim; max 5000 characters.
- generation_type must be 'text-to-music' (and text-to-music only accepts lyria-music).
- image and every entry of images, if provided, must start with https://.
- Prompts imitating real singers or copyrighted lyrics are rejected by the model; unused credits are refunded automatically.
Result
Async: poll GET /api/video-generation/{id}/status. On completed, video_url is the generated music audio file and media_kind is 'audio'. 'lyrics' (string) is present only when the track has vocals; it is raw Lyria markup (section tags like [[A0]], line timing like [12.0:]), so clean it before displaying.
Example request body
{
"mode": "lyria-music",
"generation_type": "text-to-music",
"prompt": "Upbeat lo-fi hip hop with warm piano chords, soft vinyl crackle and a relaxed 80 BPM groove. About two minutes long.",
"duration": "0"
}Lyria 3.5 (Image to Music) mode: lyria-music
text-to-music
Generates music that matches the mood of a reference image plus a text prompt with Google Lyria 3.5 (image-to-music).
| Parameter | Values & notes |
|---|---|
| promptstring · required | 1–5000Still required. Example: 'Instrumental background music that matches the mood of this image. About one minute long.' |
| imagestring | Reference image https URL; the main input for image-to-music (optional at the API level). |
| durationstringdefault "0" | "0"Send "0". |
Input files
image: 1 image https URL (jpeg/png/webp, max 10MB; upload it withPOST /api/storage/upload). Only one image is used.
Upload folders for /api/storage/upload
- image:
lyria-music
Rules
- Same mode, generation_type, validation and price as the text-to-music entry; there is no separate 'image-to-music' generation_type.
- image must be https; a prompt is still required.
Result
Same as text-to-music: video_url is the music audio file; optional 'lyrics'.
Example request body
{
"mode": "lyria-music",
"generation_type": "text-to-music",
"prompt": "Instrumental background music that matches the mood of this image. About one minute long.",
"duration": "0",
"image": "https://cdn.veevid.ai/image-to-music/i2m-sample-coffee.webp"
}Scribe V2 mode: scribe-v2
speech-to-text
Transcribes an audio file to text with speaker diarization and audio-event tags using ElevenLabs Scribe V2.
| Parameter | Values & notes |
|---|---|
| audiostring · required | https URL of the audio file to transcribe. |
| input_audio_durationnumber · requiredaffects price | 0–3600Audio length in seconds (must be > 0, max 3600); send the actual length. Billing basis: ceil(seconds/60) minutes, min 1. |
| durationstringdefault "0" | "0"Send "0". |
| diarizebooleandefault true | Identify speakers. |
| tag_audio_eventsbooleandefault true | Tag non-speech events like [laughter]. |
| keytermsarray<string>affects price | Proper nouns/brand names to bias recognition; max 100 items, each non-empty and <=50 chars. Send it only when non-empty; keyterms raise the price. |
Input files
audio: 1 audio https URL; .mp3, .wav, .m4a, .ogg, .flac, .aac only (video files such as mp4 are not supported). Max 60 minutes; upload withPOST /api/storage/upload(max 100MB).
Upload folders for /api/storage/upload
- audio:
speech-to-text/audio
Rules
- generation_type must be 'speech-to-text' and vice versa.
- audio is required and must be an https URL; audios (array) is rejected - single file only.
- If the URL has a file extension it must be one of .mp3 .wav .m4a .ogg .flac .aac; URLs without an extension are accepted.
- input_audio_duration is required, > 0 and <= 3600 seconds.
- keyterms: max 100, each non-empty and <= 50 characters after trim.
- Send
input_audio_durationequal to the actual audio length (seconds); if it is under-declared, the difference is charged after transcription. - Failed transcriptions are refunded automatically.
- To quote, send the https
audioURL andinput_audio_duration(the quote checks the URL but does not download it).
Result
Async and fast (poll GET /api/video-generation/{id}/status, e.g. every 4s). On completed: 'transcript' object {v:1, language (ISO 639-3 e.g. 'eng', or null), duration (seconds, end of last word), speakers (count), text (full text), cues: [{s: start_sec, e: end_sec, p: speaker_index_or_null, t: text}]}. Note: video_url here is the original source audio URL (not a result file) and is kept for 14 days. media_kind 'audio'.
Example request body
{
"mode": "scribe-v2",
"generation_type": "speech-to-text",
"duration": "0",
"audio": "https://example.com/interview.mp3",
"input_audio_duration": 95.4,
"diarize": true,
"tag_audio_events": true
}VEED Clean Audio mode: veed-clean-audio
remove-background-noise
Removes background noise from an audio or video file and returns a cleaned, loudness-normalized audio file with VEED Clean Audio.
| Parameter | Values & notes |
|---|---|
| audiostring · required | https URL of the audio OR video file to clean (video: only the audio track is used). |
| input_audio_durationnumber · requiredaffects price | 0–1800Media length in seconds (must be > 0, max 1800); send the actual length. |
| durationstringdefault "0" | "0"Send "0". |
| strengthnumberdefault 0.874 | 0–1How much of the original background is kept under speech; suppression floor = 1 - strength (0.874 = -18 dB). Silence between words is always cleaned. |
| target_lufsnumber|nulldefault -19 | -40–-8Output integrated loudness (LUFS). null = disable loudness normalization (keep original volume); omitted = -19. Gain only, peak capped at -1.1 dBTP so high targets may not be reached. |
| audio_output_formatstringdefault "flac" | "flac", "wav"Output is always 48kHz mono 16-bit |
Input files
audio: 1 https URL; audio .mp3 .wav .m4a .aac .ogg .opus .flac or video .mp4 .mov .m4v .webm. Max 30 minutes and 512MB (upload it withPOST /api/storage/presign-upload, see File Uploads). Video without an audio track fails and credits are refunded automatically.
Rules
- generation_type must be 'remove-background-noise' and vice versa.
- audio is required and must be an https URL; audios (array) is rejected - single file only.
- If the URL has a file extension it must be one of .mp3 .wav .m4a .aac .ogg .opus .flac .mp4 .mov .m4v .webm; URLs without an extension are accepted.
- input_audio_duration is required, > 0 and <= 1800 seconds (30 minutes).
- strength must be in [0,1]; target_lufs must be null or in [-40,-8]; audio_output_format must be flac or wav.
- Send
input_audio_durationequal to the actual media length (seconds); if it is under-declared, the difference is charged after processing. - Failed generations (e.g. 'File has no audio track') are refunded automatically.
- To quote, send the https
audioURL andinput_audio_duration(the quote checks the URL but does not download it).
Result
Async (processing takes about 0.2-0.35x the media length; poll GET /api/video-generation/{id}/status). On completed, video_url is the cleaned audio file (.flac or .wav) and media_kind is 'audio'. No lyrics/transcript/timestamps.
Example request body
{
"mode": "veed-clean-audio",
"generation_type": "remove-background-noise",
"duration": "0",
"audio": "https://example.com/noisy-recording.mp3",
"input_audio_duration": 162.3,
"strength": 0.874,
"target_lufs": -19,
"audio_output_format": "flac"
}Video Tools & Effects
Upscale, resize, edit, extend, insert shots, swap characters, lip-sync, motion control and video effects.
| Model | mode | generation_type |
|---|---|---|
| Kling 3.0 Motion Control | motion-control-standard | motion-control |
| Kling 3.0 Motion Control | motion-control-pro | motion-control |
| AI Hug Video Generator | ai-hug | ai-hug |
| AI Kissing Video Generator | ai-kissing | ai-kissing |
| AI Video Upscaler | upscale | video-upscale |
| AI Video Resizer | reframe | video-reframe |
| Wan 2.7 Edit | wan-2.7-edit | edit-video |
| Wan 3.0 | wan-3.0 | edit-video, video-extend |
| Seedance 2.0 | seedance-2.0 | edit-video, video-extend |
| Seedance 2.5 | seedance-2.5 | edit-video, video-extend |
| LTX 2.3 | ltx-2-3 | video-extend |
| MiniMax H3 Max | minimax-h3-max-extend | video-extend |
| MiniMax H3 Max | minimax-h3-max-insert | video-insert |
| MiniMax H3 Max | minimax-h3-max-recast | character-swap |
| Video Lip Sync | video-lip-sync | lip-sync |
| Google Veo 3.1 | veo3 | video-extend |
| Grok Imagine | grok-imagine | video-extend |
Kling 3.0 Motion Control mode: motion-control-standard
motion-control
Transfers the motion from a driving video onto the character in a reference image (Kling 3.0 Motion Control, Standard).
| Parameter | Values & notes |
|---|---|
| imagestring · required | Public URL of the character image. |
| videostring · required | Public URL of the driving (motion reference) video. |
| durationstring (whole seconds) · requiredaffects price | Driving video length as a whole number of seconds, rounded up (e.g. "5"; decimals are rejected); it determines the price. When the driving video is on Veevid storage the server also measures it and charges by the measured length if that is longer. |
| character_orientationstring · required | "image", "video"'image' keeps the character orientation from the image (video up to 10s); 'video' follows the video (up to 30s). |
| model_versionstringaffects pricedefault "motion-control-standard" | "motion-control-standard", "motion-control-pro"Optional; the Pro tier is used when either mode or model_version is motion-control-pro. |
| promptstring | Optional text guidance, e.g. 'No distortion, the character's movements are consistent with the video.' Omit if blank. |
| keep_original_soundbooleandefault true | Keep the driving video's audio. |
| elementsarray | At most 1 element: {frontal_image_url?: string, reference_image_urls: string[1-3]}. Only allowed when character_orientation is 'video'. Send [] when unused. |
| aspect_ratiostring | "original"Send 'original'; output follows the inputs. |
Input files
image: 1 character image URL (JPEG/PNG/WebP, max 10MB)video: 1 driving video URL; up to 10s when character_orientation=image, up to 30s when =video (mp4/webm/mov, max 200MB)images: optional element images via elements[0].frontal_image_url and elements[0].reference_image_urls (1-3)
Upload folders for /api/storage/upload
- image:
kling-motion-control/reference-images - video:
kling-motion-control/reference-videos - images:
kling-motion-control/reference-images
Rules
- image, video and character_orientation are required (400 otherwise).
- elements are only allowed when character_orientation=video; at most 1 element with 1-3 reference_image_urls.
- Send
durationequal to the driving video length in whole seconds (rounded up). - duration must be a whole number of seconds (decimals are rejected); max 10 when character_orientation=image, max 30 when =video (400 otherwise).
- mode motion-control-standard / motion-control-pro must be used with generation_type motion-control, and vice versa (400 otherwise).
- If the output video is longer than the charged duration, the difference is charged after generation; if your balance cannot cover it, the result is not delivered and the generation fails without a refund.
Example request body
{
"mode": "motion-control-standard",
"generation_type": "motion-control",
"model_version": "motion-control-standard",
"prompt": "No distortion, the character's movements are consistent with the video.",
"image": "https://cdn.veevid.ai/sample/motion-control-image-1.webp",
"video": "https://cdn.veevid.ai/sample/motion-control-video-1.mp4",
"duration": "5",
"aspect_ratio": "original",
"character_orientation": "image",
"keep_original_sound": true,
"elements": []
}Kling 3.0 Motion Control mode: motion-control-pro
motion-control
Transfers the motion from a driving video onto the character in a reference image at Pro quality (Kling 3.0 Motion Control, Pro).
| Parameter | Values & notes |
|---|---|
| imagestring · required | Public URL of the character image. |
| videostring · required | Public URL of the driving (motion reference) video. |
| durationstring (whole seconds) · requiredaffects price | Driving video length as a whole number of seconds, rounded up (e.g. "5"; decimals are rejected); it determines the price. When the driving video is on Veevid storage the server also measures it and charges by the measured length if that is longer. |
| character_orientationstring · required | "image", "video"'image': video up to 10s; 'video': video up to 30s. |
| model_versionstringaffects pricedefault "motion-control-pro" | "motion-control-standard", "motion-control-pro"Optional; the Pro tier is used when either mode or model_version is motion-control-pro. |
| promptstring | Optional text guidance (omit if blank). |
| keep_original_soundbooleandefault true | Keep the driving video's audio. |
| elementsarray | At most 1 element: {frontal_image_url?: string, reference_image_urls: string[1-3]}; only with character_orientation='video'. Send [] when unused. |
| aspect_ratiostring | "original"Send 'original'; output follows the inputs. |
Input files
image: 1 character image URL (JPEG/PNG/WebP, max 10MB)video: 1 driving video URL; up to 10s when character_orientation=image, up to 30s when =video (max 200MB)images: optional element images via elements[0].frontal_image_url and elements[0].reference_image_urls (1-3)
Upload folders for /api/storage/upload
- image:
kling-motion-control/reference-images - video:
kling-motion-control/reference-videos - images:
kling-motion-control/reference-images
Rules
- image, video and character_orientation are required (400 otherwise).
- elements are only allowed when character_orientation=video; at most 1 element with 1-3 reference_image_urls.
- Send
durationequal to the driving video length in whole seconds (rounded up). - duration must be a whole number of seconds (decimals are rejected); max 10 when character_orientation=image, max 30 when =video (400 otherwise).
- mode motion-control-standard / motion-control-pro must be used with generation_type motion-control, and vice versa (400 otherwise).
- If the output video is longer than the charged duration, the difference is charged after generation; if your balance cannot cover it, the result is not delivered and the generation fails without a refund.
Example request body
{
"mode": "motion-control-pro",
"generation_type": "motion-control",
"model_version": "motion-control-pro",
"prompt": "No distortion, the character's movements are consistent with the video.",
"image": "https://cdn.veevid.ai/sample/motion-control-image-1.webp",
"video": "https://cdn.veevid.ai/sample/motion-control-video-1.mp4",
"duration": "5",
"aspect_ratio": "original",
"character_orientation": "video",
"keep_original_sound": true,
"elements": []
}AI Hug Video Generator mode: ai-hug
ai-hug
Generates a video of two people (or a person and a pet) from two separate photos hugging in a described scene.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Only the scene sentence (max 500 chars). It is automatically wrapped in a fixed template that references the two photos as @Element1/@Element2; do not write the full prompt. |
| imagesstring[] · required | Exactly 2 image URLs: [person A, person B]; images[0] = @Element1. |
| resolutionstringaffects pricedefault "720p" | "720p", "1080p", "4k"Uses resolution, not video_quality. |
| durationstringaffects pricedefault "5" | 3–15Integer seconds. |
| aspect_ratiostringdefault "16:9" | "16:9", "9:16", "1:1" |
| generate_audiobooleanaffects pricedefault false | Generate ambient audio. |
Input files
images: exactly 2 image URLs, one subject each (up to 10MB per photo)
Upload folders for /api/storage/upload
- image:
ai-hug
Rules
- mode and generation_type must both be 'ai-hug'
- images must contain exactly 2 URLs
- prompt (scene) is required and at most 500 characters
- resolution must be 720p, 1080p or 4k; duration integer 3-15; aspect_ratio 16:9, 9:16 or 1:1
- scene_template is rejected for this mode (only ai-kissing accepts it)
Example request body
{
"mode": "ai-hug",
"generation_type": "ai-hug",
"prompt": "They spot each other in a sunlit park, walk closer and share a warm, tight hug.",
"images": [
"https://example.com/person-a.jpg",
"https://example.com/person-b.jpg"
],
"resolution": "720p",
"duration": "5",
"aspect_ratio": "16:9",
"generate_audio": false
}AI Kissing Video Generator mode: ai-kissing
ai-kissing
Generates a video of two adults kissing, either from two separate photos or from one photo of both people.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Only the scene sentence (max 500 chars); may refer to the two people only as @Element1 and @Element2 (exact spelling). It is automatically wrapped in a fixed template. |
| imagesstring[] · required | 1 URL (one photo of both people, used as first frame) or 2 URLs ([person 1, person 2], images[0] = @Element1). Count selects the input mode. |
| scene_templatestringdefault "default" | "default", "stylized"'stylized' for anime/3D style scenes (drops realism constraints). |
| resolutionstringaffects pricedefault "720p" | "720p", "1080p", "4k"Uses resolution, not video_quality. |
| durationstringaffects price | 3–15Integer seconds. Default 5 with 2 photos, 8 with 1 photo. |
| aspect_ratiostringdefault "16:9" | "16:9", "9:16", "1:1"No effect in one-photo mode (output follows the photo). |
| generate_audiobooleanaffects pricedefault false |
Input files
images: 1 image URL (couple photo) or 2 image URLs (one per person); up to 10MB per photo
Upload folders for /api/storage/upload
- image:
ai-kissing
Rules
- mode and generation_type must both be 'ai-kissing' (legacy mode 'normal' is rejected)
- images must contain 1 or 2 URLs
- prompt (scene) is required, max 500 characters; only @Element1/@Element2 references allowed
- scene_template must be default or stylized
- resolution 720p/1080p/4k; duration integer 3-15; aspect_ratio 16:9, 9:16 or 1:1
- When quoting with
/api/quote, sendimagesorinput_image_count(1 or 2) — the input mode decides the default duration (422 otherwise).
Example request body
{
"mode": "ai-kissing",
"generation_type": "ai-kissing",
"prompt": "Standing face to face in a sunlit park at golden hour, @Element1 tilts their head up and closes their eyes while @Element2 gently cups @Element1's face with one hand and leans in; they share a soft, tender kiss on the lips.",
"images": [
"https://example.com/person-1.jpg",
"https://example.com/person-2.jpg"
],
"scene_template": "default",
"resolution": "720p",
"duration": "5",
"aspect_ratio": "16:9",
"generate_audio": false
}AI Video Upscaler mode: upscale
video-upscale
Upscales an existing video to 720p, 1080p or 4K and optionally raises its frame rate.
| Parameter | Values & notes |
|---|---|
| videostring · required | URL of the source video (singular field). |
| video_qualitystring · requiredaffects price | "720p", "1080p", "4K"Target resolution. Always send video_quality. |
| target_fpsnumberaffects pricedefault 30 | 15–60Target frame rate. Above 30 fps costs more at 1080p and 4K. |
| durationstringaffects pricedefault "5" | Source video length in whole seconds, as a string. Send the actual length of the source video; it determines the price (5 when omitted). |
| video_sourcestring | "upload", "url"'upload' for files uploaded with POST /api/storage/upload, 'url' for an external link. With 'url' the link must answer a HEAD request with a supported video type and be <=200MB. |
| promptstringdefault "Upscale" | Optional; has no effect and can be omitted. |
Input files
video: 1 video URL (mp4, mov, webm, m4v, mkv, avi, wmv, ts, vob, mpg, flv, 3gp or gif), up to 200MB.
Upload folders for /api/storage/upload
- video:
upscale-input
Rules
- video is required (400 'Video is required for video-upscale generation').
- When sent, duration must be a positive integer (400).
- target_fps must be 15-60.
- With video_source 'url': the URL must answer a HEAD request within 5s, have a supported video content-type or file extension, and be <=200MB.
- Processing is asynchronous: the response has status 'processing' and a generation_id to poll.
Example request body
{
"mode": "upscale",
"generation_type": "video-upscale",
"prompt": "Upscale",
"video": "https://cdn.veevid.ai/upscale-input/clip.mp4",
"video_quality": "1080p",
"target_fps": 30,
"duration": "10",
"video_source": "upload"
}AI Video Resizer mode: reframe
video-reframe
Changes a video's aspect ratio by using AI to extend the frame (outpainting) instead of cropping.
| Parameter | Values & notes |
|---|---|
| videostring · required | URL of the source video (singular field). |
| aspect_ratiostringdefault "16:9" | "1:1", "16:9", "9:16", "4:3", "3:4", "21:9", "9:21"Target aspect ratio; use one of the listed values (16:9 when omitted). |
| durationstringaffects pricedefault "5" | Source video length in whole seconds, as a string. Send the actual length of the source video. |
| promptstringdefault "Reframe Video" | Optional description of the scene to guide the extended areas. The literal 'Reframe Video' means no prompt. |
| video_sourcestring | "upload", "url"Optional: 'upload' or 'url'. |
| video_qualitystringdefault "standard" | Optional; has no effect (send 'standard' or omit). |
Input files
video: 1 video URL, up to 100MB and up to 30s.
Upload folders for /api/storage/upload
- video:
reframe-input
Rules
- video is required (400 'Video is required for video-reframe generation').
- Always send
durationequal to the source video length (whole seconds). - aspect_ratio must be one of the 7 listed values.
- The source video must be at most 30s and 100MB.
- Processing is asynchronous: poll the returned generation_id.
Example request body
{
"mode": "reframe",
"generation_type": "video-reframe",
"video": "https://cdn.veevid.ai/reframe-input/clip.mp4",
"aspect_ratio": "9:16",
"prompt": "Reframe Video",
"duration": "8",
"video_source": "upload",
"video_quality": "standard"
}Wan 2.7 Edit mode: wan-2.7-edit
edit-video
Edits an existing 2-10s video according to a text instruction, with an optional reference image.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Editing instruction (required). |
| videostring · required | Source video URL (singular field). MP4 or MOV, 2-10s, <=100MB. |
| durationstring · requiredaffects price | 2–10Output length in whole seconds (2-10). To match the input, send the source length rounded up. |
| video_qualitystringdefault "1080p" | "720p", "1080p"Output resolution. Same price for both. |
| aspect_ratiostring | "auto", "16:9", "9:16", "1:1", "4:3", "3:4"Omit (or send 'auto') to keep the source ratio. |
| audio_settingstringdefault "auto" | "auto", "origin"'auto' lets the model generate audio; 'origin' keeps the source audio. |
| video_sourcestring | "upload", "url"'upload' for files uploaded with POST /api/storage/upload; 'url' for an external link, which must pass a HEAD check (supported video type, <=100MB). |
| imagestring | Optional reference image URL (http/https) for the edit. |
Input files
video: 1 source video URL (MP4/MOV, 2-10s, <=100MB)image: optional 1 reference image URL (jpg/png/webp)
Upload folders for /api/storage/upload
- video:
edit-video-input - image:
edit-video-ref
Rules
- Always send
duration. - duration must be an integer string 2-10 (400).
- video_quality, if sent, must be 720p or 1080p; aspect_ratio, if sent, must be auto, 16:9, 9:16, 1:1, 4:3 or 3:4 (400).
- video (singular) is required (400 'Video is required for edit-video generation').
- With video_source 'url' the link must pass a HEAD check (supported video type, <=100MB).
- Source videos uploaded to edit-video-input must be .mp4/.mov, up to 100MB.
- prompt is required.
Example request body
{
"mode": "wan-2.7-edit",
"generation_type": "edit-video",
"prompt": "Change the jacket to red leather",
"video": "https://cdn.veevid.ai/edit-video-input/clip.mp4",
"duration": "6",
"video_quality": "720p",
"audio_setting": "auto",
"video_source": "upload"
}Wan 3.0 mode: wan-3.0
edit-video
Edits a source video with a text instruction plus optional reference images and audio, and lets you pick the output length.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Editing instruction (required). It is automatically prefixed with 'Edit Video 1: ' when it lacks an edit intent. |
| videosstring[] · requiredaffects price | [source video URL]. Source seconds are billed together with output seconds. |
| imagesstring[] | Optional reference images, up to 10. |
| audiosstring[] | Optional reference audio files, up to 5. |
| durationstringaffects pricedefault "5" | 2–30Output length in whole seconds (2-30); -1 (smart duration) is rejected. |
| video_qualitystringaffects pricedefault "720p" | "480p", "720p", "1080p"Output resolution. |
| aspect_ratiostringdefault "adaptive" | "adaptive", "16:9", "9:16", "1:1", "4:3", "3:4"Use one of the listed values; 'adaptive' keeps the source ratio. |
| generate_audiobooleandefault true | Generate a soundtrack. |
| input_video_durationnumberaffects price | 0–15Source video length in seconds, rounded up. Always send it with the source video. |
| has_video_inputboolean | Quote-only field: send true to /api/quote together with the source video length; /api/generate-video ignores it. |
Input files
video: 1 source video URL in videos[] (1-15s, up to 100MB)images: 0-10 reference image URLs (up to 20MB each)audio: 0-5 reference audio URLs in audios[]
Upload folders for /api/storage/upload
- video:
wan30/reference-videos - image:
wan30 - audio:
wan30/audio
Rules
- Send input_video_duration equal to the source video length (seconds); values > 15 are rejected (400).
- duration must be an integer 2-30; -1 is rejected (400).
- input_video_duration + duration must be <= 30 (400).
- At most 10 images, 5 videos and 5 audios (400).
- videos[] must contain the source video (400 'Video is required for edit-video generation').
- prompt is required.
- When quoting with
/api/quote, addhas_video_input: trueandinput_video_durationif you send the source video (the quote endpoint does not read video URLs, so the price would come out too low).
Example request body
{
"mode": "wan-3.0",
"generation_type": "edit-video",
"prompt": "Replace the skateboarder with a short-haired woman in a yellow raincoat",
"videos": [
"https://cdn.veevid.ai/wan30/reference-videos/clip.mp4"
],
"duration": "5",
"video_quality": "720p",
"aspect_ratio": "adaptive",
"generate_audio": true,
"input_video_duration": 6
}video-extend
Continues an uploaded video (up to 15s), keeping its characters, style and sound.
| Parameter | Values & notes |
|---|---|
| promptstring | What happens next. It is automatically prefixed with 'Extend Video 1 by N seconds.' when no extend intent is present. |
| videosstring[] · requiredaffects price | [source video URL]. No taskId needed. |
| durationstringaffects pricedefault "5" | 2–30Length of the new footage in seconds. It must be <= 30 - source seconds. |
| video_qualitystringaffects pricedefault "720p" | "480p", "720p", "1080p"Output resolution. |
| aspect_ratiostringdefault "16:9" | "adaptive"Send 'adaptive' so the output follows the source; when omitted, 16:9 is used. |
| generate_audiobooleandefault true | Generate audio. |
| input_video_durationnumberaffects price | 0–15Source length in seconds, rounded up. Always send it with the source video. |
Input files
video: 1 source video URL in videos[] (1-15s, up to 100MB)
Upload folders for /api/storage/upload
- video:
wan30-extend
Rules
- Send input_video_duration equal to the source video length (seconds); values > 15 are rejected (400).
- duration must be an integer 2-30; -1 is rejected; source + duration must be <= 30 (400).
- At most 5 videos, 10 images and 5 audios (400); only the source video is needed.
- A source video in videos[] is required.
- When quoting with
/api/quote, addhas_video_input: trueandinput_video_durationif you send the source video (the quote endpoint does not read video URLs, so the price would come out too low).
Example request body
{
"mode": "wan-3.0",
"generation_type": "video-extend",
"prompt": "She turns around and waves at the camera",
"videos": [
"https://cdn.veevid.ai/wan30-extend/clip.mp4"
],
"duration": "5",
"video_quality": "720p",
"aspect_ratio": "adaptive",
"generate_audio": true,
"input_video_duration": 6
}Seedance 2.0 mode: seedance-2.0
edit-video
Edits a source video (up to 15s) using a prompt plus optional reference images and audio.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Editing instruction (required). |
| videosstring[] · requiredaffects price | [source video URL]. The source length is billed. |
| imagesstring[] | Optional reference images; up to 9 are used. |
| audiosstring[] | Optional reference audio; up to 3 are used. |
| durationstringaffects pricedefault "5" | 4–15Output length in seconds, 4-15. Out-of-range values are clamped to 4-15. |
| video_qualitystringaffects pricedefault "480p" | "480p", "720p", "1080p", "4k"Output resolution. 1080p and 4k require model_version 'seedance-2.0'; with mini or fast use 480p or 720p. |
| aspect_ratiostringdefault "16:9" | "auto", "16:9", "9:16", "1:1", "4:3", "3:4", "21:9"'auto' keeps the source ratio. Use one of the listed values. |
| model_versionstringaffects pricedefault "seedance-2.0" | "seedance-2.0-mini", "seedance-2.0-fast", "seedance-2.0"Quality/price tier. Send model_version explicitly. |
| generate_audiobooleandefault true | Generate audio. |
| input_video_durationnumberaffects price | 0–15Source video length in seconds (billed as at least 2s). Always send it with the source video. |
| has_video_inputboolean | Quote-only field: send true to /api/quote together with the source video length; /api/generate-video ignores it. |
Input files
video: 1 source video URL in videos[] (up to 15s and 200MB)images: 0-9 reference image URLs (up to 30MB each)audio: 0-3 reference audio URLs in audios[] (.wav/.mp3, up to 15MB each, 15s total)
Upload folders for /api/storage/upload
- video:
seedance20-edit - image:
seedance20 - audio:
seedance20-edit/audio
Rules
- Send input_video_duration equal to the source video length (seconds); values > 15 are rejected (400).
- Send
model_versionexplicitly; with seedance-2.0-mini or seedance-2.0-fast use 480p or 720p. - videos[] must contain the source video (400 'Video is required for edit-video generation').
- Only the first 9 images, 3 videos and 3 audios are used; duration is normalized to 4-15 rather than rejected.
- Not supported: output_format (mp4/mov is only valid for Seedance 2.5).
- When quoting with
/api/quote, addhas_video_input: trueandinput_video_durationif you send the source video (the quote endpoint does not read video URLs, so the price would come out too low).
Example request body
{
"mode": "seedance-2.0",
"generation_type": "edit-video",
"prompt": "Make it snow in the scene",
"videos": [
"https://cdn.veevid.ai/seedance20-edit/clip.mp4"
],
"duration": "5",
"video_quality": "480p",
"aspect_ratio": "auto",
"model_version": "seedance-2.0-fast",
"generate_audio": true,
"input_video_duration": 8
}video-extend
Continues an uploaded video by 4-15s with optional reference images and audio.
| Parameter | Values & notes |
|---|---|
| promptstring | What happens next. |
| videosstring[] · requiredaffects price | [source video URL]. No taskId needed. |
| imagesstring[] | Optional reference images; up to 9 are used. |
| audiosstring[] | Optional reference audio; up to 3 are used. |
| durationstringaffects pricedefault "5" | 4–15Length of the new footage in seconds; clamped to 4-15. |
| video_qualitystringaffects pricedefault "480p" | "480p", "720p", "1080p", "4k"1080p and 4k require model_version 'seedance-2.0'; with mini or fast use 480p or 720p. |
| aspect_ratiostringdefault "16:9" | "auto", "16:9", "9:16", "1:1", "4:3", "3:4", "21:9"'auto' keeps the source ratio. |
| model_versionstringaffects pricedefault "seedance-2.0" | "seedance-2.0-mini", "seedance-2.0-fast", "seedance-2.0"Quality/price tier. Send model_version explicitly. |
| generate_audiobooleandefault true | Generate audio. |
| input_video_durationnumberaffects price | 0–15Source length in seconds (billed as at least 2s). Always send it with the source video. |
| has_video_inputboolean | Quote-only field: send true to /api/quote together with the source video length; /api/generate-video ignores it. |
Input files
video: 1 source video URL in videos[] (<=15s; mp4/mov/webm up to 200MB)images: 0-9 reference image URLsaudio: 0-3 reference audio URLs in audios[] (.wav/.mp3, up to 15MB each, 15s total)
Upload folders for /api/storage/upload
- video:
seedance20-extend - image:
seedance20-extend - audio:
seedance20-extend/audio
Rules
- Send input_video_duration equal to the source video length (seconds); values > 15 are rejected (400).
- Send
model_versionexplicitly; with seedance-2.0-mini or seedance-2.0-fast use 480p or 720p. - The source video is required in videos[].
- Only the first 9 images, 3 videos and 3 audios are used.
- Reference images uploaded to the seedance20-extend folder must be .jpg/.jpeg/.png/.tiff/.gif (not .webp).
- When quoting with
/api/quote, addhas_video_input: trueandinput_video_durationif you send the source video (the quote endpoint does not read video URLs, so the price would come out too low).
Example request body
{
"mode": "seedance-2.0",
"generation_type": "video-extend",
"prompt": "The camera follows the runner across the bridge",
"videos": [
"https://cdn.veevid.ai/seedance20-extend/clip.mp4"
],
"duration": "5",
"video_quality": "480p",
"aspect_ratio": "auto",
"model_version": "seedance-2.0-fast",
"generate_audio": true,
"input_video_duration": 8
}Seedance 2.5 mode: seedance-2.5
edit-video
Edits a source video (up to 30s) with a prompt and optional references; the output keeps the source's length and aspect ratio.
| Parameter | Values & notes |
|---|---|
| promptstring · required | Editing instruction. It is automatically prefixed with 'Edit the video: ' when no editing keyword (edit/add/remove/replace/change/...) is present. |
| videosstring[] · requiredaffects price | [source video URL]. |
| imagesstring[] | Optional reference images; up to 30 are used. |
| audiosstring[] | Optional reference audio; up to 10 are used. |
| durationstring · requiredaffects price | "-1"Send "-1". Edit output always matches the source length. |
| video_qualitystringaffects pricedefault "480p" | "480p", "720p", "1080p"Output resolution. Use one of the listed values; 1080p output is 10-bit HEVC. |
| aspect_ratiostringdefault "auto" | "auto"Send 'auto' or omit; edits always keep the source aspect ratio. |
| generate_audiobooleandefault true | Generate audio. |
| input_video_durationnumberaffects price | 0–30Source video length in seconds (billed as at least 2s). Always send it with the source video. |
| output_formatstringdefault "mp4" | "mp4", "mov"Optional output container: mp4 or mov. |
| has_video_inputboolean | Quote-only field: send true to /api/quote together with the source video length; /api/generate-video ignores it. |
Input files
video: 1 source video URL in videos[] (up to 30s and 200MB)images: 0-30 reference image URLs (up to 30MB each)audio: 0-10 reference audio URLs in audios[] (.wav/.mp3, up to 15MB each)
Upload folders for /api/storage/upload
- video:
seedance25-edit - image:
seedance25 - audio:
seedance25-edit/audio
Rules
- duration must be an integer 4-30 or -1 (400); for edits send -1.
- Send input_video_duration equal to the source video length (seconds); values > 30 are rejected (400).
- output_format accepts only mp4 or mov (400 otherwise).
- videos[] must contain the source (400 'Video is required for edit-video generation').
- 480p results are usually drafts: the status response then includes
seedance25_draft, and the draft can be upgraded once to a 1080p version within 7 days withPOST /api/video-generation/{id}/upgrade-draft. - When quoting with
/api/quote, addhas_video_input: trueandinput_video_durationif you send the source video (the quote endpoint does not read video URLs, so the price would come out too low).
Example request body
{
"mode": "seedance-2.5",
"generation_type": "edit-video",
"prompt": "Replace the car with a red vintage convertible",
"videos": [
"https://cdn.veevid.ai/seedance25-edit/clip.mp4"
],
"duration": "-1",
"video_quality": "480p",
"aspect_ratio": "auto",
"generate_audio": true,
"input_video_duration": 10
}video-extend
Continues an uploaded video by 4-30s with optional reference images and audio.
| Parameter | Values & notes |
|---|---|
| promptstring | What happens next. It is automatically prefixed with 'Continue the video: ' when there is no extend keyword (extend/continue/continuation). |
| videosstring[] · requiredaffects price | [source video URL]. No taskId needed. |
| imagesstring[] | Optional reference images; up to 30 are used. |
| audiosstring[] | Optional reference audio; up to 10 are used. |
| durationstringaffects pricedefault "5" | 4–30Length of the new footage in seconds. -1 lets the model choose and is billed as the source length (min 4). |
| video_qualitystringaffects pricedefault "480p" | "480p", "720p", "1080p"Output resolution. |
| output_formatstringdefault "mp4" | "mp4", "mov"mov uses H.264 + PCM audio and usually cannot play in browsers. |
| aspect_ratiostringdefault "auto" | "auto"The output follows the source aspect ratio. |
| generate_audiobooleandefault true | Generate audio. |
| input_video_durationnumberaffects price | 0–30Source length in seconds, rounded up (billed as at least 2s). Always send it with the source video. |
| has_video_inputboolean | Quote-only field: send true to /api/quote together with the source video length; /api/generate-video ignores it. |
Input files
video: 1 source video URL in videos[] (<=30s; mp4/mov/webm up to 200MB)images: 0-30 reference image URLsaudio: 0-10 reference audio URLs in audios[] (.wav/.mp3, up to 15MB each, 30s total)
Upload folders for /api/storage/upload
- video:
seedance25-extend - image:
seedance25-extend - audio:
seedance25-extend/audio
Rules
- duration must be an integer 4-30 or -1 (400).
- Send input_video_duration equal to the source video length (seconds); values > 30 are rejected (400).
- output_format accepts only mp4 or mov (400).
- A source video in videos[] is required.
- Reference images uploaded to the seedance25-extend folder must be .jpg/.jpeg/.png/.tiff/.gif (not .webp).
- 480p results are usually drafts: the status response then includes
seedance25_draft, and the draft can be upgraded once to a 1080p version within 7 days withPOST /api/video-generation/{id}/upgrade-draft. - When quoting with
/api/quote, addhas_video_input: trueandinput_video_durationif you send the source video (the quote endpoint does not read video URLs, so the price would come out too low).
Example request body
{
"mode": "seedance-2.5",
"generation_type": "video-extend",
"prompt": "Continue the video: the dog runs to the beach",
"videos": [
"https://cdn.veevid.ai/seedance25-extend/clip.mp4"
],
"duration": "5",
"video_quality": "480p",
"output_format": "mp4",
"aspect_ratio": "auto",
"generate_audio": true,
"input_video_duration": 8
}LTX 2.3 mode: ltx-2-3
video-extend
Extends an uploaded video by 1-20s at its start or end, with optional context length control.
| Parameter | Values & notes |
|---|---|
| videostring · required | Source video URL (singular field). |
| promptstring | Optional description of the extension. |
| durationstringaffects pricedefault "5" | 1–20Seconds to add (1-20). Decimals are allowed (e.g. 0.5s steps). |
| extend_modestringdefault "end" | "start", "end"Where to add footage. |
| contextnumber | 1–20Seconds of the source to use as context. Omit to let the model maximize context. |
| video_sourcestring | "upload", "url"'upload' for files uploaded with POST /api/storage/upload; 'url' for an external link, which must pass a HEAD check (supported video type, <=200MB). |
| aspect_ratiostringdefault "original" | "original"Optional; send 'original' or omit. Output keeps the source aspect ratio. |
Input files
video: 1 source video URL (mp4/mov/webm/m4v/gif, up to 20s and 200MB)
Upload folders for /api/storage/upload
- video:
video-generator-ltx23-extend
Rules
- duration must be a number between 1 and 20 (400).
- context, if sent, must be 1-20.
- video (singular) is required (400 'Video is required for LTX 2.3 extend generation').
- No taskId is used; model_version and video_quality have no effect for extends.
Example request body
{
"mode": "ltx-2-3",
"generation_type": "video-extend",
"prompt": "The waves keep rolling in at sunset",
"video": "https://cdn.veevid.ai/video-generator-ltx23-extend/clip.mp4",
"duration": "5",
"extend_mode": "end",
"aspect_ratio": "original",
"video_source": "upload"
}MiniMax H3 Max mode: minimax-h3-max-extend
video-extend
Continues an uploaded clip (1.625-60s) by 5-15s at up to 2K resolution.
| Parameter | Values & notes |
|---|---|
| promptstring · required | What happens next (required). |
| videosstring[] · required | [source video URL]; exactly one. The singular 'video' is also accepted. |
| durationstringaffects pricedefault "5" | 5–15Seconds of new footage (whole seconds). |
| video_qualitystringaffects pricedefault "768p" | "480p", "768p", "1080p", "2k", "2K"Output resolution. |
| aspect_ratiostringdefault "auto" | "auto", "21:9", "16:9", "4:3", "1:1", "3:4", "9:16"auto (or adaptive) follows the source. A fixed ratio crops the output. If omitted, auto is used, not 16:9. |
| extend_outputstringdefault "extended" | "extended", "continuation"'extended' returns source + new footage; 'continuation' returns only the new footage plus ~1.6s of the source tail. |
| prompt_expansion_modestringdefault "balanced" | "balanced", "disabled"'balanced' enables prompt expansion; 'disabled' turns it off. |
| input_video_durationnumber | 1.625–60Source length in seconds (up to 2 decimals). Used for validation only; the source is not billed. |
| seednumber | Optional integer >= 0. |
Input files
video: 1 source video URL in videos[] (1.625-60s, <=50MB, aspect ratio 0.4-2.5; mp4/mov/webm/m4v)
Upload folders for /api/storage/upload
- video:
minimax-h3-max-extend
Rules
- The source length and extend_output do not affect price.
- generation_type must be video-extend; duration must be an integer 5-15 (400).
- Exactly one source video; prompt is required; images and audio are rejected (400).
- aspect_ratio must be auto/adaptive/21:9/16:9/4:3/1:1/3:4/9:16; video_quality must be 480p/768p/1080p/2k/2K (400).
- A declared input_video_duration must be 1.625-60s (400).
- The source video must be <=50MB with an aspect ratio between 0.4 and 2.5; out-of-spec sources fail during generation. Upload the source with POST /api/storage/upload (the upload only checks the 50MB size).
Example request body
{
"mode": "minimax-h3-max-extend",
"generation_type": "video-extend",
"prompt": "The cat jumps onto the windowsill and looks outside",
"videos": [
"https://cdn.veevid.ai/minimax-h3-max-extend/clip.mp4"
],
"duration": "5",
"video_quality": "768p",
"aspect_ratio": "auto",
"extend_output": "extended",
"prompt_expansion_mode": "balanced",
"input_video_duration": 6.04
}MiniMax H3 Max mode: minimax-h3-max-insert
video-insert
Inserts a newly generated 5-13s scene between two time points of a source video.
| Parameter | Values & notes |
|---|---|
| videostring · required | Source video URL (singular field; 'videos' means reference videos here). |
| promptstring | Describes the new scene. Required unless images or videos are given. |
| durationstringaffects pricedefault "5" | 5–13Length of the new scene in seconds; decimals allowed. |
| video_qualitystringaffects pricedefault "768p" | "480p", "768p"Output resolution. |
| start_timenumber · required | 1.625–60Cut-out point in seconds (24 fps frame grid). |
| resume_timenumber · required | 1.667–60Resume point in seconds. It must be at least 1 frame (1/24s) after start_time and leave >= 33/24s (+0.25s safety margin) of the source after it. Source between start and resume is replaced. |
| imagesstring[]affects price | Reference images for the new scene, up to 9. |
| videosstring[]affects price | Reference videos, up to 3. Each must be >= 2s, <= 15s in total, and uploaded with POST /api/storage/upload. |
| reference_video_durationnumberaffects price | 0–15Send the total reference-video length in seconds. The server also measures the files and bills the measured value if it is more than 0.5s longer; if it is omitted with reference videos, the 15-second maximum is billed. |
| input_video_durationnumber | 0–60Source video length in seconds; used to validate resume_time. The source is not billed. |
| color_matchbooleandefault true | Match the new scene's color grading to the source. |
| prompt_expansion_modestringdefault "balanced" | "balanced", "disabled"Only 'disabled' turns prompt expansion off. |
| seednumber | Optional integer >= 0. |
Input files
video: 1 source video URL in 'video' (about 3.3-60s, <=50MB; mp4/mov/webm/m4v)images: 0-9 reference image URLs in 'images' (jpg/png/webp)videos: 0-3 reference video URLs in 'videos' (mp4/mov, each >=2s, total <=15s, <=100MB each, uploaded with POST /api/storage/upload)
Upload folders for /api/storage/upload
- video:
minimax-h3-max-insert - images:
minimax-h3-max-insert/reference-images - videos:
minimax-h3-max-insert/reference-videos
Rules
- When sending reference videos, always send reference_video_duration equal to their total length (seconds).
- generation_type must be video-insert; video_quality must be 480p or 768p (400).
- duration must be 5-13 (400).
- Exactly one source in 'video'. Reference images must use 'images'; a singular 'image' is rejected (400).
- A prompt, a reference image or a reference video is required; audio is rejected (400).
- Reference videos must be uploaded with POST /api/storage/upload first; external URLs are rejected (400). Each must be >= 2s, the total <= 15s, and at most 3 videos and 9 images.
- start_time must be 1.625-60. resume_time must be > start_time by >= 1/24s and <= 60. With a declared source length, resume_time must be <= floor((source - 1.375 - 0.25) x 24)/24 (400).
- A declared source length must be ~3.29-60s (400).
- When quoting, send
input_image_countfor reference images, and for reference videos addhas_video_input: true,input_video_countandreference_video_duration(the quote endpoint does not read these files, so the price would come out too low).
Example request body
{
"mode": "minimax-h3-max-insert",
"generation_type": "video-insert",
"prompt": "A hot air balloon drifts across the sky",
"video": "https://cdn.veevid.ai/minimax-h3-max-insert/clip.mp4",
"duration": "5",
"video_quality": "768p",
"start_time": 3,
"resume_time": 3.042,
"input_video_duration": 10,
"color_match": true,
"prompt_expansion_mode": "balanced"
}MiniMax H3 Max mode: minimax-h3-max-recast
character-swap
Swaps the people in a 5-30s video with people from 1-4 reference photos, keeping motion, camera and original audio.
| Parameter | Values & notes |
|---|---|
| videostring · required | Source video URL (singular field). Must be uploaded with POST /api/storage/upload, MP4/MOV, 5-30s, and no single shot over 15s. |
| imagesstring[] · required | 1-4 reference photos, one per new person, mapped left to right by default. |
| promptstring | Optional (<=2000 chars), e.g. which person becomes whom. |
| video_qualitystringaffects pricedefault "1080p" | "768p", "1080p"Output resolution. Always send video_quality. |
| input_video_durationnumberaffects price | 5–30Source length in seconds (up to 2 decimals). Send the actual length; the server measures the file and bills the measured length. |
| seednumber | Optional integer >= 0. |
Input files
video: 1 source video URL in 'video' (MP4/MOV, 5-30s, <=50MB, uploaded with POST /api/storage/upload)images: 1-4 reference photo URLs in 'images' (jpg/png/webp, <=20MB each)
Upload folders for /api/storage/upload
- video:
minimax-h3-max-recast - images:
minimax-h3-max-recast/reference-images
Rules
- Seconds = server-measured source length (minimum 5). Reference photos are free.
- generation_type must be character-swap; video_quality must be 768p or 1080p (400). Always send video_quality.
- Input files must be uploaded with POST /api/storage/upload first (external URLs are rejected).
- If the source video's duration cannot be read from its MP4/MOV header, the request returns 422 and nothing is charged.
- A measured length outside 5-30s (+-0.1s) is rejected (400).
- If the measured length exceeds the declared input_video_duration by >0.1s, it returns 409 with measured_duration and required_credits and charges nothing; resubmit with input_video_duration = measured_duration.
- 1-4 images required; a singular 'image', reference videos or audio are rejected (400); prompt <= 2000 chars.
- Duration and aspect ratio are not inputs: the output matches the source.
Example request body
{
"mode": "minimax-h3-max-recast",
"generation_type": "character-swap",
"prompt": "Replace the man on the left with the person in photo 1",
"video": "https://cdn.veevid.ai/minimax-h3-max-recast/clip.mp4",
"images": [
"https://cdn.veevid.ai/minimax-h3-max-recast/reference-images/person1.jpg"
],
"video_quality": "768p",
"input_video_duration": 8.04
}Video Lip Sync mode: video-lip-sync
lip-sync
Re-syncs the mouth movements in an existing video to a new voice track; output length equals the audio length.
| Parameter | Values & notes |
|---|---|
| videosstring[] · required | [source video URL]; only the first is used. The singular 'video' is also accepted. Must be uploaded with POST /api/storage/upload. |
| audiosstring[] · required | [voice audio URL]; only the first is used. The singular 'audio' is also accepted. Must be uploaded with POST /api/storage/upload; use clean vocals. |
| input_audio_durationnumberaffects price | 0–61Send the audio length in seconds. |
| lip_sync_modestringdefault "lite" | "lite", "basic"lite: single front-facing speaker, faster. basic: complex scenes, multiple shots. |
| lip_sync_separate_vocalbooleandefault false | Isolate vocals / suppress background noise. |
| lip_sync_open_scenedetbooleandefault true | Shot and speaker detection (basic only). |
| lip_sync_align_audiobooleandefault true | Loop the video when the audio is longer (lite only; always on in basic). |
| lip_sync_align_audio_reversebooleandefault false | Loop by playing in reverse (lite only; requires align_audio). |
| lip_sync_templ_start_secondsnumberdefault 0 | 0–3600Start offset in the source video, floored to whole seconds (lite only). |
Input files
video: 1 source video URL (mp4/mov, <=100MB, short side >=360p, uploaded with POST /api/storage/upload)audio: 1 voice audio URL (mp3/wav/m4a/aac/ogg, <=10MB, <=60s, uploaded with POST /api/storage/upload)
Upload folders for /api/storage/upload
- video:
video-lip-sync - audio:
video-lip-sync/audio
Rules
- Always send input_audio_duration equal to the audio length (seconds); if it is omitted you are charged for the 60-second maximum. If the output is longer than declared, the difference is charged after processing (an over-declared length is not refunded).
- generation_type must be lip-sync (422).
- The video and audio must be uploaded with POST /api/storage/upload first; external URLs are rejected (422).
- input_audio_duration > 61 is rejected (422; the limit is 60s plus 1s tolerance).
- No prompt, duration or aspect ratio: the output length equals the audio and the resolution follows the source.
Example request body
{
"mode": "video-lip-sync",
"generation_type": "lip-sync",
"videos": [
"https://cdn.veevid.ai/video-lip-sync/clip.mp4"
],
"audios": [
"https://cdn.veevid.ai/video-lip-sync/audio/voice.mp3"
],
"input_audio_duration": 12.4,
"lip_sync_mode": "lite",
"lip_sync_separate_vocal": false,
"lip_sync_open_scenedet": true,
"lip_sync_align_audio": true,
"lip_sync_align_audio_reverse": false,
"lip_sync_templ_start_seconds": 0
}Google Veo 3.1 mode: veo3
video-extend
Extends a Veo 3.1 video you generated earlier, identified by its task ID.
| Parameter | Values & notes |
|---|---|
| taskIdstring · requiredaffects price | The provider_task_id of one of your own earlier Veo 3.1 generations (returned when it was created and by the status endpoint). |
| promptstring · required | What happens in the extension. Required. |
| generation_modestringaffects pricedefault "fast" | "lite", "fast", "quality"Extension quality tier. |
| watermarkstring | Optional custom watermark text, <=200 chars. |
| durationstringdefault "8" | "4", "6", "8"Must be 4, 6 or 8 if sent; send "8". |
| aspect_ratiostringdefault "original" | "original"Send "original" (the value is not used). |
Rules
- taskId is required (400). It must match one of your generations (404 'Task not found') owned by you (403).
- The source must be a Veo 3.1 or Grok Imagine video (400). A Grok Imagine taskId is run and priced as a Grok Imagine extend.
- 1080p and 4k Veo sources can be extended.
- When taskId belongs to a Grok Imagine video, the request is run and billed as a Grok Imagine extend: send duration "6" or "10" (or omit it), not "8".
Example request body
{
"mode": "veo3",
"generation_type": "video-extend",
"taskId": "veo_task_abc123",
"prompt": "The drone keeps flying over the city at night",
"generation_mode": "fast",
"duration": "8",
"aspect_ratio": "original"
}Grok Imagine mode: grok-imagine
video-extend
Extends a Grok Imagine video you generated earlier by 6 or 10 seconds from a chosen point.
| Parameter | Values & notes |
|---|---|
| taskIdstring · requiredaffects price | The provider_task_id of one of your own earlier Grok Imagine generations. |
| promptstring · required | What happens in the extension. Required. |
| durationstringaffects pricedefault "6" | "6", "10"Seconds to add. |
| extend_atnumberdefault 0 | Start position in seconds (>= 0); 0 extends from the end of the video. |
| aspect_ratiostringdefault "original" | "original"Send "original" (the value is not used). |
Rules
- duration must be '6' or '10' (400 'Grok Imagine Extend only supports 6s or 10s durations').
- taskId is required, must exist (404) and belong to you (403). The source must be a Grok Imagine or Veo 3.1 video; a Veo 3.1 taskId is run and priced as a Veo 3.1 extend.
Example request body
{
"mode": "grok-imagine",
"generation_type": "video-extend",
"taskId": "grok_task_abc123",
"prompt": "The knight raises his sword as the dragon lands",
"duration": "6",
"extend_at": 0,
"aspect_ratio": "original"
}Code Examples
Generate an image (curl)
curl -X POST https://veevid.ai/api/generate-video \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"mode": "nano-banana-2",
"generation_type": "text-to-image",
"prompt": "A cozy coffee shop interior, morning light, watercolor style",
"aspect_ratio": "1:1",
"video_quality": "1K"
}'To edit your own photo, upload it to /api/storage/upload first and send "generation_type": "image-to-image" with "images": ["UPLOADED_URL"].
Text to speech (curl)
curl -X POST https://veevid.ai/api/generate-video \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"mode": "eleven-v4-turbo",
"generation_type": "text-to-speech",
"prompt": "Welcome to Veevid. Let us make something amazing today.",
"voice": "Rachel"
}'Poll the status endpoint; the audio file is in video_url.
Python
import requests
import time
API_KEY = "vv_sk_your_key_here"
BASE = "https://veevid.ai/api"
HEADERS = {"Authorization": f"Bearer {API_KEY}"}
# 1. Upload a first frame
with open("first_frame.jpg", "rb") as f:
upload = requests.post(
f"{BASE}/storage/upload",
headers=HEADERS,
files={"file": f},
data={"folder": "video-generator-veo3"},
).json()
body = {
"mode": "veo3",
"generation_type": "image-to-video",
"prompt": "The camera slowly pushes in, leaves rustle in the wind",
"images": [upload["url"]],
"video_quality": "standard",
}
# 2. Quote
quote = requests.post(f"{BASE}/quote", headers=HEADERS, json=body).json()
print(f"Cost: {quote['required_credits']} credits")
# 3. Generate
result = requests.post(f"{BASE}/generate-video", headers=HEADERS, json=body).json()
gen_id = result["generation_id"]
# 4. Poll
while True:
status = requests.get(
f"{BASE}/video-generation/{gen_id}/status", headers=HEADERS
).json()
if status["status"] == "completed":
print(f"Ready: {status['video_url']}")
break
if status["status"] == "failed":
print(f"Failed: {status['error_message']}")
break
time.sleep(10)JavaScript / TypeScript
const API_KEY = "vv_sk_your_key_here";
const BASE = "https://veevid.ai/api";
const headers = {
Authorization: `Bearer ${API_KEY}`,
"Content-Type": "application/json",
};
async function generate(body: Record<string, unknown>) {
// 1. Quote
const quote = await fetch(`${BASE}/quote`, {
method: "POST",
headers,
body: JSON.stringify(body),
}).then((r) => r.json());
console.log(`Cost: ${quote.required_credits} credits`);
// 2. Generate
const result = await fetch(`${BASE}/generate-video`, {
method: "POST",
headers,
body: JSON.stringify(body),
}).then((r) => r.json());
// 3. Poll
while (true) {
const status = await fetch(
`${BASE}/video-generation/${result.generation_id}/status`,
{ headers }
).then((r) => r.json());
if (status.status === "completed") return status.video_url;
if (status.status === "failed") throw new Error(status.error_message);
await new Promise((r) => setTimeout(r, 10000));
}
}
await generate({
mode: "kling-3",
generation_type: "text-to-video",
prompt: "A drone shot flying over a neon-lit city at night",
duration: "5",
generate_audio: true,
});OpenClaw Agent
Install the Veevid skill:
npx clawhub@latest install veevidThen tell your agent: "Generate a 10-second product video with Kling 3.0"
The agent handles quoting, confirmation, generation, and polling automatically.
Error Codes
| Code | Meaning | What to Do |
|---|---|---|
| 400 | Invalid parameters | Check allowed values for the model |
| 401 | Invalid or missing API key | Verify your key at /settings/api-keys |
| 402 | Insufficient credits | Top up at /pricing |
| 403 | Account suspended | Contact support |
| 404 | Generation not found | Check the generation_id |
| 409 | Conflict: the input video is longer than the declared length (character swap), or a draft is not ready / already upgraded | Nothing was charged; read the error — for character swap, resend with the measured_duration from the response |
| 410 | The draft has expired | Generate again |
| 422 | Parameter combination not supported by the model | Read the error message and the model rules |
| 500 | Server error | Wait a few seconds and retry |
| 502 | A draft upgrade could not be submitted (credits refunded) | Wait a few seconds and retry |
Error response format:
{
"error": "Insufficient credits. You need 140 credits but only have 12.",
"required": 140,
"balance": 12
}Limits
| Limit | Value |
|---|---|
| API keys per account | 5 active keys |
| Prompt length | Up to 20,000 characters (model-specific) |
| Upload size | Per folder: 10MB for images by default, larger for video/audio folders, and at most 100MB per upload request; 512MB via presigned upload (clean audio) |
| Request rate | No fixed per-key rate limit today — your credit balance is the limit. Please poll status every 5–10 seconds rather than continuously. |