Skip to main content
POST
cURL
Provide at least one reference asset. The per-media limits are defined by the API schema.

Asset and Scene Limits

Seedance 2.0 Fast reference-material mode supports 9 reference images, 3 reference videos, and 3 audio clips. It is a full-capability route that supports real-person content and faces; materials are submitted to asset review before task creation.
  • Text prompt: For Chinese, keep it under 500 characters; for English, under 1,000 words.
  • Image input: Supports jpeg, png, webp, bmp, tiff, gif, heic, and heif; each image under 30 MB; aspect ratio (0.4, 2.5); width and height (300px, 6000px).
  • Multimodal reference scenes: No more than 9 images.
  • Video input: Supports mp4 and mov; each video 2–15 seconds; up to 3 reference videos; total video duration must not exceed 15 seconds; each video under 50 MB; frame rate [24, 60].
  • Audio input: Supports wav and mp3; each clip 2–15 seconds; up to 3 reference audio clips; total audio duration must not exceed 15 seconds; each clip under 15 MB.
  • Audio input restriction: Audio cannot be provided on its own; at least one reference image or video is required.
Note: Video input has a lower unit price because the billing formula differs: without video = unit price × output volume; with video = unit price × (input volume + output volume).

Authorizations

Authorization
string
header
required

Use your APIPod API key as a Bearer token in the Authorization header.

Body

application/json
model
string
required

Public APIPod model ID.

Allowed value: "seedance-2.0-fast-r2v"
prompt
string
required

Generation or editing instructions, up to 30000 characters. It is recommended to keep the prompt to no more than 500 Chinese characters or 1,000 English words. Lengthy text will lead to scattered information, and the model may ignore details and only focus on key points, resulting in missing elements in the generated video.

Maximum string length: 30000
duration
enum<integer>
default:5
Available options:
4,
5,
6,
7,
8,
9,
10,
11,
12,
13,
14,
15
Required range: 4 <= x <= 15
audio_urls
string[]

Please note that audio files cannot be uploaded alone; at least one reference video or image must be included.

image_urls
string[]
resolution
enum<string>
default:720p
Available options:
480p,
720p
video_urls
string[]

After enabling, the model will independently decide whether to search Internet content (such as products, weather, etc.) based on the user's prompt. This can improve the timeliness of generated videos but will also introduce a certain degree of latency.

aspect_ratio
enum<string>
default:adaptive
Available options:
adaptive,
1:1,
3:4,
4:3,
9:16,
16:9,
21:9
generate_audio
boolean
default:true

Controls whether the generated video contains audio synchronized with the visuals.

Response

200 - application/json

Task accepted

Standard response for task submission

code
integer
required
data
Task Submit Data · object
required

Data content for task submission

message
string
required