Skip to main content
POST
cURL
Provide at least one reference asset. The per-media limits are defined by the API schema.

Asset and Scene Limits

Seedance 2.0 reference-material mode supports 9 reference images, 3 reference videos, and 3 audio clips. It is a full-capability route that supports real-person content and faces; materials are submitted to asset review by default (real_person=true), and setting real_person=false passes materials through without review when they are confirmed face-free.
  • Text prompt: For Chinese, keep it under 500 characters; for English, under 1,000 words.
  • Image input: Supports jpeg, png, webp, bmp, tiff, gif, heic, and heif; each image under 30 MB; aspect ratio (0.4, 2.5); width and height (300px, 6000px).
  • Multimodal reference scenes: No more than 9 images.
  • Video input: Supports mp4 and mov; each video 2–15 seconds; up to 3 reference videos; total video duration must not exceed 15 seconds; each video under 50 MB; frame rate [24, 60].
  • Audio input: Supports wav and mp3; each clip 2–15 seconds; up to 3 reference audio clips; total audio duration must not exceed 15 seconds; each clip under 15 MB.
  • Audio input restriction: Audio cannot be provided on its own; at least one reference image or video is required.
Note: Video input has a lower unit price because the billing formula differs: without video = unit price × output volume; with video = unit price × (input volume + output volume).

Real-person review (real_person)

When a request carries reference materials, the boolean real_person field controls whether materials are submitted to face review first:
  • real_person: true (default): materials are submitted to the asset review service first, and the task is created with assetId://{assetId} references; the task waits during review and fails if review rejects the materials.
  • real_person: false: review is skipped and material URLs are passed through to task creation as-is.
Set it to false only when the materials are confirmed face-free to save the review round trip; keep the default true for materials containing real people.

Asset library references

Besides publicly accessible URLs, the image_url / image_urls, video_urls, and audio_url / audio_urls fields also accept asset library references in the form assetId://{assetId}. Asset IDs are shared across the platform: any account may reference a known assetId, regardless of who submitted it — the same sharing semantics as a public URL. Only ACTIVE assets pass validation. Submit and track assets via asset upload and asset detail.

Authorizations

Authorization
string
header
required

Use your APIPod API key as a Bearer token in the Authorization header.

Body

application/json
model
string
required

Public APIPod model ID.

Allowed value: "seedance-2.0-r2v"
prompt
string
required

Generation or editing instructions, up to 30000 characters. It is recommended to keep the prompt to no more than 500 Chinese characters or 1,000 English words. Lengthy text will lead to scattered information, and the model may ignore details and only focus on key points, resulting in missing elements in the generated video.

Maximum string length: 30000
image_urls
string[]
video_urls
string[]
audio_urls
string[]

Please note that audio files cannot be uploaded alone; at least one reference video or image must be included.

duration
enum<integer>
default:5
Available options:
4,
5,
6,
7,
8,
9,
10,
11,
12,
13,
14,
15
Required range: 4 <= x <= 15
resolution
enum<string>
default:720p
Available options:
480p,
720p,
1080p,
4k
aspect_ratio
enum<string>
default:adaptive
Available options:
adaptive,
1:1,
3:4,
4:3,
9:16,
16:9,
21:9
watermark
boolean
default:false

After enabling, the model will independently decide whether to search Internet content (such as products, weather, etc.) based on the user's prompt. This can improve the timeliness of generated videos but will also introduce a certain degree of latency.

generate_audio
boolean
default:true

Controls whether the generated video contains audio synchronized with the visuals.

return_last_frame
boolean
default:false

Whether to return the last frame image of the generated video. Using this parameter enables the generation of multiple consecutive videos: taking the last frame of the previous generated video as the first frame of the next video task to quickly generate multiple consecutive videos, and the first item in the returned result is this image.

real_person
boolean
default:true

Whether the reference materials contain real people (human faces). Defaults to true: materials are submitted to the asset review service and replaced with assetId:// references before task creation. Set to false only when the materials are confirmed face-free to skip review and pass material URLs through as-is.

Response

200 - application/json

Task accepted

Standard response for task submission

code
integer
required
data
Task Submit Data · object
required

Data content for task submission

message
string
required