Seedance 2.5 Multimodal Reference Guide: 50 Assets, One Shot

The single biggest lever in Seedance 2.5 is not the prompt — it is the reference tray. One job can ingest up to 50 assets: 30 images, 10 video clips and 10 audio clips, plus white-model blockouts and green-screen footage. Used well, references turn AI video from a slot machine into a controllable production tool.
What each modality controls
- Images (max 30) — character identity, wardrobe, props, locations, color palette, lighting references. Three to five angles per character is the sweet spot for consistency.
- Videos (max 10) — camera moves, motion style, choreography, pacing. Give it a dolly path and it reproduces the move on new content.
- Audio (max 10) — music tracks to perform to, voice timbre for lip-synced delivery, ambience to match.
Label every asset’s job
The most common failure is blending: you upload a lighting reference and a character reference, and the model copies the lighting reference’s face. Fix it by stating roles explicitly in the prompt: “Images 1-4 define the character. Image 5 defines the location. Image 6 is lighting mood only. Video 1 defines the camera path.” One asset, one purpose.
White-model and green-screen input
New in 2.5: hand the model a white-model blockout (clay render) and it fills in photoreal materials, light and motion. A car-assembly blockout becomes a realistic factory sequence; an architecture massing model becomes a finished flythrough. Official Maya and Blender plugins send blockouts straight to Seedance, which slots directly into previz and industrial pipelines — XCMG uses this for SOP training videos and Xpeng for design visualization. Green-screen footage works as a reference too: shoot an actor on green, then let the model build the world around the performance.
[INSERT_IMAGE: screenshot of a Blender viewport with a white-model blockout next to the Seedance 2.5 photoreal render of the same scene]
Recipes that work
- Brand ad: 4 product photos + 2 location images + 1 camera-move video + 1 music track → a 30s spot with matching product, place, motion and beat.
- Short drama: 5 images per character + 3 set photos → consistent cast across separate generations, then stitched.
- Music performance: audio track + performer photos → finger and lip movement synced to the actual recording.
- Industrial training: CAD/Blender blockout + 1 real photo of the machine → realistic procedure video without a shoot.
Limits to respect
- Video-conditioned jobs are heavier and scale with input length — keep reference videos as short as the move allows.
- More assets is not better past the point of coverage; 8-12 well-labeled assets usually beat 30 noisy ones.
- Reference video dictates motion, not content — expect the model to borrow the move, not the subject.
Next steps
Put references to work with the prompt guide, and learn surgical fixes in the video editing guide.