Seedance 2.5 Multimodal Reference Guide: 50 Assets, One Shot

📅 Aug 10, 2026 👁 1 views #2026 #ai video #Bytedance #Multimodal #Reference #Seedance 2 5
Seedance 2.5 Multimodal Reference Guide: 50 Assets, One Shot

The single biggest lever in Seedance 2.5 is not the prompt — it is the reference tray. One job can ingest up to 50 assets: 30 images, 10 video clips and 10 audio clips, plus white-model blockouts and green-screen footage. Used well, references turn AI video from a slot machine into a controllable production tool.

What each modality controls

  • Images (max 30) — character identity, wardrobe, props, locations, color palette, lighting references. Three to five angles per character is the sweet spot for consistency.
  • Videos (max 10) — camera moves, motion style, choreography, pacing. Give it a dolly path and it reproduces the move on new content.
  • Audio (max 10) — music tracks to perform to, voice timbre for lip-synced delivery, ambience to match.

Label every asset’s job

The most common failure is blending: you upload a lighting reference and a character reference, and the model copies the lighting reference’s face. Fix it by stating roles explicitly in the prompt: “Images 1-4 define the character. Image 5 defines the location. Image 6 is lighting mood only. Video 1 defines the camera path.” One asset, one purpose.

White-model and green-screen input

New in 2.5: hand the model a white-model blockout (clay render) and it fills in photoreal materials, light and motion. A car-assembly blockout becomes a realistic factory sequence; an architecture massing model becomes a finished flythrough. Official Maya and Blender plugins send blockouts straight to Seedance, which slots directly into previz and industrial pipelines — XCMG uses this for SOP training videos and Xpeng for design visualization. Green-screen footage works as a reference too: shoot an actor on green, then let the model build the world around the performance.

[INSERT_IMAGE: screenshot of a Blender viewport with a white-model blockout next to the Seedance 2.5 photoreal render of the same scene]

Recipes that work

  • Brand ad: 4 product photos + 2 location images + 1 camera-move video + 1 music track → a 30s spot with matching product, place, motion and beat.
  • Short drama: 5 images per character + 3 set photos → consistent cast across separate generations, then stitched.
  • Music performance: audio track + performer photos → finger and lip movement synced to the actual recording.
  • Industrial training: CAD/Blender blockout + 1 real photo of the machine → realistic procedure video without a shoot.

Limits to respect

  • Video-conditioned jobs are heavier and scale with input length — keep reference videos as short as the move allows.
  • More assets is not better past the point of coverage; 8-12 well-labeled assets usually beat 30 noisy ones.
  • Reference video dictates motion, not content — expect the model to borrow the move, not the subject.

Next steps

Put references to work with the prompt guide, and learn surgical fixes in the video editing guide.