Longer single takes
Generate up to 30 seconds in one pass. Complete a narrative beat, product demo or social clip without stitching fragments together.
Turn prompts, first and last frames, and up to 20 reference assets into expressive AI video with native audio. Generate 2 to 30 second clips in one pass with adaptive framing and cinematic camera control.
Turn prompts, first and last frames, and up to 20 reference assets into expressive AI video with native audio. Generate 2 to 30 second clips in one pass with adaptive framing and cinematic camera control.
From a first idea to a finished clip, the workflow stays expressive and practical.
Generate up to 30 seconds in one pass. Complete a narrative beat, product demo or social clip without stitching fragments together.
Feed up to 10 reference images, 5 reference videos and 5 audio files. Characters, styles, camera moves and sound direction stay aligned with your assets.
Pin the opening and closing frames of your shot. Wan 3.0 handles the motion in between while respecting your composition.
Dialogue, ambience and effects are generated with the picture, so lip movement and sound land in sync — no separate sound pass.
Describe push-ins, orbits and tracking shots in natural language. Motion stays stable and perspective holds through the move.
Adaptive framing follows your source media, or lock 16:9, 9:16, 1:1, 4:3 or 3:4 for the channel you target.
Steps from idea to downloadable result.
Describe the subject, action, camera and sound. Reference your attached assets with @Image1, @Video1 or @Audio1 tokens in the prompt.
Upload a first frame with an optional last frame, or attach reference images, videos and audio. Choose one workflow — frames and references cannot be combined.
Pick resolution, aspect ratio and duration from 2 to 30 seconds. The studio estimates your cost live as you configure the shot.
Start the generation and watch the preview panel. Your clip is saved to My Creation, so you can close the page while it renders.
Refine the prompt, adjust duration or swap references. Longer takes benefit from beat-by-beat direction in the prompt.
Download the finished MP4 and ship it — vertical for shorts, widescreen for the web, square for feeds.
Answers to common questions about inputs, settings and billing.
Wan 3.0 is Alibaba’s next-generation all-in-one video model. It generates video with native audio from text prompts, first and last frame images, and multimodal references including images, video clips and audio files.
You can start from a text prompt alone, add a first frame (optionally with a last frame), or use reference media: up to 10 reference images, 5 reference videos and 5 reference audio files in a single request.
Clips run from 2 to 30 seconds in one generation. When reference videos are attached, the total reference video duration plus your output duration must stay within 30 seconds.
Yes. Wan 3.0 produces a native audio track — dialogue, ambience and sound effects — in the same pass as the picture. You can switch audio off if you only need silent footage.
No. The frames workflow and the reference-media workflow are mutually exclusive, following the provider API rules. Reference audio also requires at least one reference image or video.
Cost is a flat per-second rate set by your chosen resolution (480p, 720p or 1080p) multiplied by the output duration. Audio does not change the price, and reference media does not add to the billed seconds.
Start with a prompt or upload a frame and create a 30-second story with Wan 3.0 today.
Start creating with Wan 3.0