Native 30-second storytelling
Generate an integer duration from 2 to 30 seconds, or use smart duration so the model selects a suitable length.
A practical hub for Alibaba's all-in-one video model: understand the 30-second multimodal workflow, build requests, estimate cost, write stronger prompts, and turn documents into production-ready plans.
Wan 3.0 replaces a menu of task-specific endpoints with one all-in-one model. The media type and the intent in your prompt route text-to-video, first-frame animation, first-and-last-frame control, multimodal reference, document understanding, editing, or extension.
The release expands both the creative canvas and the amount of source material a single request can understand.
Generate an integer duration from 2 to 30 seconds, or use smart duration so the model selects a suitable length.
Combine up to 10 images, five video clips, and five audio clips while addressing each asset by type and order.
Use a file or a public web page as structured source material for explainers, ads, presentations, and narrative video.
Reference guidance targets faces, props, spaces, style, and relationships across a longer sequence.
Dialogue, effects, ambience, and music can be generated with the picture; disabling audio does not change the price.
Modify visuals, plot, dialogue, or extend a supplied video through prompt intent without choosing a separate model family.
The same model ID routes work from the media types supplied in input.media and from clear action words in the prompt.
| Task | How it is triggered | Best practice |
|---|---|---|
| Text to video | Prompt only | Set ratio, resolution, duration, and audio directly. |
| First frame | first_frame | Use adaptive ratio to preserve the image composition. |
| First and last frame | first_frame + last_frame | Describe the transition and motion between the two anchors. |
| Omni reference | reference_image / video / audio | Refer to Image 1, Video 1, and Audio 1 by their type-specific order. |
| File or web page | file or link | One public file or page per request; file and link are mutually exclusive. |
| Edit or extend video | reference_video + edit intent | Use adaptive ratio; smart duration is useful for preserving or extending timing. |
Each page solves a different planning decision without consuming model quota.
The launch is substantial, but the operational boundary matters as much as the headline features.
Official generation requires a regional Model Studio workspace and API key; the resource pages here do not proxy paid requests.
The 30-second free quota applies only in the Beijing region, combines input and output video seconds, and expires after 90 days under the official rules.
Alibaba notes that audio texture and on-screen text accuracy are still improving, so review dialogue, labels, charts, and typography before delivery.
Generated result URLs and task IDs are available for 24 hours; production systems should download outputs to durable storage.
The model is pay-as-you-go. Beijing accounts may receive a one-time 30-second quota valid for 90 days; the international region currently has no official free quota. These planning tools are free and make no model call.
Yes. Without input video, duration accepts any integer from 2 to 30. With input video, the combined input-video and output-video duration cannot exceed 30 seconds.
Yes. Official formats include common office documents, PDF, text, Markdown, Keynote, Pages, and Numbers. One file up to 100 MB and up to 50 validated pages is accepted.
Yes. Audio is enabled by default. Prompts can direct dialogue, voice character, effects, ambience, and music; audio on or off does not change billing.
Facts and prices on these pages are dated snapshots from Alibaba Cloud and Model Studio. Confirm the console before a paid production run.