Native audio-video

Native audio-video means one sampler emits picture and a synchronized soundtrack (speech, effects, music). That capability is advertised on Wan 2.5 and later product APIs. It is not a property of the Wan2.2-T2V/I2V open weights.

Wan 2.5API-onlysync
On this page 5
Wan2.2 open silent video versus Wan 2.5+ native audio-video APIs
Keep the product line and the open checkpoints in separate boxes.

What 2.2 open models emit

  • T2V-A14B / I2V-A14B / TI2V-5B — silent video.
  • S2V-14B — you bring the audio; the model animates.
  • Animate-14B — motion transfer; audio not generated.

What 2.5+ products add

Public 2.5 write-ups (fal, DashScope-style APIs, aggregators) describe joint generation of dialogue, ambience, and music, 1080P, and clips longer than the 2.2 ~5 s recipes. Later 2.6/2.7/3.0 keep native AV and add multi-shot or reference controls. None of those are published as Wan2.2-style GitHub checkpoints.

Open-weight substitutes

To stay on Apache-2.0 / public weights: generate silent Wan2.2 video, then ThinkSound (or another V2A) for foley; or drive S2V with CosyVoice for talking heads. That is two (or three) models, not native AV.

FAQ

Will a 2.2 GGUF magically add audio?
No. Quantization changes weight dtype, not the output modality.
Does S2V count as native AV?
No. Native AV synthesizes sound. S2V consumes sound.

Primary sources

Checked against Wan-Video/Wan2.2 README · Aug 28, 2026

© 2026 wan2.video