Voice Clone
In Voice Clone, enter your text and upload or record a clear 3–10 second reference clip. Add its transcript if available, then generate and download 24 kHz audio.
OmniVoice · k2-fsa
Turn text into speech with OmniVoice by k2-fsa. Try the official demo for zero-shot voice cloning and voice design across 646 languages.
Loading the demo…
The public ZeroGPU demo has shared GPU quotas and queues. Hugging Face may require sign-in or ask you to try later when the allowance is used. Wan2.Video does not charge for access to this page.
In Voice Clone, enter your text and upload or record a clear 3–10 second reference clip. Add its transcript if available, then generate and download 24 kHz audio.
In Voice Design, create a voice without a reference recording. Combine gender, age, pitch, accent or dialect, and whisper attributes. Voice design works most reliably in Chinese and English.
Add [laughter] or [sigh] to the text for non-verbal sounds. Correct Chinese pronunciation with tone-number pinyin and English pronunciation with bracketed CMU phonemes.
The model has about 0.6B parameters. The team reports RTF as low as 0.025 (about 40× real time); public demo speed also depends on the queue and hardware.
The public ZeroGPU demo has shared GPU quotas and queues. Hugging Face may require sign-in or ask you to try later when the allowance is used. Wan2.Video does not charge for access to this page.
The selected Hugging Face Space receives your inputs and runs the model. Its maintainer controls processing and retention. The surrounding guide is translated; the demo uses the languages provided by its maintainer.
Reload the workspace or choose a backup server if available. If your GPU quota is used up, wait for it to reset or try a related tool.