OmniVoice · k2-fsa

Free OmniVoice Text-to-Speech & Voice Cloning

Turn text into speech with OmniVoice by k2-fsa. Try the official demo for zero-shot voice cloning and voice design across 646 languages.

Try the tool 646 languages24 kHz

Try the tool

OmniVoice
k2-fsa/OmniVoice

Loading the demo…

The public ZeroGPU demo has shared GPU quotas and queues. Hugging Face may require sign-in or ask you to try later when the allowance is used. Wan2.Video does not charge for access to this page.

How to use it

Voice Clone

In Voice Clone, enter your text and upload or record a clear 3–10 second reference clip. Add its transcript if available, then generate and download 24 kHz audio.

Voice Design

In Voice Design, create a voice without a reference recording. Combine gender, age, pitch, accent or dialect, and whisper attributes. Voice design works most reliably in Chinese and English.

Set up the result

Add [laughter] or [sigh] to the text for non-verbal sounds. Correct Chinese pronunciation with tone-number pinyin and English pronunciation with bracketed CMU phonemes.

The model has about 0.6B parameters. The team reports RTF as low as 0.025 (about 40× real time); public demo speed also depends on the queue and hardware.

Common questions

Is the demo free to use?

The public ZeroGPU demo has shared GPU quotas and queues. Hugging Face may require sign-in or ask you to try later when the allowance is used. Wan2.Video does not charge for access to this page.

Where are my uploads processed?

The selected Hugging Face Space receives your inputs and runs the model. Its maintainer controls processing and retention. The surrounding guide is translated; the demo uses the languages provided by its maintainer.

What if the demo does not load?

Reload the workspace or choose a backup server if available. If your GPU quota is used up, wait for it to reset or try a related tool.

Official resources