Choose a portrait
Upload a clear character image. The Space preprocesses and fits it to the selected orientation.
Free audio-driven avatar video
Combine one character image with a short voice clip to create a speaking avatar with synchronized lips, expressions and adaptive body motion—powered by public Hugging Face ZeroGPU Spaces.
Official OmniAvatar showcase
Audio controls speech while the prompt guides behavior, emotion and scene.
Use the recommended Space first or switch manually to the backup. ZeroGPU Spaces can sleep, wake, queue or change status at any time.
Checking demo availability…
alexnasa · OmniAvatar
ZeroGPU configuration checked August 26, 2026
The public demo turns three simple inputs into a short avatar performance. Clear references and concise prompts work best.
Upload a clear character image. The Space preprocesses and fits it to the selected orientation.
Upload audio, record with your microphone or use a sample. The free workflow currently keeps the first five seconds.
Keep the adaptive prompt or describe expression, body movement and background, then choose the number of steps.
OmniAvatar conditions the face, body and surrounding scene together, making it useful for presenters, podcasts, singing and character performances.
Wav2Vec audio conditioning aligns mouth movement and facial expression with the supplied voice.
The character can gesture and move beyond the face instead of remaining a static talking head.
Prompt text can guide actions, emotion, camera framing and optional background details.
The community UI accepts audio files, browser microphone recording and cached voice examples.
Hugging Face meters shared GPU time by account tier. Remaining quota also affects queue priority, and xlarge tasks cost twice as much as large tasks.
| Account | Included daily GPU | Queue priority |
|---|---|---|
| Anonymous | 2 min | Low |
| HF Free | 5 min | Medium |
| HF PRO | 40 min | Highest |
The current inference function trims uploaded audio to its first five seconds before generation. This keeps a 14B avatar job short enough for shared ZeroGPU capacity.
Current alexnasa code requests a 96 GB xlarge GPU. At the default four steps, its estimated reservation and 2× quota multiplier can exceed the anonymous daily allowance. This is an estimate from the live code and can change.
Unlike API-backed demos, these Spaces load the open OmniAvatar stack and request ZeroGPU only for the inference function.
Your browser
Public HF Space
ZeroGPU
Avatar MP4
Every public Gradio Space exposes an API, but OmniAvatar uses several preprocessing and state steps. Read the live schema with view_api() instead of hard-coding a generic /predict call.
Client.connect("alexnasa/OmniAvatar") The current generation event is based on infer_scene and accepts image, audio, prompt, orientation, steps and session state. Treat it as an experimental endpoint.
The official 14B and 1.3B weights and inference code use Apache-2.0, but no Hugging Face Inference Provider currently hosts the 14B model. Stable products need their own deployment and capacity plan.
The public ZeroGPU interface is free within Hugging Face quotas. A new 14B generation is expensive, so anonymous quota may be insufficient and queues can be long.
The Space can be opened without signing in. Generation depends on the anonymous quota remaining and the requested GPU duration; Hugging Face may ask you to sign in for more quota.
You can select a longer file, but the current public inference path intentionally processes only the first five seconds for a fresh generation.
The public Gradio API is useful for experiments, but Space status, queue, code and daily quota can change. It is not a production uptime commitment.
They leave Wan2.Video and are processed by the selected third-party Hugging Face Space. Use only material you are authorized to upload.
Explore motion transfer, sound generation and other character-video tools available on Wan2.Video.
Transfer motion or replace a character using a reference video.
ExploreLearn about synchronized audio generation for video.
ExploreCreate synchronized audio-video clips through public Spaces.
ExploreTry a browser-based face replacement workflow.
Explore