Free audio-driven avatar video

OmniAvatar Free Online Demo

Combine one character image with a short voice clip to create a speaking avatar with synchronized lips, expressions and adaptive body motion—powered by public Hugging Face ZeroGPU Spaces.

Real ZeroGPU Image + Audio 5-second demo Prompt control
Create an avatar

Official OmniAvatar showcase

Audio controls speech while the prompt guides behavior, emotion and scene.

Create a talking avatar now

Use the recommended Space first or switch manually to the backup. ZeroGPU Spaces can sleep, wake, queue or change status at any time.

Checking demo availability…

alexnasa · OmniAvatar

ZeroGPU configuration checked August 26, 2026

ZeroGPU · dynamic
Waking or loading the OmniAvatar Space…

A portrait, a voice and a direction

The public demo turns three simple inputs into a short avatar performance. Clear references and concise prompts work best.

1

Choose a portrait

Upload a clear character image. The Space preprocesses and fits it to the selected orientation.

2

Add a short voice

Upload audio, record with your microphone or use a sample. The free workflow currently keeps the first five seconds.

3

Direct the performance

Keep the adaptive prompt or describe expression, body movement and background, then choose the number of steps.

More than a lip-sync tool

OmniAvatar conditions the face, body and surrounding scene together, making it useful for presenters, podcasts, singing and character performances.

Audio-synchronized speech

Wav2Vec audio conditioning aligns mouth movement and facial expression with the supplied voice.

Adaptive body animation

The character can gesture and move beyond the face instead of remaining a static talking head.

Text-directed behavior

Prompt text can guide actions, emotion, camera framing and optional background details.

Upload or record audio

The community UI accepts audio files, browser microphone recording and cached voice examples.

ZeroGPU is free, but not unlimited

Hugging Face meters shared GPU time by account tier. Remaining quota also affects queue priority, and xlarge tasks cost twice as much as large tasks.

AccountIncluded daily GPUQueue priority
Anonymous2 minLow
HF Free5 minMedium
HF PRO40 minHighest

Why the public demo keeps five seconds

The current inference function trims uploaded audio to its first five seconds before generation. This keeps a 14B avatar job short enough for shared ZeroGPU capacity.

Current alexnasa code requests a 96 GB xlarge GPU. At the default four steps, its estimated reservation and 2× quota multiplier can exceed the anonymous daily allowance. This is an estimate from the live code and can change.

The model really runs on shared Hugging Face GPUs

Unlike API-backed demos, these Spaces load the open OmniAvatar stack and request ZeroGPU only for the inference function.

Your browser

Public HF Space

ZeroGPU

Avatar MP4

Public Gradio API

Every public Gradio Space exposes an API, but OmniAvatar uses several preprocessing and state steps. Read the live schema with view_api() instead of hard-coding a generic /predict call.

Client.connect("alexnasa/OmniAvatar")

The current generation event is based on infer_scene and accepts image, audio, prompt, orientation, steps and session state. Treat it as an experimental endpoint.

Production and self-hosting

The official 14B and 1.3B weights and inference code use Apache-2.0, but no Hugging Face Inference Provider currently hosts the 14B model. Stable products need their own deployment and capacity plan.

14B 1.3B Apache-2.0 480p

OmniAvatar demo FAQ

Is the OmniAvatar demo free?

The public ZeroGPU interface is free within Hugging Face quotas. A new 14B generation is expensive, so anonymous quota may be insufficient and queues can be long.

Can I generate without signing in?

The Space can be opened without signing in. Generation depends on the anonymous quota remaining and the requested GPU duration; Hugging Face may ask you to sign in for more quota.

Can I upload a long audio track?

You can select a longer file, but the current public inference path intentionally processes only the first five seconds for a fresh generation.

Can I use this as a free production API?

The public Gradio API is useful for experiments, but Space status, queue, code and daily quota can change. It is not a production uptime commitment.

Where are my image and voice processed?

They leave Wan2.Video and are processed by the selected third-party Hugging Face Space. Use only material you are authorized to upload.