Real-time AI avatars,
streamed live.
Create a lip-synced, conversational talking-head from your own photo, image or video — then stream it live to any app over WebRTC. Tone-aware, self-hostable, and resident on your own Sydney GPUs.
Everything you need for a live, believable avatar
From source media to a streaming, tone-aware presenter in a few lines of code — with the models running where your data already lives.
Upload one headshot and get a natural, gently-moving talking head — idle motion, blinks and all — with consent gating built in.
Turn a rendered character or brand portrait into a presenter. Any face-forward image becomes a drivable avatar.
Shoot a few seconds of footage for the most expressive, true-to-life result — bundle tone variants under one avatar.
Frames are generated on a GPU and streamed live over WebRTC, audio and video kept in lock-step. Barge-in stops mid-sentence.
Replies are spoken by a tone-matched face; emotion events fire as the avatar speaks so your UI can react in real time.
Run a full model → TTS → avatar turn with send_message(), or speak exact words with talk() and incremental talk streams.
Three steps to a live avatar
Upload a photo, image or short video in the portal. We prepare a drivable, natural-motion avatar — gated behind consent.
Point ZeliClient at your avatar server, call connect(), and receive synchronized audio + video frames over WebRTC.
send_message() for a full conversational turn, or talk() to speak exact words. Emotion and transcript events stream back.
Your avatars never leave your infrastructure
Unlike HeyGen, Synthesia or D-ID, Zeli Avatar is self-hosted. The models run on your GPUs in Sydney, so faces, footage and transcripts stay within Australian data residency — the difference between a demo and something compliance will actually sign off.
How the pipeline worksBring a face to your product.
Create your first avatar and stream it live today. No offshore cloud, no per-seat lock-in — just your GPUs and a few lines of Python.