Lip sync driven by your own audio
Kling AI Avatar reads your uploaded voice track and animates the mouth and face to match it, so every word lines up with the speech you recorded.
A portrait and a voice become your spokesperson.
Sign in and your generations will appear here.
Browse the gallery for ideasFAQ
One portrait image (JPG, PNG or WebP, up to 10 MB) and one voice file (MP3 or WAV, up to about 30 seconds and 20 MB). A text prompt is optional.
No. Kling AI Avatar uses the audio you upload and adds lip sync to it. Record your own voice or use a voiceover you have the rights to.
An AI avatar video starts at 297 credits, and the exact cost is shown before you start. Failed generations are refunded automatically. New accounts receive 50 free credits, the Free plan adds 50 credits monthly, and Pro includes 1,000 credits per month from $9.90.
Yes. Content you generate can be used in ads, stores and social media, subject to the terms of service and the model providers' policies. Only use faces and voices you own or have permission to use.
A sharp, front-facing portrait with the mouth and teeth area unobstructed, neutral expression and soft light. Avoid hands, masks or microphones covering the face.
Video generations usually take about 1–5 minutes, depending on queue. You can run up to 8 generations at the same time, so several scripts can render in parallel.
Yes. Generations are private unless you share them to the community. Adult or explicit content is not allowed and is blocked.
Kling AI Avatar reads your uploaded voice track and animates the mouth and face to match it, so every word lines up with the speech you recorded.
Start from a single photo in JPG, PNG or WebP, up to 10 MB. A real headshot, a brand character or an image you created with Nano Banana 2 or GPT Image 2 all work as the starting frame.
Upload MP3 or WAV audio, up to about 30 seconds and 20 MB. That covers a product hook, a short explainer or a full UGC-style testimonial script.
Each AI avatar video starts at 297 credits and shows the exact cost before you begin. If a generation fails, the credits return to your balance automatically.
Create the portrait, make the talking clip and add B-roll with 13 models on one credit balance. Up to 8 generations can run at the same time per account.
An AI talking avatar is a video in which a still portrait speaks your words. You provide two inputs: a face and a voice. The model studies the audio, predicts the mouth shapes for each sound, and animates the face with matching lip sync. The output is a talking-head clip that looks as if the person in the photo recorded the message on camera.
On NeoMundo, the AI talking avatar tool at /create/avatar runs on Kling AI Avatar from Kuaishou. It needs exactly one image and one audio file. A text prompt is optional and can guide the mood or framing, but the speech itself always comes from the audio you upload, so pronunciation, tone and pacing stay the way you recorded them.
This makes the format useful whenever you need a presenter but not a film shoot. A marketer can turn a 20-second voiceover into an AI avatar video for a product ad. A course creator can add a friendly face to a lesson intro. A small shop can publish several short explainers a week without booking talent, lights or a studio each time.
Because every model on NeoMundo shares one credit balance, you can generate a portrait, animate it, and cut in product shots or B-roll from the video models without switching tools. Videos usually finish in 1–5 minutes, and results stay private until you choose to share them.
Choose a clear, front-facing photo with the mouth visible and even lighting. JPG, PNG and WebP up to 10 MB are accepted. You can also generate a new character in the image generator first.
Upload an MP3 or WAV file of up to about 30 seconds. Record in a quiet room or use a clean voiceover; clear speech gives the most accurate mouth timing.
Describe expression or mood in one short line, such as "friendly, confident, slight smile", or leave the prompt empty and let the audio drive the performance.
Check the credit cost shown before you start, then generate. When the clip is ready, usually within a few minutes, download the MP4 or share it to the community.
Test several scripts with the same presenter. Record 3 different 15-second hooks, create 3 versions of one AI avatar video, and compare them in your ad account.
Pair a spokesperson clip with product photos and videos from the Marketing Studio to explain features, sizing or setup in under 30 seconds.
Give a creator-style portrait a natural, casual script and publish vertical talking-head clips for TikTok, Reels or Shorts.
Add a consistent host to lesson intros, onboarding steps or weekly team updates, keeping the same face across every episode.