Lip sync driven by your audio
Mouth shapes and timing follow the speech in your file, so the portrait says exactly what you recorded, in the language you recorded it in.
Kling
Turn a portrait and a voice into a talking video
Kling AI Avatar lip-syncs a portrait to your audio, producing natural talking-head videos for ads, product explainers and UGC.
Try AI AvatarFAQ
One portrait image and one audio file. A text prompt is optional and only guides expression and mood.
The mouth movement follows the speech in your audio file. Upload a clear recording in the language you need and check the result before publishing.
It starts at 297 credits and depends on duration. You see the exact cost before starting. Failed generations are refunded automatically. New accounts get 50 free credits, and the Free plan adds 50 credits each month.
Content you generate can be used in ads, stores and social media, subject to NeoMundo's terms of service and the model provider's policies. Only use portraits and voices you own or have clear permission to use.
A front-facing, evenly lit face at a good size in the frame, with the mouth visible and a neutral or slightly open expression. Busy backgrounds and strong side angles make lip sync harder.
You can upload any portrait image that shows a clear face and mouth. Results are most consistent with realistic or semi-realistic faces, so test a short audio clip first.
No. The spoken words always come from your audio. The prompt only shapes expression, for example smiling, serious or energetic.
Upload a portrait, add your voice and see the credit cost before you generate. New accounts start with 50 free credits.
Try AI AvatarMouth shapes and timing follow the speech in your file, so the portrait says exactly what you recorded, in the language you recorded it in.
One image (JPG, PNG or WebP up to 10 MB) and one audio file (MP3, WAV or M4A up to 20 MB). No camera, studio or video footage is needed.
Leave the prompt empty for a neutral delivery, or add a short line such as "warm smile, friendly and relaxed" or "serious, confident presenter" to steer the mood.
There are no resolution, aspect ratio or duration menus for this model. Your portrait and audio define the result, which keeps the workflow to a single upload screen.
Generations start at 297 credits. The studio displays the exact cost before you start, and failed runs are refunded automatically.
Kling AI Avatar is the portrait-to-speech model from Kling, the video model family by Kuaishou. It takes a still image of a face and an audio track, then animates the face so the lips, jaw and expression move in time with the voice. On NeoMundo it lives in the avatar studio at /create/avatar and is listed as AI Avatar.
The result is a lip sync video that looks like someone recorded a message to camera. That makes it practical for content where a person speaking is the whole point: product explainers, short ads, course intros, FAQ answers and UGC-style clips, all produced from a photo and a voice recording instead of a shoot.
Inputs: exactly 1 portrait image and exactly 1 audio file, plus an optional text prompt for expression and mood. Pricing starts at 297 credits and depends on duration; the studio shows the exact cost before you generate. Videos usually take about 1 to 5 minutes, and up to 8 generations can run at the same time on one account.
The quality of a talking head video depends mostly on your two files. A clear, front-facing portrait and clean speech without music or background noise give the model the best material to work from; the tips below explain what to look for.
Use a sharp, well-lit photo where the face looks toward the camera, the mouth is fully visible and not covered by hands, hair or a microphone. Head-and-shoulders or waist-up crops work well; avoid tiny faces in wide shots.
Record or export a voice track in MP3, WAV or M4A under 20 MB. One speaker, no background music, and a normal speaking pace help the lip sync stay accurate. Trim silence at the start and end.
Describe the delivery in a few words, for example "enthusiastic, nods occasionally" or "calm and trustworthy". Keep it about mood and expression; the words spoken always come from the audio.
Confirm the credit cost, generate, and download the clip. For a series, reuse the same portrait with new audio files so the presenter stays consistent across videos.
Pair a presenter portrait with a 20 to 30 second voice-over that walks through features, then place the clip on a product page or in a store listing.
Create talking head video ads from a creator photo and a script read, and test several scripts with the same face without booking a new shoot.
Record the same message in different languages and generate one lip sync video per language from the same portrait.
Give lessons, onboarding steps or frequent customer answers a face, updating a clip simply by recording new audio.