AI Video

AI Lip Sync: How to Make a Talking Video Look Natural

AI lip sync from phonemes to export: the portrait and audio that work, a step-by-step LazyKiwi walkthrough, tool prices checked in 2026 and artifact fixes.

Daniel OkaforDaniel Okafor 9 min read
Share
AI Lip Sync: How to Make a Talking Video Look Natural

Lip sync AI takes speech (a recording or text read by a synthetic voice) and reshapes the mouth in a photo or a video so it moves with each sound. Done well, a single headshot becomes a presenter; done badly, the jaw floats and the teeth flicker. This guide covers how the models map sounds to mouth shapes, the exact steps in the LazyKiwi Avatar Talking flow, which source images and audio behave, the lip sync AI tools worth a test with prices read on 5 September 2026, the free options, the fixes for the usual artifacts, and the disclosure rules on YouTube, TikTok and Meta.

What is AI lip sync and how does it work?

Speech is a stream of phonemes (the sounds), and every phoneme maps to a viseme (the mouth shape that makes it). The model listens to the audio, predicts the viseme sequence, and paints the lower face to match, frame by frame, while keeping the rest of the head as it was.

Concept: matching mouth movement to audio
Concept: matching mouth movement to audio

Why did older models drift out of time?

The research turning point was Wav2Lip, published as 'A Lip Sync Expert Is All You Need for Speech to Lip Generation In The Wild' by Prajwal, Mukhopadhyay, Namboodiri and Jawahar; the authors trained the generator against a dedicated lip-sync discriminator and reported accuracy 'almost as good as real synced videos', per the arXiv abstract (2020). Earlier approaches optimised for a pretty frame rather than for timing, which is why the mouth in old face-swap clips lags the audio by a few frames.

Which failures does the model still make?

Three remain common. Plosives (p, b, m) need a fully closed mouth for a frame or two, and weak models skip the closure. Sustained vowels stretch the jaw open longer than a human would. And silence is hard: the model has to decide whether a pause is a closed mouth or a relaxed open one, and it often keeps chewing. Commercial vendors now ship separate speed and precision modes for this reason; HeyGen's API, for example, lists 'Lipsync - Speed' and 'Lipsync - Precision' as distinct endpoints in its developer docs (2026).

Photo-based systems add a second stage: they first animate the still (blink, head turn) and then sync the mouth. That is what the LazyKiwi Avatar Talking page does with a headshot and a typed script, and what the AI video generator does not do on its own, since its models produce silent motion.

How to make an AI lip sync video step by step

The Avatar Talking flow in LazyKiwi is text-driven: you type the lines and pick a voice; there is no audio upload on the page. The following steps are the ones on screen on 5 September 2026.

  1. 1

    Drop in a headshot

    Use Upload avatar. A straight-on face with the lips visible and no hand, microphone or mug in front of the chin gives the model a clean mouth region to repaint.

  2. 2

    Write for the mouth, not the eye

    Keep lines short and end them with full stops. Spell numbers as words if a digit must be crisp ('twenty-nine' rather than '29'). The field accepts 1,000 characters; a 25-second clip is about 60 words.

  3. 3

    Leave the action on Speech for a first pass

    Speech is the formal, mostly still preset. Head-heavy presets (Conversation, Emotional, Singing) move the face more, which hides sync errors but also creates new ones at the mouth corners.

  4. 4

    Pick a voice and start with Neutral emotion

    The emotion setting changes how wide the mouth opens on stressed words. Neutral gives the flattest, easiest-to-check motion; switch to Happy or Serious only after the timing is right.

  5. 5

    Generate and review at quarter speed

    Scrub the result frame by frame on 'p', 'b' and 'm' sounds and on every full stop. If the lips do not close on the plosives, cut the sentence in half and run it again before touching the headshot.

  6. 6

    Add motion and cutaways elsewhere

    For a moving-camera version of the same face, run the headshot through Hailuo 2.3 Fast image-to-video for a silent 6-second turn, then cut it between talking segments.

Workflow: portrait rules, audio prep, sync
Workflow: portrait rules, audio prep, sync

Credits

Talking clips and image-to-video renders draw from the same balance. The free plan carries 40 credits a month, per the LazyKiwi pricing page, so keep test scripts under ten seconds until the timing is right.

Which portrait and audio give the best lip sync results?

The input decides more than the model. The table below is the order we would try sources in, from safest to riskiest.

SourceSync qualityWhat goes wrongUse it for
Front-facing headshot, mouth closedBestAlmost nothing; teeth are invented, so check themPresenters, explainers
Three-quarter viewGoodFar corner of the mouth smears on wide vowelsCharacter clips with a turn
Smiling with teeth showingFairTooth flicker as the model repaints the smileShort, upbeat lines only
Profile or tilted headPoorOne lip hidden; the model guesses the shapeAvoid
Existing video of the speakerBest with a video toolHead motion fights the new mouth if the tool is photo-onlyDubbing, re-voicing

What makes audio easy to sync?

Clean, dry speech with no music underneath. Synthetic voices are the easiest of all, because they never mumble; ElevenLabs, the voice most creators route through, lists 74 languages for text to speech and a $0 tier with 10k credits a month, per its pricing page (2026). If you record yourself, keep the microphone off-axis so plosives do not pop, and leave half a second of silence at both ends.

For head motion, the Hailuo route is image-only: the model has no text-to-video mode, keeps the frame shape of your upload, runs 6 to 60 seconds at 768p or 1080p, and takes prompts up to 2,000 characters, per the Hailuo model page. Prompt for 'slight nod, lips closed' so the silent clip does not fight the talking one when you intercut them.

Which lip sync AI tools are worth testing in 2026?

Seven tools cover the field, and they split by input: photo plus script, video plus audio, or performance capture. Prices are as shown on each site on 5 September 2026; where a page publishes none, the cell says so rather than guessing.

ToolInputEntry price (5 September 2026)Free tierBest for
LazyKiwi Avatar TalkingPhoto + typed scriptCredit plans40 credits a monthTalking headshots without an audio file
Sync Labs (sync-3)Video + audioHobbyist $5 a month plus $0.05 per secondPay-per-use with watermarkRe-voicing real footage, 4K
HeyGenPhoto or stock avatar + script or audioCreator $29 a month3 videos a month, 1 minutePresenter videos with translation
D-IDPhoto + scriptLite $4.7 a month billed annually14-day trial, watermarkedCheap talking photos
HedraAvatar models in a studioBasic $15 a month, 1,500 creditsLimited watermarked generationsCharacter clips with other models nearby
Runway Act-OnePerformance video + character imagePaid Runway plansNone for Act-OneActing transfer, not dubbing
Kling AIInside the Kling appNot stated on the public home pageNot statedUsers already generating in Kling

Sync Labs

The specialist. The site leads with sync-3 as 'the most intelligent lipsyncing model', takes a video and an audio track, and handles objects blocking the face, multiple faces and 4K ProRes, per the Sync Labs site (2026). Pricing is per second generated: Hobbyist $5 a month at $0.05 per second for clips up to 1 minute, Creator $19 a month for 5-minute clips with no watermark, Growth $49, Scale $249, and a free pay-per-use tier that watermarks, per the Sync Labs pricing page (2026).

Sync Labs' home page on 5 September 2026, leading with the sync-3 lip-sync model.
Sync Labs' home page on 5 September 2026, leading with the sync-3 lip-sync model.

HeyGen and D-ID

Both are avatar platforms with lip sync as a component. HeyGen's Creator plan is $29 a month and the free tier stops at three one-minute videos, per the HeyGen pricing page (2026). D-ID describes its AI Avatars as digital humans built 'from images or video for both offline videos and real-time experiences', per the D-ID site (2026), and its Studio pricing with the Annual toggle showed Lite at $4.7 a month ($56 a year) with 40 credits and a watermarked 14-day trial, per the D-ID pricing page (2026).

D-ID Studio pricing on 5 September 2026 with the Annual toggle: Trial, Lite, Pro and Advanced.
D-ID Studio pricing on 5 September 2026 with the Annual toggle: Trial, Lite, Pro and Advanced.

Hedra, Runway Act-One and Kling

Hedra sells credits (Basic $15 a month for 1,500), per its pricing page (2026), and positions avatar models inside a wider studio. Runway's Act-One is a different animal: it takes a driving performance video and a character image and transfers 'eye-lines, micro expressions, pacing and delivery', and is available on all paid plans, per Runway Research (2026). Kling AI's public site describes audio-visual synchronisation research but publishes no lip-sync spec or price on its home page, per Kling AI (2026), so evaluate it inside the app.

For the best lip sync video ai at a given budget, match input first: if you have footage, Sync Labs; if you have a photo and no audio, LazyKiwi or D-ID; if you have an actor, Act-One. A wider comparison of the avatar platforms is in our avatar generator roundup.

Is there a free lip sync AI?

Yes, with strings attached. Every free tier we checked on 5 September 2026 either watermarks (Sync Labs pay-per-use, D-ID Trial), caps length (HeyGen, one minute) or meters credits (LazyKiwi, 40 a month on the free plan). Searches for lip sync ai free and for an ai lip sync video generator free usually end at one of those three.

  • LazyKiwi: 40 credits a month, no watermark step, text-driven only; enough for a handful of ten-second tests.
  • Sync Labs: free pay-per-use with a watermark; the $5 Hobbyist tier removes the cap on jobs, not the watermark (Creator does).
  • HeyGen: three one-minute videos a month at $0, with the custom video avatar included.
  • D-ID: 14-day trial, watermarked, then Lite from $4.7 a month billed annually.
  • Open source: Wav2Lip's code is public; expect to run it yourself on a GPU.

An ai lip sync generator that is free and unwatermarked and unmetered does not exist, and any page promising all three is either metering something else or reselling one of the tools above. Budget one paid month once the test clips prove the source image works.

How to fix common AI lip sync artifacts

Most artifacts trace back to the input, so the fix is usually upstream of the model.

ArtifactWhat it looks likeLikely causeFix
DriftMouth runs early or late by a few framesMusic or reverb in the audio; long run-on sentenceDry audio, shorter sentences, regenerate
Teeth flickerTeeth appear and vanish between framesSmiling source with visible teethClosed-mouth headshot
Jaw over-motionChin drops on every syllableEnergetic voice or Happy emotion settingNeutral emotion, calmer voice
Chewing in silenceMouth keeps moving during pausesNo silence padding; model fills gapsAdd half a second of silence, end lines with full stops
Corner smearMouth corners blur on wide vowelsThree-quarter or tilted sourceStraight-on portrait
Head fights mouthNeck moves against the new lipsPhoto-only tool on a video sourceUse a video-input tool such as Sync Labs

Two workbench tricks help after the fact. If only the presenter is wrong, run the finished clip through AI character swap with a reference photo of the new face rather than regenerating the speech. If the whole style is wrong, video to video AI restyles the clip while keeping the mouth timing you already fixed.

Check plosives at quarter speed before you check anything else; if the lips close on 'p' and 'b', the rest of the clip is usually fine.

— LazyKiwi editorial test note, September 2026

Do lip-synced videos need an AI disclosure?

Often, yes. YouTube requires creators to disclose when they use AI to 'meaningfully alter or generate photorealistic content', and its first listed example is content that 'makes a real person appear to say or do something they didn't do'; cloning your own voice for a voice-over is listed as not requiring disclosure, per the YouTube Help policy (2026).

YouTube's 'Disclosing use of GenAI content' policy page, whose first example of content that must be disclosed is a real person made to say something they did not.
YouTube's 'Disclosing use of GenAI content' policy page, whose first example of content that must be disclosed is a real person made to say something they did not.

TikTok introduced its AI-generated content label in September 2023 and requires labels on AI content with 'realistic images, audio or video', per the TikTok Newsroom (2023). Meta began adding 'AI info' labels to video, audio and images when it detects industry-standard AI indicators or when people disclose, per the Meta Newsroom (2024).

Result: a natural talking-head frame
Result: a natural talking-head frame
  • Your own face, your own words: label it anyway on YouTube if the result is photorealistic; it costs nothing.
  • Someone else's face: written consent first, disclosure always, and no scripts they did not approve.
  • Fictional or clearly stylised character: usually exempt, but a label still avoids takedowns.
  • Translated versions of a real speaker: disclose the dubbing; our AI video dubbing guide covers the workflow.

Key takeaways

  • Lip sync models map phonemes to visemes; Wav2Lip (2020) fixed timing with a lip-sync discriminator, and vendors now ship separate speed and precision modes.
  • In LazyKiwi, Avatar Talking is text-driven (headshot plus a 1,000-character script); the free plan carries 40 credits a month.
  • A straight-on, closed-mouth headshot and dry audio prevent most artifacts; teeth flicker and jaw over-motion come from smiling sources and Happy voices.
  • Sync Labs is the specialist for video-plus-audio at $0.05 per second; D-ID Lite at $4.7 a month billed annually is the cheapest paid talking photo on 5 September 2026.
  • YouTube, TikTok and Meta all label realistic synthetic speech; disclose whenever a real person appears to say something new.
Daniel Okafor

Video Workflows Editor

Daniel Okafor

I run the same clip through every video model we ship and write down what actually changes: motion, timing, cost, and where each one falls apart.

FAQ

Common questions

How do I lip sync an AI video?

With a photo: upload a straight-on headshot to LazyKiwi Avatar Talking, type the lines, pick a voice, generate, then check plosives frame by frame. With existing footage: upload the clip and a clean audio track to a video-input tool such as Sync Labs, which repaints the mouth to the new audio. Shorter sentences and dry audio fix most timing faults.

Does Hailuo AI have lip sync?

Not in LazyKiwi. Hailuo 2.3 Fast is image-to-video only and produces silent motion at 768p or 1080p for 6 to 60 seconds. Use it for head turns and cutaways, then create the talking segment in Avatar Talking, which is the text-driven lip-sync flow, and cut the two together.

Can AI lip sync to any language?

Most tools sync to whatever audio they receive, because visemes are shared across languages; the limit is the voice. ElevenLabs lists 74 text-to-speech languages on its pricing page, and Sync Labs names English, Spanish, Hindi, French, German and Korean on its site. Check tone languages and fast speech by ear before publishing.

Why is my lip sync out of time?

The usual causes are music or reverb under the voice, a run-on sentence with no full stops, and a source image with the mouth open or teeth showing. Strip the audio to dry speech, split the script into short lines, and regenerate from a closed-mouth headshot. If you used a photo tool on a video, switch to a video-input model.

Is AI lip sync free?

Partly. On 5 September 2026 LazyKiwi's free plan gave 40 credits a month with no watermark step, Sync Labs offered watermarked pay-per-use, HeyGen allowed three one-minute videos, and D-ID a watermarked 14-day trial. Unwatermarked, unmetered output needs a paid tier; Sync Labs Creator ($19 a month) is the first without a watermark.

LazyKiwi

Make your first lip-synced clip

Upload a headshot, type the lines, pick a voice. Test on a ten-second script before you spend credits on a full one.

Make a lip-synced video