- Seedance Blog: AI Video Tutorials & Guides
- Magic Hour Lip Sync Tutorial: From Clean Source to Finished Video
Magic Hour Lip Sync Tutorial: From Clean Source to Finished Video

AI Overview
How do you lip sync a video in Magic Hour?
Upload a video with one clearly visible face, add a clean audio file, choose the available lip-sync mode, and generate. Review the complete clip before downloading because a convincing still frame cannot prove timing.
What source video gives the best lip-sync result?
Use a stable, well-lit shot where the mouth stays unobstructed and large enough to inspect. Moderate head motion and one speaker make errors easier to diagnose than fast cuts or crowded scenes.
How should the audio be prepared?
Trim the voice track to the intended line, remove noise and long silences, and keep natural pacing. Match the emotion and approximate duration of the performance already visible in the source video.
Can one video be lip synced into several languages?
Yes, prepare one approved audio file per language and generate separate versions. Check pronunciation, mouth timing, identity, captions, and final duration independently for every market.
Prepare Source Video and Voice Before Lip Sync
A reliable Magic Hour lip sync tutorial begins before the upload. The source clip should already contain the framing, acting, camera movement, wardrobe, and background you want to keep. Lip sync is a mouth-and-performance transformation, not a rescue operation for a broken face, a hidden jaw, or a camera that loses the speaker.
Use a short representative test first. Choose one speaker, a front-facing or gentle three-quarter angle, even light, and a face large enough to inspect on a phone. Avoid hair crossing the lips, hands covering the mouth, extreme profile turns, rapid cuts, strong motion blur, and repeated entries or exits. If the starting point is a still portrait, animate it first in the image-to-video workspace; Magic Hour's own help guidance separates talking-photo work from lip sync on an existing video.

Editorial output frame for judging face visibility, mouth detail, and a clean single-speaker composition; it is not a Magic Hour benchmark.
Prepare audio as a delivery asset, not an afterthought. Export one clean file with the intended voice, pronunciation, emotion, and pauses. Remove the slate, countdown, unused silence, clicks, and heavy room noise. Do not time-stretch a line so aggressively that consonants smear or the speaker sounds unnatural. The picture may be visually persuasive while the voice reveals the failure immediately.
Use this source-readiness gate before spending a generation:
| Check | Ready signal | Fix before upload |
|---|---|---|
| Face | one identity, both eyes and full mouth readable | crop closer or choose another take |
| Motion | moderate head and camera movement | stabilize or shorten the source |
| Audio | clean single voice with final wording | edit noise, pauses, and mispronunciations |
| Timing | line fits the visible performance window | shorten copy or choose a longer shot |
| Rights | consent and usage rights are recorded | obtain approval before processing |
Run the Magic Hour Lip Sync Workflow
Open Magic Hour's Lip Sync tool and select the video workflow. Upload the approved face video, then upload the prepared audio. The current product surface presents a standard lip-sync mode; available choices, limits, and credit estimates can change, so treat the interface shown in your account as the operational source of truth.
Before generating, confirm that you selected the final files rather than a rough recording with a similar name. A simple version pair prevents mistakes: presenter-en-picture-v03.mp4 and presenter-en-voice-v05.wav. Write down the source duration and audio duration. If the difference is large, do not assume the model will invent a natural pause or accelerate speech gracefully.
Generate one short test and watch it at normal speed. Then replay the opening word, the fastest phrase, any closed-lip sounds such as M or B, and the final breath. Inspect the transition into speech and the resting mouth after speech ends. A mouth that begins moving early or continues after the line is a timing failure even if the middle looks strong.

A finished editorial frame that illustrates the value of a readable mouth, stable angle, and clean single-speaker audio.
The free path is useful for proving the input pair, but do not plan a client delivery around an assumed allowance. Check the live length, format, resolution, mode, and credit information before a production batch. Save the accepted output immediately with its source-video and source-audio version numbers.
Write Audio and Performance Directions That Sync
Lip-sync quality depends partly on whether the new line belongs to the visible performance. A calm speaker with a still torso is a poor match for shouted, breathless audio. A smiling source can make serious narration feel false even when phonemes align. Match energy, cadence, facial tension, and pause structure before asking the mouth to carry the entire illusion.
Use a timed performance brief next to the files:
0.0–0.6 s: relaxed eye contact, mouth at rest. 0.6–4.8 s: deliver one 12–16-word sentence at conversational pace; emphasize the product benefit, with one natural blink. 4.8–5.6 s: close the final consonant, hold eye contact, return to a relaxed mouth. Audio: clean voice only, no music or room echo. Locks: preserve face, hair, wardrobe, camera, background, and product geometry.
The brief does not replace the uploaded audio; it makes review objective. If the output begins speech at 0.2 seconds, the problem is visible. If a translated line runs seven seconds, the producer knows to rewrite or select another shot instead of forcing the words into the original window.
For longer deliverables, split copy at natural edit points. One short approved sentence is easier to correct than a paragraph with multiple emotional turns. Keep a clean master without music so dialogue can be replaced, captions can be timed, and the edit can breathe.
Fix Mouth, Timing, and Identity Failures
Do not solve every bad result by generating again with the same inputs. Classify the failure first, change the smallest likely cause, and compare against the previous version. That preserves evidence and keeps credits focused on a decision.
| Visible failure | Likely cause | Next action |
|---|---|---|
| lips lag the voice throughout | audio lead-in or duration mismatch | trim the head, align start time, regenerate a short test |
| mouth looks soft or unstable | face too small, blurred, or obstructed | use a closer, sharper source with stable lighting |
| identity changes around hard syllables | source angle or expression is too extreme | choose a calmer take and reduce profile movement |
| final mouth keeps moving | trailing silence, breath, or undefined ending | trim audio tail and require a resting end pose |
| only one word looks wrong | pronunciation or compressed consonant cluster | rerecord that line with cleaner diction and a natural pause |
| a second face moves | crowded frame or ambiguous speaker | crop to one speaker or split the shot |
Review at full speed before using slow motion. Frame-by-frame scrutiny can make normal articulation look strange, while viewers experience the line in time. After the real-time pass, inspect only the failed phrase. If the face degrades because it is too distant, the distance-face repair guide explains why source scale and composition should be corrected before polishing.
This existing Seedance media-library clip is an inspection example, not a Magic Hour output. Watch the full motion and listen for mouth timing, voice stability, and room ambience together.
Keep the original and revised outputs side by side. Record the changed input and the reason. If you replace the video, audio, crop, and processing mode at once, the next result cannot tell you which correction mattered.
Localize One Performance Across Languages
Create a separate approved script and audio master for each market. Translation length rarely matches the source exactly: German or Spanish may expand, while another language may fit more compactly. Preserve meaning first, then rewrite for spoken duration. Do not speed the voice until it sounds like an advertisement disclaimer.

Use one clearly framed performance as the visual master, then evaluate every language as its own audio-and-mouth deliverable.
Use this localization matrix for each version:
| Market | Script duration | Pronunciation approved | Mouth timing | Voice identity | Captions | Status |
|---|---|---|---|---|---|---|
| English | 5.2 s | yes | pass | baseline | checked | approved |
| Russian | record | reviewer | review | compare | localize | pending |
| Japanese | record | reviewer | review | compare | localize | pending |
| German | record | reviewer | review | compare | localize | pending |
Ask a fluent reviewer to approve names, numbers, emphasis, and tone. Mouth motion can look synchronized while the language sounds awkward. Keep the visual master, clean dialogue, caption file, and mixed export separate. The voice-consistency checklist helps when the same speaker appears across several clips rather than one isolated line.
Turn the Lip-Synced Clip into a Finished Video
A generated lip-sync clip is one production unit. Finish it by trimming dead frames, checking the delivery aspect ratio, adding verified captions, balancing dialogue against music, and reviewing the exported master on phone speakers and headphones. Do not hide a timing error under loud music; fix the dialogue version first.

The approval frame should preserve the person, product, environment, and resting mouth after the final word.
Use three approval passes. First, watch without sound for face identity, mouth anatomy, eye motion, hands, and background stability. Second, listen without watching for pronunciation, pacing, noise, clipping, and unwanted ambience. Third, watch normally and decide whether picture and voice feel like one performance. For frame-accurate trimming and audio replacement, follow the AI editor versus timeline editor guide.
When one campaign contains several source takes, languages, reviewers, and reruns, the coordination cost becomes larger than the single generation step. In Seedance Agent, keep the approved picture, voice files, language brief, acceptance notes, and rejected versions together; route only the failed shot for revision instead of rebuilding the entire deliverable.
Name the released master with the market and approval version, archive its inputs, and keep a short manifest containing speaker consent, script, source hashes, generation date, and reviewer. That record makes future copy changes safer and prevents an old voice from being paired with a newer visual.
Conclusion
The practical Magic Hour lip sync workflow is simple to run but strict to approve: begin with a clear single-speaker video, prepare a clean voice track that fits the visible performance, generate one short test, inspect the beginning and end of speech, and diagnose the smallest failure before rerunning. For localization, treat every language as a separate performance with its own pronunciation, timing, voice, caption, and export checks. When those versions multiply, organize the complete lip-sync production in Seedance so references, approvals, and selective reruns remain attached to the finished video.
Ready to try it yourself?
Put the steps from this guide into practice with Seedance and turn prompts or images into polished videos in minutes.
Free credits on signup. Plans from $20/month.
Related Articles
More posts in the same locale you may want to read next.

Seedance App Preview Video Generator 2026: Create App Store and Product Launch Clips
Use Seedance to turn app screenshots, feature copy, and launch goals into App Store previews, Google Play promo videos, and product launch clips.
Read article
Grok Imagine Video Extension Prompts: Build Longer Clips Without Losing the Story
Copy practical Grok Imagine video extension prompts, preserve characters and camera logic, fix failed continuations, and assemble longer AI video scenes.
Read article
LibTV Seedance Task ID Verification: Know When a Render Is Really Done
Verify LibTV Seedance task IDs correctly: wait for terminal JSON, confirm canvas write-back, diagnose failures, and approve the finished video.
Read article