HeyGen Translation Speed vs Precision: Which Mode Should You Use?

E
Emma Chen·9 min read·Sep 13, 2026
Share on X
HeyGen Translation Speed vs Precision: Which Mode Should You Use?

AI Overview

What is the difference between HeyGen Speed and Precision translation?

Both translate speech and include lip sync. Speed is aimed at simpler, mostly front-facing footage; Precision is designed for harder mouth visibility, camera angles and speaker changes. Choose by the source shot, not the mode name alone.

Is Speed mode always faster to finish processing?

Not necessarily. “Speed” names a translation-engine option, not a guaranteed delivery time. Processing also depends on source length, account concurrency and queue conditions, so compare completed jobs instead of assuming a fixed time saving.

When should I use Audio Only instead?

Use Audio Only when the speaker is offscreen, small in frame or covered by B-roll, and visible mouth alignment is unnecessary. It translates and revoices speech without paying for lip sync that the viewer cannot inspect.

How should I test a multilingual video before scaling it?

Translate one representative language first, check terminology and the whole clip with sound, then approve or correct its script before launching more languages. A short pilot can expose mouth, timing and speaker-assignment problems early.

What HeyGen's Three Translation Engines Actually Do

HeyGen's Video Translation workflow currently presents Audio Only, Speed and Precision as different engine choices. Audio Only translates and revoices the soundtrack without lip sync. Speed and Precision both attempt visible mouth synchronization; the distinction is the kind of footage each is intended to handle. HeyGen's help center positions Speed for straightforward front-facing speakers and Precision for profile views, occlusions, camera-angle changes and complex conversation. That is a selection guide, not a guarantee that every difficult clip will pass automatically.

Its published credit table lists Audio Only at 4 credits per source minute, Speed at 6 and Precision at 10. Check account pricing before ordering: plans and API billing can differ. The mode label does not guarantee delivery time. HeyGen's separate guidance estimates around five minutes of processing per source minute, with concurrency and demand affecting the wait.

Speed is a source-shot decision

The cleanest Speed candidate is one visible speaker looking toward camera, with an unobstructed face and few abrupt cuts. Pilot Speed on such footage. Its lower credit use than Precision does not remove the need to review difficult words, head turns or subtitles.

Illustrative front-facing speaker in a quiet pottery studio with mouth fully visible

Illustrative generated frame, not a HeyGen output: this simple, unobstructed face is the kind of source shot to pilot with Speed.

Precision is for visible complexity

Choose Precision for side profiles, mouth occlusion, rapid angle changes or two people trading lines. The question is whether those visible risks justify extra credits. If the translation wording is wrong, changing engines will not fix the script.

Illustrative three-quarter-profile kitchen speaker with foreground occlusion and visible mouth

An illustrative profile/occlusion case: inspect lip edges across the entire line, not one flattering still.

Choose a Mode From Your Source Footage

Use the footage itself as your decision tree. Mark each speaking section as mouth not visible, simple visible face, or difficult visible face. If most of the piece is a product montage with narration, Audio Only may be sufficient. If one presenter addresses camera throughout, begin with Speed. If face angle, overlap or speaker switching is central to the deliverable, test Precision on that exact segment rather than a convenient easy clip.

Source pattern First mode to test What to inspect
Narration under B-roll; mouth absent Audio Only Translation, voice, pacing and mix
One centered presenter, unobstructed face Speed Consonant timing, teeth, final syllables
Profile, hands or props crossing the mouth Precision Lip edge stability through obstruction
Two speakers and frequent angle changes Precision Correct voice-to-face assignment at every cut

This table follows HeyGen's published mode descriptions, not a controlled benchmark. If the first candidate fails, fix the source or proofread the translation before another run. A clearer crop or cleaner source mix may help more than a higher mode. Our HeyGen alternatives guide covers platform choice.

Audio Only can be the professional choice

Lip sync matters only when a speaking mouth is visible. In a product tutorial, translation quality and captions may matter more. Dubbing without face alteration preserves approved imagery. Do not spend Precision credits on B-roll; isolate the visible talking-head sections for separate review.

Multiple speakers raise two distinct risks

Speaker switching creates language and visual risks: preserve each person's meaning and tone, then attach the right voice to the right face. Review every handoff. If speech overlaps, consider a cleaner edit or separately approved dialogue tracks.

Illustrative two-person podcast conversation with clear speaker positions and separate microphones

Illustrative generated frame, not a mode test: two visible speakers make assignment at turn-taking the key review point.

Run a One-Language Test Before a Batch

Choose a 20–40 second excerpt containing the hardest condition: a side angle, speaker handoff or long sentence. An easy pilot understates risk. Note source frame rate, aspect ratio and languages. Ensure music does not drown the voice.

Prepare names, terms and permissions

List product names, proper nouns, acronyms, legal phrases and words that must remain untranslated. HeyGen's Brand Glossary can help keep terminology consistent. Decide whether the speaker has consented to translation or voice reproduction and whether the resulting localized message still matches the approved original. Review the entire spoken script for meaning before using lip sync as the success measure. If you already have subtitles, know their role: an input SRT can guide translation, but HeyGen says it may still retranslate the wording rather than reproduce that file verbatim.

Select the mode and inspect the first output

In Video Translation, upload the source, select one target language and choose Audio Only, Speed or Precision according to the shot table. Generate the pilot, then watch with sound at normal speed. Check three moments: the first spoken phrase, a middle turn or camera change, and the final phrase. Pause around hard consonants and speaker switches. Note errors by timestamp and category—translation wording, voice assignment, timing, mouth deformation or background-audio loss. Do not call a clip approved because its thumbnail looks convincing.

The real moving Seedance dialogue sample below is included to demonstrate how to inspect a finished talking-head clip, not to claim a HeyGen result or compare engines. Listen for voice continuity, then watch mouth timing and the cut into the next beat. The same inspection method applies to your actual localized render.

Seedance dialogue output for reviewing voice and visible mouth timing

This is a Seedance output used as a review example, not an official HeyGen translation test.

Proofread before scaling

If the sentence is mistranslated, use HeyGen's Proofread or Review & Edit workflow where your plan supports it. That is a script-control step, not a substitute for visual review. HeyGen says an output SRT used in Proofread mode can be spoken as written, whereas a source SRT submitted before translation is only a guide. Once the pilot passes, record the approved wording and only then expand to additional languages. For a broader checklist of voice continuity across generated clips, see our Seedance voice-consistency guide.

Calculate Credits and Rework, Not Just the First Render

At the help center's listed credit rates, a two-minute source translated into one language would be 8 credits with Audio Only, 12 with Speed or 20 with Precision. Five target languages multiply those initial totals by five. The useful budget is not just the first pass: allow for proofread changes, rejected lip-sync segments and any language that needs a new cut. Confirm current account pricing and minimum billing behavior before purchasing; the example is simple arithmetic, not a quote for every plan.

A practical worksheet has four columns: source minutes, number of target languages, chosen mode and expected reruns. Do one representative language first. If Speed passes the difficult section, a lower-cost batch may be sensible. If it fails only on a side-profile insert, edit that insert or reserve Precision for a separately handled version. If the source itself is unclear, rerendering at a higher tier can multiply cost without fixing the cause.

Fix Problems Before Switching Modes

Mouth timing is off but wording is right. Check face visibility, frame cuts and speech pace. Test Precision on a difficult visible-face excerpt if you began with Speed; use Audio Only if the mouth is not actually part of the deliverable. Meaning or pronunciation is wrong. Correct the script, glossary or source audio first. A lip-sync engine cannot rescue incorrect text. The wrong person seems to speak. Review speaker handoffs and consider a cleaner edit with separated turns.

On-screen text deserves its own pass. HeyGen's Video Translation help says the tool translates spoken audio and optionally lip movement, not embedded titles, graphics or hard-coded captions. Plan localized title cards and captions as separate assets. For a campaign that needs new localized shots rather than a dub of fixed footage, text-to-video creation may be a better starting route than asking a translation engine to rewrite the original picture.

Where Seedance Agent Fits a Multilingual Campaign

HeyGen Speed and Precision address a specific job: translating an existing video while synchronizing speech and, when selected, a visible mouth. Seedance Agent addresses a different production problem. If each market needs a different opening, product shot, background or call to action, the work becomes a sequence of briefs, references, approvals and reruns. Use the approved transcript and shot list to plan those localized variants; ask the agent to keep product geometry, character identity and pacing consistent while you review each market's deliverable. It does not turn HeyGen's Precision switch into a Seedance control.

The Seedance supported-languages guide explains native multilingual generation, which is distinct from dubbing a finished source. If you already have a strong still frame that needs motion, the image-to-video workflow is another useful production path. Choose the route that matches the actual job: translate and preserve an approved performance, or create new localized scenes with coordinated review.

Conclusion

For HeyGen translation, start with Audio Only when a speaking mouth is irrelevant, Speed for simple visible presenters, and Precision for profile views, occlusion or speaker switching. Pilot the hardest 20–40 seconds in one language, proofread meaning, inspect lip and voice continuity, then scale only the approved version. Mode names do not guarantee processing times, and additional credits cannot repair a poor source or incorrect script. When the campaign requires new market-specific scenes rather than only a dub, plan the production with Seedance Agent and carry the approved transcript into the creative brief.

Ready to try it yourself?

Put the steps from this guide into practice with Seedance and turn prompts or images into polished videos in minutes.

Free credits on signup. Plans from $20/month.