- Seedance Blog: AI Video Tutorials & Guides
- Wan Animate 2 Tutorial: Move Mode, Mix Mode, and ComfyUI Workflow Guide
Wan Animate 2 Tutorial: Move Mode, Mix Mode, and ComfyUI Workflow Guide

AI Overview
What is Wan Animate 2 and what can it do?
Wan Animate 2 is an open-source character animation system built on Wan 2.2. It transfers full-body motion, facial expression, and lip movement from a driver video to a reference character through Move or Mix mode.
What is the difference between Move mode and Mix mode in Wan Animate 2?
Move animates a still character and lets the model create the surrounding background. Mix replaces the person in a driver video while retaining the original scene, background, and audio.
Can I run Wan Animate 2 locally in ComfyUI?
Yes. The workflow uses Wan 2.2 Animate weights, WanVideoAnimateEmbeds, a sampler, and a VAE. Choose bf16 when memory allows or int8-convrot for a lighter local setup.
Ready to try it yourself?
Free credits on signup. Plans from $20/month.
What inputs does Wan Animate 2 require?
Both modes need a reference image and driver video. Reliable face crops improve identity; Mix mode additionally benefits from clean background frames and a solid SAM2 person mask.
What Wan Animate 2 Actually Does — and Why It's Different
Ordinary image-to-video models invent motion from a still and a prompt. Wan Animate 2 instead treats a real performance as structured guidance: the driver supplies pose, timing, facial expression, and mouth movement, while the reference image supplies the character identity. That makes it useful when the performance matters as much as the look.
This is also different from endpoint-controlled generation. A first-and-last-frame workflow defines where a shot begins and ends, but it does not reproduce every beat between them. See the MiniMax H3 first-and-last-frame guide when the destination frame matters more than performance transfer.
The three-panel clip keeps the still reference, driving performance, and generated result visible at the same time. It is a more useful quality check than a single polished frame because you can compare gesture timing and identity directly.
Move Mode vs Mix Mode — Choosing the Right One
Both modes transfer a performance, but they solve different production problems. Use this decision table before building masks or preparing extra frames:
| Decision | Move mode | Mix mode |
|---|---|---|
| Inputs | Reference image + driver video | Reference image + driver video |
| Background | Regenerated by the model | Original video background is preserved |
| Best use | Character action and image animation | Person replacement inside an existing scene |
| One-line rule | “Make the picture move” | “Replace the person in the video” |
Choose Move when the character and movement are the creative core and the environment may change. Choose Mix when blocking, props, lighting, camera movement, and sound already work in the source clip. For a browser-first version of the same idea, try the motion control workflow before committing to a local graph.
This example shows why silhouette and body proportion matter. A driver and character can look completely different, yet the result remains readable when the reference has a clear full-body pose and the driver stays inside frame.
What You Need Before Starting — Inputs and Model Files
Prepare the assets before opening ComfyUI. A clean input set prevents more failures than adding sampling steps later.
- Driver video: one visible subject, stable framing, clear limbs, and deliberate facial movement. Trim dead time before export.
- Reference image: a sharp full-body or three-quarter character with visible hands, feet, and an unobstructed face.
- Face crops: close crops of the reference and driver faces for stronger identity and expression alignment.
- SAM2 mask for Mix: a solid person mask without holes around hair, hands, or moving clothing.
- Background frames for Mix: clean plates from the original shot; use them whenever the replaced actor reveals previously hidden areas.
- Wan 2.2 Animate weights: bf16 offers the cleanest full-precision path; int8-convrot lowers memory pressure with a possible quality tradeoff.
Hardware, quantization, and decode time still shape the local experience. The tradeoffs in our LTX 2.5 vs MiniMax H3 comparison are a useful checklist when deciding whether to run locally or use a hosted workflow.

The official architecture diagram shows how the reference image, reference video, and target-video tokens enter the animation pipeline.
Step-by-Step ComfyUI Workflow
- Load the Wan 2.2 Animate weights. Select bf16 for quality when VRAM permits, or the int8-convrot build for a lighter graph. Confirm that the VAE and text encoder match the workflow.
- Connect
WanVideoAnimateEmbeds. Feed it the reference identity, driver pose sequence, face crops, and SAM2 mask. Mix mode needs the cleanest mask because the original background must survive the replacement. - Match the driver dimensions. Set width, height, and
num_framesto the prepared driver video. Use dimensions divisible by 16 and test a short segment first. - Configure
WanVideoSampler. Add a direct scene prompt, any compatible LoRA, the sampler, scheduler, seed, and step count. The prompt should describe appearance and environment rather than fighting the driver’s action. - Choose Mix or Move. Route Mix with the original scene and background frames; route Move when the model should rebuild the environment around the animated reference.
- Decode and save. Inspect the first result at full size, then compare face, fingers, mask edges, floor contact, and lip timing against the driver before raising duration or resolution.
If you want to validate the character and framing without installing models, start with the reference-to-video generator, then carry the approved input pair into ComfyUI.
Getting Better Results — Common Fixes and Tips
- Background drift: provide clean background frames in Mix mode and use a solid SAM2 mask. Feathering too aggressively can create a halo around the actor.
- Face drift: use sharp face crops, even lighting, and a reference angle close to the driver. Avoid hands covering the face during the first test.
- Resolution mismatch: make width, height, and
num_framesmatch the driver exactly. Resize once before the graph rather than at several nodes. - Unnatural lips: choose a driver with a visible mouth, clear syllables, and limited motion blur. Extreme profile angles reduce usable lip detail.
- Slow iteration: apply a compatible lightx2v step-distillation LoRA and begin with a short clip. Fewer steps are valuable for tests, but always compare the final run against the undistilled baseline.
Change one variable at a time. A fixed seed, short driver segment, and saved input set turn every retry into a useful comparison instead of a new experiment.
Use Cases — What You Can Build with Wan Animate 2
- AI host or digital human — Move: animate a clean presenter portrait from a recorded performance when the set can be regenerated around the host.
- E-commerce model replacement — Mix: keep product staging and camera motion while replacing the performer with a campaign character.
- Custom IP animation — Move: transfer human gesture and facial timing to a mascot, illustration, or stylized 3D character.
- UGC dance and action — Move: turn one approved image into a movement-led social clip without hand-animating key poses.
- Batch ad character replacement — Mix: preserve the same edit, location, props, and audio while adapting the on-screen talent for multiple variants.
For recurring brand characters, dataset consistency matters more than adding random footage. The preparation and validation practices in our MiniMax H3 LoRA training guide transfer well to custom animation projects.
Conclusion
The simplest rule is Move = animate the image; Mix = replace the person in the video. Start with a clean reference and short driver, lock dimensions and masks, then scale only after identity and motion remain stable. If you want the same reference-led workflow without maintaining local weights and nodes, create your next AI video on Seedance →.
Ready to try it yourself?
Put the steps from this guide into practice with Seedance and turn prompts or images into polished videos in minutes.
Free credits on signup. Plans from $20/month.
Related Articles
More posts in the same locale you may want to read next.

Seedance App Preview Video Generator 2026: Create App Store and Product Launch Clips
Use Seedance to turn app screenshots, feature copy, and launch goals into App Store previews, Google Play promo videos, and product launch clips.
Read article
LTX 2.5 Review: Speed, Multi-Shot, and How It Compares to MiniMax H3
An honest LTX 2.5 review covering speed, native multi-shot video, ComfyUI setup, open weights, image quality, and a direct MiniMax H3 comparison.
Read article
MiniMax H3: 30 vs 50 Steps — What Actually Changes in Quality and Speed
Compare MiniMax H3 at 20, 30, and 50 steps for image detail, generation time, audio quality, Turbo LoRA speed, and the best setting for each workflow.
Read article