- Seedance Blog: AI Video Tutorials & Guides
- MiniMax H3 ComfyUI: Complete Setup Guide for T2V, I2V, and R2V Workflows
MiniMax H3 ComfyUI: Complete Setup Guide for T2V, I2V, and R2V Workflows

AI Overview
Can you run MiniMax H3 locally in ComfyUI?
Yes. MiniMax H3 has open weights and native support in ComfyUI 0.30.0 or later. Load an H3 template, download the required diffusion model, text encoder, and two VAEs, then run it on your own hardware.
What workflows does MiniMax H3 support in ComfyUI?
ComfyUI ships H3 templates for Text-to-Video, Image-to-Video with optional first/last frames, and Reference-to-Video. All three can generate native stereo dialogue, effects, and music with the video.
How much VRAM do you need to run MiniMax H3 in ComfyUI?
There is no single hard minimum. A 24GB GPU is a practical target for the pruned int8 workflow; reduced 480p community tests have run on 12GB with 32GB system RAM, offloading, and fast storage.
Ready to try it yourself?
Free credits on signup. Plans from $20/month.
Is there an easier alternative to running MiniMax H3 locally in ComfyUI?
Yes. Seedance provides browser-based audio-video generation without local model downloads, node setup, or GPU memory planning, making it simpler when you want a finished clip rather than a custom local graph.
What Is MiniMax H3 and Why Run It in ComfyUI?
MiniMax H3 is an open-weights, omni-modal generation model that understands text, images, video, and audio in one context. It can produce video at up to 2K, 24fps, and about 15 seconds, with voice, sound effects, and music generated together as native stereo audio.
ComfyUI turns those capabilities into a visible node graph. You can choose the model precision, seed, resolution, duration, scheduler, references, first or last frame, and output path—then save the graph as a repeatable production workflow. Native templates cover T2V, I2V, and R2V, while the core H3 nodes can be recombined for more specialized graphs.
Run H3 locally when parameter control, reusable nodes, private inputs, or batch automation matters more than setup time. It fits creators who already understand ComfyUI and teams that want to avoid per-call API quotas. The tradeoff is substantial storage, GPU memory planning, long downloads, and more troubleshooting than a browser workflow.
System Requirements and Model Downloads
Update ComfyUI to 0.30.0 or later, then open Workflow → Browse Workflow Templates → Video and select a MiniMax H3 template. ComfyUI can prompt for missing files, but manual placement makes problems easier to diagnose.
| Required file | Put it in | Purpose |
|---|---|---|
minimax_h3_fl2va_pruned_int8_convrot.safetensors |
ComfyUI/models/diffusion_models/ |
T2V, I2V, first/last-frame generation |
qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors |
ComfyUI/models/text_encoders/ |
Multimodal prompt and reference understanding |
minimax_h3_video_vae_fp16.safetensors |
ComfyUI/models/vae/ |
Video latent encoding and decoding |
minimax_h3_audio_vae_fp32.safetensors |
ComfyUI/models/vae/ |
Native audio encoding and decoding |
R2V uses minimax_h3_ref2va_pruned_int8_convrot.safetensors in diffusion_models/ instead of the FL2VA diffusion file. Keep the full filenames; after adding them, refresh models or restart ComfyUI.
Treat 24GB VRAM as a comfortable planning target for the pruned int8 path, not a published minimum. Lower-memory cards need smaller frames, shorter clips, aggressive offloading, and enough system RAM plus NVMe space for model swapping. If the graph cannot load or spends most of its time swapping, reduce resolution first or move the workflow to cloud hardware.
MiniMax H3 Text-to-Video (T2V) Workflow Setup
- Open Template Library and load MiniMax H3 T2V.
- Confirm the diffusion model, Qwen3-VL text encoder, video VAE, and audio VAE are selected.
- In Resolution Selector, choose aspect ratio and megapixels; keep
multipleat32. - Enter one structured prompt, set duration and seed, then queue the graph.
- Review picture, dialogue, effects, music, and the final frame before raising resolution.
H3's native working canvas uses a 768px short edge and caps a side at roughly 1344px. Around 1.0 megapixel at 16:9 resolves to about 1344×768. Start smaller for prompt tests, then scale only the approved setup. Duration snaps to H3's frame-block grid at 24fps, so ComfyUI may adjust the requested value slightly.
Write the scene first, then timed action, camera movement, and audio in the same prompt block. A clean T2V brief is more reliable than a list of unrelated cinematic adjectives. For a faster hosted starting point, the Seedance Text to Video workflow removes the local graph and download steps.
Official ComfyUI template output for MiniMax H3 T2V. It is a real workflow sample, rehosted on Seedance R2 for reliable playback.
MiniMax H3 Image-to-Video (I2V) Workflow Setup
The I2V template uses MiniMaxH3ImageToVideo. Connect an image to first_frame, optionally connect another to last_frame, and describe the motion between them. Both inputs are optional at node level, so the same FL2VA weights can support text-led, first-frame, last-frame, or first-and-last-frame workflows.
I2V works best for product showcases, character continuity, controlled transitions, and style-led animation. Match the output to H3's 768px short-edge canvas, keep dimensions divisible by 32, and avoid feeding a tiny or heavily compressed source. Describe what must remain unchanged before describing camera motion.

The real input image distributed with ComfyUI's MiniMax H3 I2V template.
The official I2V result from that source image. Compare product shape, materials, wheel placement, camera path, and sound—not just first-frame impact.
If you want the same input-to-motion idea without maintaining local weights, try Seedance Image to Video.
MiniMax H3 Reference-to-Video (R2V) Workflow Setup
R2V uses MiniMaxH3ReferenceToVideo and the separate ref2va diffusion weights. It accepts mixed images, videos, and audio so a target shot can borrow character identity, visual style, subject motion, camera movement, or voice from different sources.
Connect references in a deliberate order and name them exactly in the prompt: <Picture 1>, <Video 1>, and <Audio 1>. Assign one job to each reference. For example: preserve the character from Picture 1, use the dolly move from Video 1, and match the voice from Audio 1. ComfyUI's current workflow supports up to nine images, three videos, and three standalone audio clips.
Use ref_image_size=match for faster tests or max to preserve a reference short edge up to 2048px when identity matters more than speed. Too many overlapping references can weaken instruction following, so prove a two-reference graph before adding more.
Official R2V template output using the distributed character and style references.
For a hosted reference workflow, open Seedance Reference to Video. The MiniMax H3 reference-video guide provides a deeper prompt and reference-assignment walkthrough.
Prompt Writing Tips for Better MiniMax H3 Output
Use three explicit fields for H3 prompts: integrated_multimodal_description, overall_soundscape, and non_diegetic_music. Inside the scene description, write timed shots, visible action, camera direction, and exact dialogue. Give recurring speakers stable IDs such as (S1) and place only spoken words inside <d>[English] ...</d>.
T2VA — weak: A cinematic city scene with cool sound.
T2VA — better: 0–3s wide rainy crosswalk, slow dolly forward; 3–7s courier turns toward camera. Footsteps and traffic in stereo; low percussion begins at 4s.
I2VA/FL2VA — weak: Animate this image dramatically.
I2VA/FL2VA — better: Preserve the product geometry, wheel, buttons, transparent shell, and blue-orange lighting. One slow 30-degree orbit; no new text or parts. End aligned to the supplied last frame.
R2V — weak: Use all references and make an action video.
R2V — better: Keep identity and costume from <Picture 1>; copy only camera motion from <Video 1>; match the voice tone from <Audio 1>. Do not transfer the video subject or background.
T2VA needs a complete world description. I2VA should protect the source before adding motion. FL2VA and L2VA must describe the transition into the supplied boundary frame. R2V must explain the relationship between every reference and the target. The MiniMax H3 prompting guide expands these mode-specific patterns.
MiniMax H3 ComfyUI vs. Browser-Based AI Video Tools
| Factor | Local H3 in ComfyUI | Seedance | Comfy Cloud |
|---|---|---|---|
| Setup | Install/update ComfyUI and place large weights | Open in a browser | Load a hosted ComfyUI template |
| Local VRAM | Required; settings depend on precision and output | Not required | Not required locally |
| Workflow control | Full node-level control and automation | Streamlined generation interface | ComfyUI graph without local hardware |
| Native audio | H3 stereo audio in the same generation | Joint audio-video workflow | Same H3 workflow on cloud compute |
| Best for | Deep customization, local inputs, reusable graphs | Fast ideation and production without setup | ComfyUI users who lack local GPU capacity |
Choose local ComfyUI when the graph itself is part of your production system. Choose cloud-hosted ComfyUI when you want its nodes but not the hardware burden. Choose Seedance when you want to move from prompt or reference to a usable clip with fewer infrastructure decisions. Try Seedance in your browser →
Conclusion
MiniMax H3 is one of ComfyUI's most capable open-weights options for generating video and native audio together. Start in three steps: update to ComfyUI 0.30.0+, place the diffusion model, text encoder, and two VAEs in their correct folders, then load the matching T2V, I2V, or R2V template. Keep early tests small, assign every reference a job, and raise resolution only after the prompt works; if local downloads, VRAM, and node maintenance outweigh the benefit of deep control, a browser-based Seedance workflow is the faster route to production.
Ready to try it yourself?
Put the steps from this guide into practice with Seedance and turn prompts or images into polished videos in minutes.
Free credits on signup. Plans from $20/month.
Related Articles
More posts in the same locale you may want to read next.

Seedance App Preview Video Generator 2026: Create App Store and Product Launch Clips
Use Seedance to turn app screenshots, feature copy, and launch goals into App Store previews, Google Play promo videos, and product launch clips.
Read article
MiniMax H3 Gibberish Fix: Why It Happens and How to Solve It
Fix MiniMax H3 gibberish audio with cleaner dialogue tags, shorter prompts, stable speaker IDs, single-language scripts, and practical fallback workflows.
Read article
MiniMax H3 Turbo LoRA: 5× Faster Video Generation Setup Guide
Install MiniMax H3 Turbo LoRA in ComfyUI, choose v4 or v1, wire the sampler, and tune 4–8 step generation for faster video with stereo audio.
Read article