MiniMax H3 ComfyUI: Complete Setup Guide for T2V, I2V, and R2V Workflows

E
Emma Chen·7 min read·Aug 17, 2026
Share on X
MiniMax H3 ComfyUI: Complete Setup Guide for T2V, I2V, and R2V Workflows

AI Overview

Can you run MiniMax H3 locally in ComfyUI?

Yes. MiniMax H3 has open weights and native support in ComfyUI 0.30.0 or later. Load an H3 template, download the required diffusion model, text encoder, and two VAEs, then run it on your own hardware.

What workflows does MiniMax H3 support in ComfyUI?

ComfyUI ships H3 templates for Text-to-Video, Image-to-Video with optional first/last frames, and Reference-to-Video. All three can generate native stereo dialogue, effects, and music with the video.

How much VRAM do you need to run MiniMax H3 in ComfyUI?

There is no single hard minimum. A 24GB GPU is a practical target for the pruned int8 workflow; reduced 480p community tests have run on 12GB with 32GB system RAM, offloading, and fast storage.

Ready to try it yourself?

Free credits on signup. Plans from $20/month.

Try Seedance free

Is there an easier alternative to running MiniMax H3 locally in ComfyUI?

Yes. Seedance provides browser-based audio-video generation without local model downloads, node setup, or GPU memory planning, making it simpler when you want a finished clip rather than a custom local graph.

What Is MiniMax H3 and Why Run It in ComfyUI?

MiniMax H3 is an open-weights, omni-modal generation model that understands text, images, video, and audio in one context. It can produce video at up to 2K, 24fps, and about 15 seconds, with voice, sound effects, and music generated together as native stereo audio.

ComfyUI turns those capabilities into a visible node graph. You can choose the model precision, seed, resolution, duration, scheduler, references, first or last frame, and output path—then save the graph as a repeatable production workflow. Native templates cover T2V, I2V, and R2V, while the core H3 nodes can be recombined for more specialized graphs.

Run H3 locally when parameter control, reusable nodes, private inputs, or batch automation matters more than setup time. It fits creators who already understand ComfyUI and teams that want to avoid per-call API quotas. The tradeoff is substantial storage, GPU memory planning, long downloads, and more troubleshooting than a browser workflow.

System Requirements and Model Downloads

Update ComfyUI to 0.30.0 or later, then open Workflow → Browse Workflow Templates → Video and select a MiniMax H3 template. ComfyUI can prompt for missing files, but manual placement makes problems easier to diagnose.

Required file Put it in Purpose
minimax_h3_fl2va_pruned_int8_convrot.safetensors ComfyUI/models/diffusion_models/ T2V, I2V, first/last-frame generation
qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors ComfyUI/models/text_encoders/ Multimodal prompt and reference understanding
minimax_h3_video_vae_fp16.safetensors ComfyUI/models/vae/ Video latent encoding and decoding
minimax_h3_audio_vae_fp32.safetensors ComfyUI/models/vae/ Native audio encoding and decoding

R2V uses minimax_h3_ref2va_pruned_int8_convrot.safetensors in diffusion_models/ instead of the FL2VA diffusion file. Keep the full filenames; after adding them, refresh models or restart ComfyUI.

Treat 24GB VRAM as a comfortable planning target for the pruned int8 path, not a published minimum. Lower-memory cards need smaller frames, shorter clips, aggressive offloading, and enough system RAM plus NVMe space for model swapping. If the graph cannot load or spends most of its time swapping, reduce resolution first or move the workflow to cloud hardware.

MiniMax H3 Text-to-Video (T2V) Workflow Setup

  1. Open Template Library and load MiniMax H3 T2V.
  2. Confirm the diffusion model, Qwen3-VL text encoder, video VAE, and audio VAE are selected.
  3. In Resolution Selector, choose aspect ratio and megapixels; keep multiple at 32.
  4. Enter one structured prompt, set duration and seed, then queue the graph.
  5. Review picture, dialogue, effects, music, and the final frame before raising resolution.

H3's native working canvas uses a 768px short edge and caps a side at roughly 1344px. Around 1.0 megapixel at 16:9 resolves to about 1344×768. Start smaller for prompt tests, then scale only the approved setup. Duration snaps to H3's frame-block grid at 24fps, so ComfyUI may adjust the requested value slightly.

Write the scene first, then timed action, camera movement, and audio in the same prompt block. A clean T2V brief is more reliable than a list of unrelated cinematic adjectives. For a faster hosted starting point, the Seedance Text to Video workflow removes the local graph and download steps.

MiniMax H3 T2V · official ComfyUI template output

Official ComfyUI template output for MiniMax H3 T2V. It is a real workflow sample, rehosted on Seedance R2 for reliable playback.

MiniMax H3 Image-to-Video (I2V) Workflow Setup

The I2V template uses MiniMaxH3ImageToVideo. Connect an image to first_frame, optionally connect another to last_frame, and describe the motion between them. Both inputs are optional at node level, so the same FL2VA weights can support text-led, first-frame, last-frame, or first-and-last-frame workflows.

I2V works best for product showcases, character continuity, controlled transitions, and style-led animation. Match the output to H3's 768px short-edge canvas, keep dimensions divisible by 32, and avoid feeding a tiny or heavily compressed source. Describe what must remain unchanged before describing camera motion.

Official transparent gaming-mouse input used by the MiniMax H3 I2V template

The real input image distributed with ComfyUI's MiniMax H3 I2V template.

MiniMax H3 I2V · official ComfyUI template output

The official I2V result from that source image. Compare product shape, materials, wheel placement, camera path, and sound—not just first-frame impact.

If you want the same input-to-motion idea without maintaining local weights, try Seedance Image to Video.

MiniMax H3 Reference-to-Video (R2V) Workflow Setup

R2V uses MiniMaxH3ReferenceToVideo and the separate ref2va diffusion weights. It accepts mixed images, videos, and audio so a target shot can borrow character identity, visual style, subject motion, camera movement, or voice from different sources.

Connect references in a deliberate order and name them exactly in the prompt: <Picture 1>, <Video 1>, and <Audio 1>. Assign one job to each reference. For example: preserve the character from Picture 1, use the dolly move from Video 1, and match the voice from Audio 1. ComfyUI's current workflow supports up to nine images, three videos, and three standalone audio clips.

Use ref_image_size=match for faster tests or max to preserve a reference short edge up to 2048px when identity matters more than speed. Too many overlapping references can weaken instruction following, so prove a two-reference graph before adding more.

MiniMax H3 R2V · official ComfyUI template output

Official R2V template output using the distributed character and style references.

For a hosted reference workflow, open Seedance Reference to Video. The MiniMax H3 reference-video guide provides a deeper prompt and reference-assignment walkthrough.

Prompt Writing Tips for Better MiniMax H3 Output

Use three explicit fields for H3 prompts: integrated_multimodal_description, overall_soundscape, and non_diegetic_music. Inside the scene description, write timed shots, visible action, camera direction, and exact dialogue. Give recurring speakers stable IDs such as (S1) and place only spoken words inside <d>[English] ...</d>.

T2VA — weak: A cinematic city scene with cool sound.

T2VA — better: 0–3s wide rainy crosswalk, slow dolly forward; 3–7s courier turns toward camera. Footsteps and traffic in stereo; low percussion begins at 4s.

I2VA/FL2VA — weak: Animate this image dramatically.

I2VA/FL2VA — better: Preserve the product geometry, wheel, buttons, transparent shell, and blue-orange lighting. One slow 30-degree orbit; no new text or parts. End aligned to the supplied last frame.

R2V — weak: Use all references and make an action video.

R2V — better: Keep identity and costume from <Picture 1>; copy only camera motion from <Video 1>; match the voice tone from <Audio 1>. Do not transfer the video subject or background.

T2VA needs a complete world description. I2VA should protect the source before adding motion. FL2VA and L2VA must describe the transition into the supplied boundary frame. R2V must explain the relationship between every reference and the target. The MiniMax H3 prompting guide expands these mode-specific patterns.

MiniMax H3 ComfyUI vs. Browser-Based AI Video Tools

Factor Local H3 in ComfyUI Seedance Comfy Cloud
Setup Install/update ComfyUI and place large weights Open in a browser Load a hosted ComfyUI template
Local VRAM Required; settings depend on precision and output Not required Not required locally
Workflow control Full node-level control and automation Streamlined generation interface ComfyUI graph without local hardware
Native audio H3 stereo audio in the same generation Joint audio-video workflow Same H3 workflow on cloud compute
Best for Deep customization, local inputs, reusable graphs Fast ideation and production without setup ComfyUI users who lack local GPU capacity

Choose local ComfyUI when the graph itself is part of your production system. Choose cloud-hosted ComfyUI when you want its nodes but not the hardware burden. Choose Seedance when you want to move from prompt or reference to a usable clip with fewer infrastructure decisions. Try Seedance in your browser →

Conclusion

MiniMax H3 is one of ComfyUI's most capable open-weights options for generating video and native audio together. Start in three steps: update to ComfyUI 0.30.0+, place the diffusion model, text encoder, and two VAEs in their correct folders, then load the matching T2V, I2V, or R2V template. Keep early tests small, assign every reference a job, and raise resolution only after the prompt works; if local downloads, VRAM, and node maintenance outweigh the benefit of deep control, a browser-based Seedance workflow is the faster route to production.

Ready to try it yourself?

Put the steps from this guide into practice with Seedance and turn prompts or images into polished videos in minutes.

Free credits on signup. Plans from $20/month.