Feedback
AI Ad Video Example
Loading...
comfyui minimax h3
Build 2K/24fps clips with the comfyui minimax h3 workflow — text, image, and reference prompts plus synced stereo audio, all inside ComfyUI nodes.
All Tools
Discover our comprehensive AI-powered animation toolkit
MiniMax H3
MiniMax H3 AI Video Generator

Seedance 2.0
The Future of AI Video Is Here.

Kling 3.0
Next-Gen AI Video Generator
Grok Video Generator
Create Videos from Text or Images with AI
MiniMax H3 video generator
MiniMax H3 AI Video Generator

Seedance 2.0
The Future of AI Video Is Here.

Kling Motion Control
Turn reference images into amazing motion videos in minutes

Suno AI Music Generator
Create Professional Music with AI
A Closer Look at comfyui minimax h3 in ComfyUI
comfyui minimax h3 brings MiniMax's omni-modal engine into ComfyUI as downloadable open weights. One shared context reads text, stills, footage, and audio together, so dialogue, effects, and music are synthesized alongside the picture in a single forward pass. Clips land at up to 2K and 24fps for roughly fifteen seconds, and every diffusion parameter stays editable at the node level.
- Stereo Sound Baked InSpeech, effects, and score are synthesized alongside the picture and muxed into one MP4, perfectly in sync straight out of the comfyui minimax h3 nodes.
- Runs on Your Own MachineLoad the comfyui minimax h3 checkpoint locally and tune resolution, clip length, and diffusion settings yourself — nothing is capped by a hosted API.
- Mix Text, Image, Video & VoiceFeed prompts, stills, footage, and voice samples into a single run to hold a face, art direction, movement, camera path, or timbre steady across the comfyui minimax h3 graph.
Three Steps to Run comfyui minimax h3 in ComfyUI
Follow three short steps to produce open-weight clips with built-in audio through the comfyui minimax h3 pipeline.
Capabilities Built Into the comfyui minimax h3 Nodes
From three ready-made ComfyUI templates to open-weight multimodal prompting, built-in stereo sound, reference-locked control, and optional Sage Attention acceleration, the comfyui minimax h3 stack covers local video production end to end.
Three Ready-Made Templates
Open the comfyui minimax h3 template shelf and you get text-to-video, image-to-video, and reference-to-video graphs, one per generation mode, usable straight away.
One Shared Context Window
Text, stills, footage, and sound all land in the same comfyui minimax h3 context, so every reference type can steer a single render at once.
Lock Down Faces, Style & Motion
Hold a face, art style, movement, camera path, or timbre steady using supplied material — the comfyui minimax h3 R2V node accepts as many as 9 images, 3 videos, and 3 audio clips.
Crisp On-Screen Text and Logos
Lettering and brand marks come out legible under the comfyui minimax h3 model, and you can describe how references relate to each other in plain language.
Optional Sage Attention Boost
Drop a Patch Sage Attention KJ node into the comfyui minimax h3 graph to roughly halve render time while barely touching output quality.
Snapped Resolution and Duration
The comfyui minimax h3 Resolution Selector derives width and height from your ratio and megapixel target, rounding to the model's 32-pixel grid and 17-frame blocks at 24fps.
comfyui minimax h3 — Questions Answered
Answers to the questions people ask most about running MiniMax H3 inside ComfyUI.
What exactly is comfyui minimax h3?
It is ComfyUI's built-in bridge to MiniMax H3, an omni-modal generator that MiniMax published as open weights. Running it lets you turn prompts, stills, clips, and audio into a video with synced stereo sound in one forward pass.
How high can the output resolution go?
Clips can reach 2K at 24fps and run for roughly 15 seconds. The native canvas keeps a 768px short edge, tops out at 768x1344 pixels, and snaps dimensions to multiples of 32.
Which generation modes ship with it?
Three: text-to-video (T2V), image-to-video (I2V) with optional first- and last-frame anchors, and reference-to-video (R2V), which pins down a character, style, motion, camera path, or voice.
Does it produce sound as well?
It does. Speech, effects, and music are modeled alongside the picture inside the comfyui minimax h3 graph and arrive already synced in the same MP4.
What do I need to get started?
Move to ComfyUI 0.30.0 or newer, open Template Library > Video, select a comfyui minimax h3 graph, and accept the pop-up that pulls weights from the Hugging Face Comfy-Org/MiniMax-H3 repository.
Is there a way to make rendering faster?
Install SageAttention plus the KJNodes pack, then insert a Patch Sage Attention KJ node between the UNETLoader and BasicGuider in your comfyui minimax h3 graph — render time drops by about half.
Ready to Run comfyui minimax h3 Yourself?
Load MiniMax H3 on your own machine, keep every parameter open, and get stereo sound baked into each render — text, image, and reference graphs are waiting for you.
