Feedback
AI Ad Video Example
Loading...
comfyui minimax h3
Generate 2K clips at 24fps inside ComfyUI using the comfyui minimax h3 integration — with native stereo audio from text, image, or reference prompts.
All Tools
Discover our comprehensive AI-powered animation toolkit
MiniMax H3
MiniMax H3 AI Video Generator
Seedance 2.5
The Future of AI Video Is Here.

Seedance 2.0
The Future of AI Video Is Here.

Kling 3.0
Next-Gen AI Video Generator
MiniMax H3 video generator
MiniMax H3 AI Video Generator

Seedance 2.0
The Future of AI Video Is Here.

Kling Motion Control
Turn reference images into amazing motion videos in minutes
Sora2 AI
Advanced AI Video Generator for High-Quality Videos
Why Choose the comfyui minimax h3 Workflow
By leveraging the comfyui minimax h3 integration, you can run MiniMax's open-weight, omni-modal generation model inside ComfyUI. This pipeline processes text, images, video, and audio together in one context, producing clips with native stereo sound — dialogue, effects, and music all generated in the same pass. You can achieve up to 2K resolution at 24fps for roughly 15 seconds, with granular node-level control over every setting.
- Full Stereo SoundSpeech, sound effects, and music are rendered alongside the visuals into a single MP4, perfectly aligned in one run using the comfyui minimax h3 workflow.
- Complete Local AutonomyExecute the comfyui minimax h3 model on your own machine, adjusting resolution, duration, and all diffusion parameters freely — with zero API restrictions.
- Versatile Input FusionBlend textual, visual, video, and audio references in a single run, anchoring a character, aesthetic, movement, camera angle, or voice through the comfyui minimax h3 nodes.
A 3-Step Guide to Using the comfyui minimax h3 Workflow
Produce open-weight videos with coordinated stereo audio in three simple steps by following the comfyui minimax h3 workflow.
Essential Features Delivered by the comfyui minimax h3 Integration
This comfyui minimax h3 setup provides three native ComfyUI graph templates, open-weight multimodal processing, synchronized stereo audio, reference-driven generation, and an optional Sage Attention boost — everything you need for a robust local video studio.
Ready-to-Use Templates for Each Mode
The comfyui minimax h3 template set includes three separate examples for text-to-video, image-to-video, and reference-to-video generation, so each mode runs straight away.
Cross-Modal Context Understanding
The comfyui minimax h3 model can process text, images, footage, and audio simultaneously in one context, enabling you to combine any mix of reference inputs in a single generation.
Anchoring via Reference Media
You can fix a character's appearance, artistic style, action, camera angle, or voice from supplied references — the comfyui minimax h3 R2V node accepts up to 9 pictures, 3 videos, and 3 audio clips.
Precise Text and Brand Rendering
The comfyui minimax h3 model renders spelled-out words and brand logos cleanly, using natural-language instructions that clarify how references relate to each other.
Sage Attention Acceleration
Place the Patch Sage Attention KJ node inside the comfyui minimax h3 graph and you can nearly double the processing speed while retaining high visual quality.
Smart Resolution and Duration Grid
The comfyui minimax h3 Resolution Selector automatically calculates width and height from aspect ratio and megapixels, snapping to the required 32-multiple grid and 17-frame block duration at 24fps.
All About the comfyui minimax h3 Workflow: FAQ
Answers to common questions about using MiniMax H3 inside ComfyUI through the comfyui minimax h3 workflow.
What does the comfyui minimax h3 workflow do?
It is the native way to run MiniMax H3, an open-weight omni-modal generation model, inside ComfyUI. The comfyui minimax h3 integration lets you create video with built-in stereo audio from text, images, footage, and audio references in one forward pass.
What are the supported output specs?
Through the comfyui minimax h3 workflow, you can generate clips up to 2K resolution, 24fps, and about 15 seconds long. The default output canvas has a 768-pixel short edge, a maximum of 768x1344 pixels, and dimensions rounded to multiples of 32.
Which modes come with the templates?
The comfyui minimax h3 template collection offers three built-in examples: text-to-video (T2V), image-to-video (I2V) with optional first/last-frame settings, and reference-to-video (R2V) that can lock a character, style, motion, camera movement, or voice.
Does the output include audio?
Yes. The comfyui minimax h3 model generates stereo audio — speech, sound effects, and music — together with the visuals in the same pass, and all tracks are synced into one MP4.
How do I set it up for the first time?
Start by updating ComfyUI to version 0.30.0 or later. Then open Template Library > Video, select a comfyui minimax h3 template, and follow the pop-up prompts to download the required open weights from the Hugging Face repository Comfy-Org/MiniMax-H3.
Any way to improve generation speed?
Yes, you can install SageAttention and the KJNodes custom nodes. Then insert a Patch Sage Attention KJ node between the UNETLoader and BasicGuider in your comfyui minimax h3 graph to roughly double rendering speed.
Jump Into Video Creation with the comfyui minimax h3 Workflow
Operate MiniMax H3 directly within ComfyUI using the comfyui minimax h3 workflow — with native stereo audio, open-weight access, and complete parameter control, plus ready-to-run T2V, I2V, and R2V templates.
