- Overview (this page): model capabilities, workflow index, output resolution, step count, sampler and scheduler, and speedups
- Native workflows: Text-to-Video, Image-to-Video, Reference-to-Video, plus advanced native-node techniques
- Multiframe Reference: anchor reference frames at specific points along the output timeline
- Fun ControlNet Union: drive H3 with a control video, or run video inpainting with a mask
- Prompt guide: official MiniMax prompt writing guides, general tips, and prompt embeddings
- FastH3: the FastVideo FastH3 checkpoints and their sparse-attention workflow
H3’s open weights let you run the model locally. Commercial use of locally generated outputs requires a MiniMax commercial license, available through Comfy, the only official reseller. Generations on Comfy Cloud already include commercial rights.
Key features
- Native stereo audio: Dialogue, sound effects, and music are generated together with the video, synced in one MP4
- Multimodal context: Text, images, video, and audio references can be combined in one generation
- Reference-driven generation: Lock a character’s identity, a style, a motion, a camera move, or a voice from reference materials
- Instruction following: Describe the relationship between references and the target shot in natural language
- Accurate text rendering: Spelled-out text and brand elements render cleanly
- Open weights: Run locally in ComfyUI with full control over every parameter
Getting started
MiniMax H3 is supported in ComfyUI with open weights. To get started:- Update ComfyUI. The base Text to Video, Image to Video, and Reference to Video templates need 0.30.0 or later. Multiframe Reference needs 0.34.0. Fun ControlNet Union, the sparse attention node, and the Model Attention Backend node need 0.35.0, and FastH3 needs 0.36.0
- Go to Template Library > Video > choose any MiniMax H3 workflow
- Follow the pop-up to download models and run the workflow. The download is tens of gigabytes depending on the template, so allow disk space and time.
Match an H3 LoRA to the checkpoint build it was distilled on, and take the
pruned build when a repository publishes both. Pruned checkpoints replace the time embedder and the full-width adaln weights with a short shared curve basis (adaln_t_table, 1025 curve samples by 8 basis columns) that ComfyUI reads from the checkpoint itself, and the templates load the pruned builds (minimax_h3_fl2va_pruned_int8_convrot.safetensors and minimax_h3_ref2va_pruned_int8_convrot.safetensors, with bf16 and fp8 counterparts). A LoRA distilled on the full build (minimax_h3_fl2va_int8_convrot.safetensors or minimax_h3_ref2va_int8_convrot.safetensors) carries adaln weights with no matching tensor in the pruned build: ComfyUI reports a shape mismatch for them and does not merge them, so that part of the LoRA is skipped. The turbo LoRAs the workflows download carry no adaln tensors, so they load on either build.Workflow index
The template library currently ships with six example workflows. They are example templates, not an exhaustive list: the model supports more generation modes through the native MiniMax H3 nodes, and you can build additional workflows with them. The library also includes a continuation variant of the Image to Video template (video_minimax_h3_i2v_continuation).
Text to Video (T2V)
Generate videos from text prompts with native stereo audio
Image to Video (I2V)
Generate videos from an input image, with optional first/last-frame control
Reference to Video (R2V)
Lock in a character, style, motion, camera move, or voice from reference images, videos, and audio
Multiframe Reference
Anchor reference frames at specific points along the output timeline with chained Add Guide nodes
Fun ControlNet Union
Drive H3 with a Canny, Depth, HED, MLSD, or Pose control video, or run video inpainting with a mask
Setting the output resolution
Each workflow uses a Resolution Selector node to control the overall output size. The node computeswidth and height from three settings, and its outputs connect directly to the width and height inputs of the MiniMax H3 node:
- Aspect ratio: Pick a preset such as
16:9 (Widescreen),9:16 (Portrait Widescreen), or1:1 (Square) - Megapixels: Target total pixel count for the output. Higher values give larger frames; lower values run faster
- Multiple: The computed resolution is rounded to the nearest multiple of this number. Keep it at
32to match H3’s resolution grid
0.98 for H3’s native canvas (a 768px short edge, 1344x768 at 16:9), or enter 1344 x 768 directly in the MiniMax H3 node’s width and height inputs (its default). Skip the 1.0 Megapixel step: it yields 1376x768, above the model’s 768x1344 pixel area cap.
Frames well above the native canvas do not hold detail: H3 is trained at around 1 megapixel, and a later upscaling pass cannot recover what the generation step left out. For 2K output, generate at the native canvas and upscale in a separate pass.
Step count
These numbers apply to the base weights with the template’s Enable Lightning LoRA switch off, which is the default and runs 20 steps. Turning that switch on loads the turbo LoRA and drops the step count to 8 in the Text to Video and Image to Video templates (including the continuation variant), and to 4 in the templates that ship the reference turbo LoRA: Reference to Video, Multiframe Reference, and Fun ControlNet Union. In the Text to Video and Image to Video templates the switch appears as theturbo_mode widget on the subgraph node.
A short schedule costs the most on reference-driven shots. Reference tokens ride along every sampling step, so turning the switch on leaves far fewer steps for that conditioning to act on: at 4 steps a reference can end up barely applied, and a subject’s pose or face angle can drift away from the reference as the clip goes on. When a shot has to follow a reference closely, leave the Enable Lightning LoRA switch off and run the base 20-step schedule; if the reference still drifts, raise the step count to 25.
The step count a shot needs depends on what the shot contains:
- Shots with simple content hold up at 12 to 16 steps
- Shots with high-frequency detail (chainmail, filigree or floral patterns, piles of small objects) keep improving up to about 50 steps. Below that range, those areas show unstable triangular grid artifacts that swim under motion, and distilled step-reduction LoRAs and turbo checkpoints surface them first
- Higher counts also improve prompt adherence and motion, with most of the gain by step 16. Beyond about 50 the difference is hard to see
- Steps do not add sharpness the native canvas cannot resolve, so a soft frame is a resolution problem first
Sampler and scheduler
Every local MiniMax H3 workflow samples withres_multistep and the simple scheduler. The Text to Video, Image to Video, and FastH3 templates keep those two nodes inside the workflow’s subgraph, so open the subgraph to reach them. In the Reference to Video, Multiframe Reference, and Fun ControlNet Union templates, KSamplerSelect and BasicScheduler sit on the top-level canvas.
res_multistep is a second-order multistep sampler: each step reuses the denoised estimate from the step before it. The first step has no previous estimate to reuse, so it runs as an ordinary first-order (Euler) step, and the second-order steps begin with the second step. er_sde builds up its order in the same way, running one stage on its first step, two on the second, and, at its default max_stage of 3, all three stages from the third step on.
Multistep history lives inside a single sampling run rather than in the latent, so a workflow that splits the schedule across two samplers does not carry it over. In a two-sampler setup, such as a base stage followed by a second sampler after a latent upscale, the second sampler starts from an empty history: its first step runs at first order, and the second-order update resumes from its second step. A short tail is where that shows, because one of its few steps is spent at lower order. Keep the whole schedule in one sampling run, or give the second sampler enough steps to rebuild its history.
The flow-shift pair comes from the model definition rather than the workflow: ComfyUI’s H3 definition carries shift 12 and audio_shift 3, and every workflow except FastH3 reads those values, so no shift node appears in them. The FastH3 templates instead ship the built-in ModelSamplingMiniMaxH3 node (MiniMaxH3SigmaShift in the workflow JSON, category model/patch/minimax), set to shift_video 10 and shift_audio 3, the pair its distilled schedule is built around. The video shift drives the sampler’s sigma schedule, and the model inverts the video schedule onto the shared base grid to derive the audio schedule from it. Add the node when a checkpoint calls for a different pair, and stay near the values that checkpoint is built for: on the distilled 8-step builds, sampling at shift_video 3 instead of 10 shows grid artifacts in the frames.
Speeding up generation with Sage Attention
The example workflows use the standard attention implementation. You can roughly double the generation speed with Sage Attention, with minimal quality loss. Sage Attention is an optional dependency, so you need to install it yourself:- Install the
sageattentionPython package. Download the wheel that matches your PyTorch and CUDA versions from the SageAttention releases page, then install it withpip install <wheel-file>. - Install the KJNodes custom nodes, which provide the
Patch Sage Attention KJnode. Use the ComfyUI Manager, or clone the repository intoComfyUI/custom_nodes/and restart ComfyUI. - Add a
Patch Sage Attention KJnode to the workflow and connect it between theUNETLoaderand theBasicGuidernode: itsmodelinput receives the model from theUNETLoader, and itsmodeloutput feeds themodelinput of theBasicGuider. Setsage_attentiontoauto. - Run the workflow as usual. Only the guider needs the patch; the scheduler only generates the sigmas and can stay as is.
- Sage Attention requires float16 or bfloat16 tensors. MiniMax H3 runs some layers in other dtypes, so you may see “Input tensors must be in dtype of torch.float16 or torch.bfloat16, using pytorch attention instead” messages in the console. These are expected; the affected layers fall back to standard attention and generation still works.
- Alternatively, you can enable Sage Attention globally by launching ComfyUI with the
--use-sage-attentionflag instead of adding the node.
Quality degradation with INT8 attention
If you see morphing near the end of a clip or garbled on-screen text while using Sage Attention, the INT8 attention quantization is the likely cause. H3’s last blocks concentrate most of their attention-key signal into a few channels, and INT8 kernels that round each row with a single shared scale lose part of that signal. Which fix applies depends on the checkpoint the workflow loads:- The templates ship int8-convrot checkpoints:
minimax_h3_fl2va_pruned_int8_convrot.safetensorsfor T2V, I2V, and the Image to Video continuation variant, andminimax_h3_ref2va_pruned_int8_convrot.safetensorsfor R2V, Multiframe Reference, and Fun ControlNet Union. Comfy Kitchen attention does not support these checkpoints and sampling crashes with an alignment error (ComfyUI issue #15529). Keep the default attention with them, or load the bf16 counterpart in theUNETLoader(minimax_h3_fl2va_pruned_bf16.safetensorsorminimax_h3_ref2va_pruned_bf16.safetensors, which needs more VRAM) and then change the backend below. - With the bf16 checkpoints, switching the dense attention backend to Comfy Kitchen attention clears the artifacts. Its INT8 kernel applies a channel rotation before quantizing, which preserves that signal at a similar speed to Sage Attention.
model/patch) and set its backend to comfy kitchen attention. Connect it where the model reaches the BasicGuider, that is after the LoraLoaderModelOnly node and its Enable Lightning LoRA switch. Alternatively, launch ComfyUI with the --use-ck-attention flag (ComfyUI 0.32.0 or later). The backend is provided by the comfy-kitchen package that ships with ComfyUI, and it only appears in the node’s backend list when the INT8 kernels are available on your hardware.
Speeding up generation with sparse attention
Attention cost grows quickly with clip length. ComfyUI’s built-in Model Sparse Attention node (categorymodel/patch, experimental, ComfyUI 0.35.0 or later) runs block-sparse attention on eligible layers to cut that cost, and the gain grows with sequence length. Set its method to sol-attn, the mode the base H3 weights use (sla and vsa expect weights trained for those patterns). The sparse path needs CUDA and the sol_attn kernel from comfy-kitchen; without them every layer silently falls back to dense attention and you see no speedup.
- Keep sparse attention out of the first and last steps. The defaults are
start_percent0.2andend_percent1.0, so sparsity runs from 20% of the schedule through the final step. Starting later (around0.4) and ending earlier (around0.9) keeps the opening motion and the closing frames dense, at a modest cost in speed. - Leave
sink_conditioningat its default (exact_kv_and_rows) on H3. It keeps the packed text, audio, and reference rows exact and the generated audio query rows dense, so the sparse path does not degrade the audio track. taucontrols how sparsesol-attnis:1.0keeps roughly 16% of key blocks exact,1.5about 7%, and2.0about 2.7%. The default is1.3, and higher values are faster but riskier.- Short clips gain little. Sequences below
min_tokens(default12288) and blocks listed indense_blocksstay dense.
vsa mode, which is trained for its sparse pattern at keep_percent 10. See the FastH3 workflow page for details.