Skip to main content
MiniMax H3 is MiniMax’s general-purpose, omni-modal generation model, now available as open weights. It jointly understands text, images, video, and audio in a single context, and generates video with native stereo audio: voice, sound effects, and music are modeled together in a single forward pass instead of being layered on afterward. The open weights generate at a 768-pixel short edge (about 1 megapixel) at 24fps for about 15 seconds. 2K output comes from MiniMax’s hosted service and needs a separate upscale pass in ComfyUI. ComfyUI natively supports MiniMax H3. The documentation is split across six pages:
  • Overview (this page): model capabilities, workflow index, output resolution, step count, sampler and scheduler, and speedups
  • Native workflows: Text-to-Video, Image-to-Video, Reference-to-Video, plus advanced native-node techniques
  • Multiframe Reference: anchor reference frames at specific points along the output timeline
  • Fun ControlNet Union: drive H3 with a control video, or run video inpainting with a mask
  • Prompt guide: official MiniMax prompt writing guides, general tips, and prompt embeddings
  • FastH3: the FastVideo FastH3 checkpoints and their sparse-attention workflow
Make sure your ComfyUI is updated.Workflows in this guide can be found in the Workflow Templates. If you can’t find them in the template, your ComfyUI may be outdated.If nodes are missing when loading a workflow, possible reasons:
  1. You are not using the latest ComfyUI version (Nightly version)
  2. Some nodes failed to import at startup
H3’s open weights let you run the model locally. Commercial use of locally generated outputs requires a MiniMax commercial license, available through Comfy, the only official reseller. Generations on Comfy Cloud already include commercial rights.

Key features

  • Native stereo audio: Dialogue, sound effects, and music are generated together with the video, synced in one MP4
  • Multimodal context: Text, images, video, and audio references can be combined in one generation
  • Reference-driven generation: Lock a character’s identity, a style, a motion, a camera move, or a voice from reference materials
  • Instruction following: Describe the relationship between references and the target shot in natural language
  • Accurate text rendering: Spelled-out text and brand elements render cleanly
  • Open weights: Run locally in ComfyUI with full control over every parameter

Getting started

MiniMax H3 is supported in ComfyUI with open weights. To get started:
  1. Update ComfyUI. The base Text to Video, Image to Video, and Reference to Video templates need 0.30.0 or later. Multiframe Reference needs 0.34.0. Fun ControlNet Union, the sparse attention node, and the Model Attention Backend node need 0.35.0, and FastH3 needs 0.36.0
  2. Go to Template Library > Video > choose any MiniMax H3 workflow
  3. Follow the pop-up to download models and run the workflow. The download is tens of gigabytes depending on the template, so allow disk space and time.
The model files are hosted on Hugging Face in the Comfy-Org/MiniMax-H3 repository.
Match an H3 LoRA to the checkpoint build it was distilled on, and take the pruned build when a repository publishes both. Pruned checkpoints replace the time embedder and the full-width adaln weights with a short shared curve basis (adaln_t_table, 1025 curve samples by 8 basis columns) that ComfyUI reads from the checkpoint itself, and the templates load the pruned builds (minimax_h3_fl2va_pruned_int8_convrot.safetensors and minimax_h3_ref2va_pruned_int8_convrot.safetensors, with bf16 and fp8 counterparts). A LoRA distilled on the full build (minimax_h3_fl2va_int8_convrot.safetensors or minimax_h3_ref2va_int8_convrot.safetensors) carries adaln weights with no matching tensor in the pruned build: ComfyUI reports a shape mismatch for them and does not merge them, so that part of the LoRA is skipped. The turbo LoRAs the workflows download carry no adaln tensors, so they load on either build.

Workflow index

The template library currently ships with six example workflows. They are example templates, not an exhaustive list: the model supports more generation modes through the native MiniMax H3 nodes, and you can build additional workflows with them. The library also includes a continuation variant of the Image to Video template (video_minimax_h3_i2v_continuation).

Text to Video (T2V)

Generate videos from text prompts with native stereo audio

Image to Video (I2V)

Generate videos from an input image, with optional first/last-frame control

Reference to Video (R2V)

Lock in a character, style, motion, camera move, or voice from reference images, videos, and audio

Multiframe Reference

Anchor reference frames at specific points along the output timeline with chained Add Guide nodes

Fun ControlNet Union

Drive H3 with a Canny, Depth, HED, MLSD, or Pose control video, or run video inpainting with a mask
Underlying node modes: the MiniMax H3 Image to Video node covers text-to-video and first/last-frame image-to-video (t2va and fl2va), and the MiniMax H3 Reference to Video node covers reference-driven generation with images, videos, and audio (ref2va). For prompt writing resources (official MiniMax guides, general tips, and prompt embeddings), see the prompt guide.

Setting the output resolution

Each workflow uses a Resolution Selector node to control the overall output size. The node computes width and height from three settings, and its outputs connect directly to the width and height inputs of the MiniMax H3 node:
  • Aspect ratio: Pick a preset such as 16:9 (Widescreen), 9:16 (Portrait Widescreen), or 1:1 (Square)
  • Megapixels: Target total pixel count for the output. Higher values give larger frames; lower values run faster
  • Multiple: The computed resolution is rounded to the nearest multiple of this number. Keep it at 32 to match H3’s resolution grid
The template ships with a fast preview size. For full-quality output at 16:9, set the Resolution Selector’s Megapixels to 0.98 for H3’s native canvas (a 768px short edge, 1344x768 at 16:9), or enter 1344 x 768 directly in the MiniMax H3 node’s width and height inputs (its default). Skip the 1.0 Megapixel step: it yields 1376x768, above the model’s 768x1344 pixel area cap. Frames well above the native canvas do not hold detail: H3 is trained at around 1 megapixel, and a later upscaling pass cannot recover what the generation step left out. For 2K output, generate at the native canvas and upscale in a separate pass.

Step count

These numbers apply to the base weights with the template’s Enable Lightning LoRA switch off, which is the default and runs 20 steps. Turning that switch on loads the turbo LoRA and drops the step count to 8 in the Text to Video and Image to Video templates (including the continuation variant), and to 4 in the templates that ship the reference turbo LoRA: Reference to Video, Multiframe Reference, and Fun ControlNet Union. In the Text to Video and Image to Video templates the switch appears as the turbo_mode widget on the subgraph node. A short schedule costs the most on reference-driven shots. Reference tokens ride along every sampling step, so turning the switch on leaves far fewer steps for that conditioning to act on: at 4 steps a reference can end up barely applied, and a subject’s pose or face angle can drift away from the reference as the clip goes on. When a shot has to follow a reference closely, leave the Enable Lightning LoRA switch off and run the base 20-step schedule; if the reference still drifts, raise the step count to 25. The step count a shot needs depends on what the shot contains:
  • Shots with simple content hold up at 12 to 16 steps
  • Shots with high-frequency detail (chainmail, filigree or floral patterns, piles of small objects) keep improving up to about 50 steps. Below that range, those areas show unstable triangular grid artifacts that swim under motion, and distilled step-reduction LoRAs and turbo checkpoints surface them first
  • Higher counts also improve prompt adherence and motion, with most of the gain by step 16. Beyond about 50 the difference is hard to see
  • Steps do not add sharpness the native canvas cannot resolve, so a soft frame is a resolution problem first
Audio settles later than the picture does. Video and audio are denoised together in one pass, so one schedule drives both, and the audio track keeps improving at step counts where the picture has stopped changing. Speech and voice timbre converge last: at 8 steps they are the weakest part of the output, and 12 steps and above keep the track usable.

Sampler and scheduler

Every local MiniMax H3 workflow samples with res_multistep and the simple scheduler. The Text to Video, Image to Video, and FastH3 templates keep those two nodes inside the workflow’s subgraph, so open the subgraph to reach them. In the Reference to Video, Multiframe Reference, and Fun ControlNet Union templates, KSamplerSelect and BasicScheduler sit on the top-level canvas. res_multistep is a second-order multistep sampler: each step reuses the denoised estimate from the step before it. The first step has no previous estimate to reuse, so it runs as an ordinary first-order (Euler) step, and the second-order steps begin with the second step. er_sde builds up its order in the same way, running one stage on its first step, two on the second, and, at its default max_stage of 3, all three stages from the third step on. Multistep history lives inside a single sampling run rather than in the latent, so a workflow that splits the schedule across two samplers does not carry it over. In a two-sampler setup, such as a base stage followed by a second sampler after a latent upscale, the second sampler starts from an empty history: its first step runs at first order, and the second-order update resumes from its second step. A short tail is where that shows, because one of its few steps is spent at lower order. Keep the whole schedule in one sampling run, or give the second sampler enough steps to rebuild its history. The flow-shift pair comes from the model definition rather than the workflow: ComfyUI’s H3 definition carries shift 12 and audio_shift 3, and every workflow except FastH3 reads those values, so no shift node appears in them. The FastH3 templates instead ship the built-in ModelSamplingMiniMaxH3 node (MiniMaxH3SigmaShift in the workflow JSON, category model/patch/minimax), set to shift_video 10 and shift_audio 3, the pair its distilled schedule is built around. The video shift drives the sampler’s sigma schedule, and the model inverts the video schedule onto the shared base grid to derive the audio schedule from it. Add the node when a checkpoint calls for a different pair, and stay near the values that checkpoint is built for: on the distilled 8-step builds, sampling at shift_video 3 instead of 10 shows grid artifacts in the frames.

Speeding up generation with Sage Attention

The example workflows use the standard attention implementation. You can roughly double the generation speed with Sage Attention, with minimal quality loss. Sage Attention is an optional dependency, so you need to install it yourself:
  1. Install the sageattention Python package. Download the wheel that matches your PyTorch and CUDA versions from the SageAttention releases page, then install it with pip install <wheel-file>.
  2. Install the KJNodes custom nodes, which provide the Patch Sage Attention KJ node. Use the ComfyUI Manager, or clone the repository into ComfyUI/custom_nodes/ and restart ComfyUI.
  3. Add a Patch Sage Attention KJ node to the workflow and connect it between the UNETLoader and the BasicGuider node: its model input receives the model from the UNETLoader, and its model output feeds the model input of the BasicGuider. Set sage_attention to auto.
  4. Run the workflow as usual. Only the guider needs the patch; the scheduler only generates the sigmas and can stay as is.
Notes:
  • Sage Attention requires float16 or bfloat16 tensors. MiniMax H3 runs some layers in other dtypes, so you may see “Input tensors must be in dtype of torch.float16 or torch.bfloat16, using pytorch attention instead” messages in the console. These are expected; the affected layers fall back to standard attention and generation still works.
  • Alternatively, you can enable Sage Attention globally by launching ComfyUI with the --use-sage-attention flag instead of adding the node.

Quality degradation with INT8 attention

If you see morphing near the end of a clip or garbled on-screen text while using Sage Attention, the INT8 attention quantization is the likely cause. H3’s last blocks concentrate most of their attention-key signal into a few channels, and INT8 kernels that round each row with a single shared scale lose part of that signal. Which fix applies depends on the checkpoint the workflow loads:
  • The templates ship int8-convrot checkpoints: minimax_h3_fl2va_pruned_int8_convrot.safetensors for T2V, I2V, and the Image to Video continuation variant, and minimax_h3_ref2va_pruned_int8_convrot.safetensors for R2V, Multiframe Reference, and Fun ControlNet Union. Comfy Kitchen attention does not support these checkpoints and sampling crashes with an alignment error (ComfyUI issue #15529). Keep the default attention with them, or load the bf16 counterpart in the UNETLoader (minimax_h3_fl2va_pruned_bf16.safetensors or minimax_h3_ref2va_pruned_bf16.safetensors, which needs more VRAM) and then change the backend below.
  • With the bf16 checkpoints, switching the dense attention backend to Comfy Kitchen attention clears the artifacts. Its INT8 kernel applies a channel rotation before quantizing, which preserves that signal at a similar speed to Sage Attention.
To switch backends, add the built-in Model Attention Backend node (category model/patch) and set its backend to comfy kitchen attention. Connect it where the model reaches the BasicGuider, that is after the LoraLoaderModelOnly node and its Enable Lightning LoRA switch. Alternatively, launch ComfyUI with the --use-ck-attention flag (ComfyUI 0.32.0 or later). The backend is provided by the comfy-kitchen package that ships with ComfyUI, and it only appears in the node’s backend list when the INT8 kernels are available on your hardware.

Speeding up generation with sparse attention

Attention cost grows quickly with clip length. ComfyUI’s built-in Model Sparse Attention node (category model/patch, experimental, ComfyUI 0.35.0 or later) runs block-sparse attention on eligible layers to cut that cost, and the gain grows with sequence length. Set its method to sol-attn, the mode the base H3 weights use (sla and vsa expect weights trained for those patterns). The sparse path needs CUDA and the sol_attn kernel from comfy-kitchen; without them every layer silently falls back to dense attention and you see no speedup.
  • Keep sparse attention out of the first and last steps. The defaults are start_percent 0.2 and end_percent 1.0, so sparsity runs from 20% of the schedule through the final step. Starting later (around 0.4) and ending earlier (around 0.9) keeps the opening motion and the closing frames dense, at a modest cost in speed.
  • Leave sink_conditioning at its default (exact_kv_and_rows) on H3. It keeps the packed text, audio, and reference rows exact and the generated audio query rows dense, so the sparse path does not degrade the audio track.
  • tau controls how sparse sol-attn is: 1.0 keeps roughly 16% of key blocks exact, 1.5 about 7%, and 2.0 about 2.7%. The default is 1.3, and higher values are faster but riskier.
  • Short clips gain little. Sequences below min_tokens (default 12288) and blocks listed in dense_blocks stay dense.
FastH3 checkpoints use the same node in vsa mode, which is trained for its sparse pattern at keep_percent 10. See the FastH3 workflow page for details.