Skip to main content
MiniMax H3 Reference to Video creates the text conditioning and the empty audio-video latent needed for MiniMax H3 reference-to-video generation. You provide a prompt plus optional reference images, videos, and audio clips, and the node encodes these references into tokens the model can use while generating. The prompt refers to the references with <Picture i>, <Video k>, and <Audio j> tags.

Inputs

Notes:
  • The prompt refers to reference media with 1-based tags per type: <Picture i> for images, <Video k> for videos, and <Audio j> for audio. References are presented to the model in a fixed order: images, then videos (with each soundtrack’s <Audio j> label right before its <Video k>), then standalone audio.
  • Reference videos must contain at least 5 frames (~0.2 seconds at 24 fps), otherwise the node raises an error. Video frames are also capped to the selected length and trimmed to a supported frame count.
  • The requested length is aligned to a supported frame count before the latent is created.

Outputs

This documentation was AI-generated. If you find any errors or have suggestions for improvement, please feel free to contribute! Edit on GitHub

Source fingerprint (SHA-256): d9a444e712cdc255d7c56a3ab38d0523659f198b3228b9283a7028cfd0e4f3f9