Skip to main content
This node merges a video latent and an audio latent into a single joint audio-video (AV) latent, ready for AV models such as LTXV or MiniMax H3. If the video input is already an AV latent, its video stream is kept and only the audio stream is replaced with the supplied audio latent.

Inputs

Note: The samples from both inputs are combined as a pair of video and audio streams in a nested tensor. If either input contains a noise_mask, the output includes a combined one; a missing mask is replaced with an all-ones mask matching the shape of its samples. When shorter audio is padded, the padded region is left unmasked so the model can generate it. The node raises an error if the audio latent cannot be fitted to the video latent, for example when the two latents differ in more than one dimension or when they differ in the batch or channel dimensions.

Outputs

This documentation was AI-generated. If you find any errors or have suggestions for improvement, please feel free to contribute! Edit on GitHub

Source fingerprint (SHA-256): 0231f9db2ce73132d8555fbb33f295b68aa68a0c1c54e4a0c5d2e1f67b5611cb