Skip to main content
The LTXVSeparateAVLatent node splits a combined audio-visual latent into two separate latents: one containing the video data and one containing the audio data. This works with any audio-visual model, such as LTXV or MiniMax H3. The samples tensor is split along its first dimension, with the first element becoming the video latent and the second element becoming the audio latent; if a noise mask is present, it is split in the same way.

Inputs

Note: The input latent’s samples tensor is expected to have at least two elements along the first dimension (batch dimension). The first element is used for the video latent, and the second element is used for the audio latent. If a noise_mask is present, it is split in the same way.

Outputs

This documentation was AI-generated. If you find any errors or have suggestions for improvement, please feel free to contribute! Edit on GitHub

Source fingerprint (SHA-256): 22ed38bbc1b5716cee380c35c50455810f79c273f51bbe6a535c9ae33192afe6