Inputs
Note: The samples from both inputs are combined as a pair of video and audio streams in a nested tensor. If either input contains a
noise_mask, the output includes a combined one; a missing mask is replaced with an all-ones mask matching the shape of its samples. When shorter audio is padded, the padded region is left unmasked so the model can generate it. The node raises an error if the audio latent cannot be fitted to the video latent, for example when the two latents differ in more than one dimension or when they differ in the batch or channel dimensions.
Outputs
This documentation was AI-generated. If you find any errors or have suggestions for improvement, please feel free to contribute! Edit on GitHub
Source fingerprint (SHA-256):
0231f9db2ce73132d8555fbb33f295b68aa68a0c1c54e4a0c5d2e1f67b5611cb