Inputs
Note: When a reference image is provided, it is encoded into a latent that gets attached to the positive conditioning, while a zero-filled latent of the same shape is attached to the negative conditioning. When audio encoder output is provided, the audio embeddings are interpolated and attached to the positive conditioning, while a zero-filled audio embedding is attached to the negative conditioning. If the optional inputs are omitted, zero-filled placeholder tensors are used for both reference latents and audio embeddings.
Outputs
This documentation was AI-generated. If you find any errors or have suggestions for improvement, please feel free to contribute! Edit on GitHub
Source fingerprint (SHA-256):
db674a4a00729a8715988030083e2858f958cd21de73bbbe4ed6d76f5f539419