Inputs
Note: When a VAE is provided, the node generates reference latents from all provided input images. Up to three images can be processed at once. Images are scaled to a target area of 384x384 pixels (aspect ratio preserved) for vision-language processing, and to dimensions divisible by 8 (with a target area of 1024x1024 pixels) for VAE encoding.
Outputs
This documentation was AI-generated. If you find any errors or have suggestions for improvement, please feel free to contribute! Edit on GitHub
Source fingerprint (SHA-256):
5eea53a84045924b44d445244e6149b341188d22573aaaced87bac8a139dac96