Skip to main content
TextEncodeZImageOmni encodes a text prompt together with up to three optional reference images into a conditioning format for image generation models. The prompt is tokenized and encoded with the CLIP model, and each connected image can optionally be processed by a vision encoder and/or a VAE so that visual references are embedded alongside the text. This node is marked as experimental.

Inputs

Note: The node accepts a maximum of three images (image1, image2, image3). The image_encoder and vae inputs are only used when at least one image is provided; when both are connected, each image is processed by both. When auto_resize_images is True and a vae is connected, images are resized to have a total pixel area close to 1024x1024 before encoding.

Outputs

This documentation was AI-generated. If you find any errors or have suggestions for improvement, please feel free to contribute! Edit on GitHub

Source fingerprint (SHA-256): b40a3150f536b6f37e2b53e6d9992fcb4fd32dceb540c0a76773a7ba1af9a7b8