Skip to main content
This node encodes an input image using DINOv3 and the Flux2 VAE to create positive and negative conditioning data for the TripoSplat model. It also generates a fixed-size noise target (latent plus camera data) that serves as the starting point for the KSampler.

Inputs

ParameterDescriptionData TypeRequiredRange
clip_visionDINOv3 ViT-H/16+ image encoderCLIP_VISIONYes-
vaeFlux2 VAEVAEYes-
imageThe input image to encodeIMAGEYes-

Outputs

Output NameDescriptionData Type
positivePositive conditioning data containing DINOv3 features and Flux2 VAE latentCONDITIONING
negativeNegative conditioning data containing zero-filled DINOv3 features and zero-filled Flux2 VAE latentCONDITIONING
latentThe fixed size noise target (latent sequence plus camera token) for the KSamplerLATENT
This documentation was AI-generated. If you find any errors or have suggestions for improvement, please feel free to contribute! Edit on GitHub

Source fingerprint (SHA-256): 9187a4a020818b9adc762eb41e913086b59d62c47abe92d4bafdb14bc8779f51