Skip to main content
This node prepares image conditioning for the Trellis2 3D generation pipeline. It extracts visual features from the input image with a DINOv3 vision model at two resolutions, organizes them into per-stage feature maps (optionally enhanced with a NAF model), and combines them with camera data derived from the horizontal field of view. It outputs a positive and a negative conditioning pair, where the negative uses zeroed features for classifier-free guidance.

Inputs

Note: The camera_angle_x value is converted to radians internally and used to compute the camera distance for the projection transform matrix. When the supplied vision model includes a NAF component, the node additionally produces high-resolution feature maps for the shape and texture stages.

Outputs

This documentation was AI-generated. If you find any errors or have suggestions for improvement, please feel free to contribute! Edit on GitHub

Source fingerprint (SHA-256): 3eba711620f6c56a21bbf7df89f8d406ce6f90908298b1a295a1dbbddd042472