Skip to main content
Pixal3D is an open-source image-to-3D model from Tencent ARC, presented at SIGGRAPH 2026. It turns a single image into a high-fidelity 3D asset with full PBR materials. Most 3D-native generators synthesize shapes in a canonical space and inject image cues through attention, which leaves pixel-to-3D associations ambiguous. Pixal3D instead uses pixel-aligned generation: it establishes direct pixel-to-3D correspondence through back-projection, so the front of the generated model matches your input image 1:1, with no warped or misaligned textures.

Pixal3D: Image to Model

Upload a single image. Generate a high-fidelity 3D model with full PBR textures, aligned to your input view. This workflow also includes the TRELLIS.2 pipeline. The Boolean (Switch to Trellis2) node defaults to false, which runs the Pixal3D pipeline and loads the Pixal3D model automatically. Pixal3D workflow preview
Make sure your ComfyUI is updated.Workflows in this guide can be found in the Workflow Templates. If you can’t find them in the template, your ComfyUI may be outdated.If nodes are missing when loading a workflow, possible reasons:
  1. You are not using the latest ComfyUI version (Nightly version)
  2. Some nodes failed to import at startup

Run on Comfy Cloud

Run this workflow instantly on Comfy Cloud

Download Workflow

Download JSON or search “Pixal3D & TRELLIS.2: Image to Model” in Template Library
Input materials Upload this file to the LoadImage node:

viking_wolf_rune_axe.png

LoadImage node 122 · viking_wolf_rune_axe.png
Input image

How it works

Pixal3D combines camera-aware, pixel-aligned generation with a complete mesh post-processing pipeline:
  1. Background removal: BiRefNet removes the background from the input image, and the workflow crops the subject to a centered canvas. A switch lets you skip background removal
  2. Camera estimation: MoGe estimates geometry and the camera field of view from the image. The FOV drives the pixel-aligned conditioning
  3. Structure generation: a sparse structure latent is sampled and decoded into voxels, then converted into a rough mesh
  4. Shape refinement: the shape stage and the upsampling stage refine the mesh up to the target resolution (1536)
  5. Texture generation: a texture diffusion stage produces PBR material voxels (base color, metallic, roughness)
  6. Post-processing: DC remesh, QEM decimation, UV unwrapping, and baking of base color, normal, and ambient occlusion maps into the final textured mesh

Steps to run

  1. Load an image: use the LoadImage node to load a single image of the object
  2. Queue the workflow: press Ctrl (Cmd on macOS) + Enter
  3. Wait for the pipeline: the structure, shape, and texture stages run in sequence, followed by post-processing
  4. View the result: inspect the mesh in the Preview3DAdvanced node. The 3D model is saved to ComfyUI/output/3d/ComfyUI/

Model downloads

Download the models used by this workflow. Both diffusion checkpoints are required: the switch selects which one runs. Place them in the corresponding models/ subdirectories.

Pixal3D UNet

pixal3d_int8_convrot.safetensors: Pixal3D diffusion model, loaded by default

TRELLIS.2 UNet

trellis_2_int8_convrot.safetensors: TRELLIS.2 diffusion model, loaded when the switch is set to true

Shape VAE

trellis_2_shape_vae_bf16.safetensors: VAE for structure and shape decoding

Texture VAE

trellis_2_texture_vae_bf16.safetensors: VAE for texture decoding

DINOv3 CLIP vision

dino_v3_L_naf_fp32.safetensors: CLIP vision encoder for image conditioning

MoGe geometry

moge_2_vitl_normal_fp16.safetensors: depth and camera estimation for pixel-aligned conditioning

BiRefNet background removal

birefnet.safetensors: background removal model for preprocessing

Model storage location