Skip to main content
OmniGen2 is an open-source unified multimodal generation model from the VectorSpaceLab research team. A single set of weights handles text-to-image generation, instruction-guided image editing, and subject-driven in-context generation. ComfyUI supports OmniGen2 natively, so all three tasks run locally through the official example workflows. The model pairs a frozen 3B Qwen2.5-VL vision-language model with a separate 4B diffusion transformer, about 7B parameters in total. The vision-language model reads your reference images and written instructions, and the diffusion transformer turns that understanding into pixels. Because both paths keep their own parameters and image tokenizer, OmniGen2 preserves its visual understanding ability while keeping image quality high. ComfyUI ships two official workflows for the model: text to image, and image editing from natural language instructions such as “change the dress to blue”. The image editing workflow also handles multi-image composition, combining people, reference objects, and scenes from several input images.
Make sure your ComfyUI is updated.Workflows in this guide can be found in the Workflow Templates. If you can’t find them in the template, your ComfyUI may be outdated.If nodes are missing when loading a workflow, possible reasons:
  1. You are not using the latest ComfyUI version (Nightly version)
  2. Some nodes failed to import at startup
Model highlights:
  • Visual understanding: Reads and analyzes image content through the Qwen2.5-VL foundation model
  • Text-to-image generation: Creates high-fidelity images from text prompts
  • Instruction-guided image editing: Applies complex edits from natural language instructions
  • In-context generation: Combines people, reference objects, and scenes from multiple inputs into new images
Technical features:
  • Dual-path architecture: 3B Qwen2.5-VL text model plus an independent 4B diffusion transformer
  • Omni-RoPE position encoding: Multi-image spatial positioning and identity distinction
  • Parameter decoupling: Keeps text autoregressive generation from degrading image quality
  • In-image text: Generates legible text content inside the generated image

OmniGen2 Model Download

Since this article involves different workflows, the corresponding model files and installation locations are as follows. The download information for model files is also included in the corresponding workflows: Diffusion Models VAE Text Encoders File save location:

ComfyUI OmniGen2 Text-to-Image Workflow

OmniGen2: Text to Image

Generate high-quality images from text prompts using OmniGen2’s unified 7B multimodal model with dual-path architecture. OmniGen2 text-to-image workflow preview

Run on Comfy Cloud

Open and run this workflow directly in Comfy Cloud

Download Workflow

Download JSON or search “OmniGen2” in Template Library
Example output OmniGen2 text-to-image example output

1. Download Workflow File

2. Complete Workflow Step by Step

Workflow Step Guide Please follow the numbered steps in the image for step-by-step confirmation to ensure smooth operation of the corresponding workflow:
  1. Load Main Model: Ensure the Load Diffusion Model node loads omnigen2_fp16.safetensors
  2. Load Text Encoder: Ensure the Load CLIP node loads qwen_2.5_vl_fp16.safetensors
  3. Load VAE: Ensure the Load VAE node loads ae.safetensors
  4. Set Image Dimensions: Set the generated image dimensions in the EmptySD3LatentImage node (recommended 1024x1024)
  5. Input Prompts:
    • Input positive prompts in the first CLipTextEncode node (content you want to appear in the image)
    • Input negative prompts in the second CLipTextEncode node (content you don’t want to appear in the image)
  6. Start Generation: Click the Queue Prompt button, or use the shortcut Ctrl(cmd) + Enter to execute text-to-image generation
  7. View Results: After generation is complete, the corresponding images will be automatically saved to the ComfyUI/output/ directory, and you can also preview them in the SaveImage node

ComfyUI OmniGen2 Image Editing Workflow

OmniGen2 has rich image editing capabilities and supports adding text to images.

OmniGen2 Image Edit

Edit images with natural language instructions using OmniGen2’s advanced image editing capabilities and text rendering support. OmniGen2 image edit workflow preview

Run on Comfy Cloud

Open and run this workflow directly in Comfy Cloud

Download Workflow

Download JSON or search “OmniGen2 Image Edit” in Template Library
Input materials Upload this file to the matching LoadImage node:

image_omnigen2_image_edit_input_image.png

LoadImage node 16 · image_omnigen2_image_edit_input_image.png
image_omnigen2_image_edit_input_image.png
Example output
Input imageOmniGen2 image edit example output

1. Download Workflow File

2. Complete Workflow Step by Step

Workflow Step Guide
  1. Load Main Model: Ensure the Load Diffusion Model node loads omnigen2_fp16.safetensors
  2. Load Text Encoder: Ensure the Load CLIP node loads qwen_2.5_vl_fp16.safetensors
  3. Load VAE: Ensure the Load VAE node loads ae.safetensors
  4. Upload Image: Upload the provided image in the Load Image node
  5. Input Prompts:
    • Input positive prompts in the first CLipTextEncode node (content you want to appear in the image)
    • Input negative prompts in the second CLipTextEncode node (content you don’t want to appear in the image)
  6. Start Generation: Click the Queue Prompt button, or use the shortcut Ctrl(cmd) + Enter to execute text-to-image generation
  7. View Results: After generation is complete, the corresponding images will be automatically saved to the ComfyUI/output/ directory, and you can also preview them in the SaveImage node

3. Additional Workflow Instructions

  • If you want to enable the second image input, you can use the shortcut Ctrl + B to enable the corresponding node inputs for nodes that are in pink/purple state in the workflow
  • If you want to customize dimensions, you can delete the Get image size node linked to the EmptySD3LatentImage node and input custom dimensions