Skip to main content
ERNIE-Image is an open text-to-image model by Baidu, licensed under Apache-2.0. Built on an 8B parameter Diffusion Transformer (DiT), it delivers high-quality image generation with precise text rendering, strong instruction following, and structured visual generation. The model includes a built-in Prompt Enhancer (3B) that expands short inputs into richer prompts for better results. Model highlights:
  • Precise text rendering: dense and layout-sensitive text in English, Chinese, and more
  • Strong instruction following: handles complex prompts, multi-object relations, and knowledge-intensive descriptions
  • Structured visual generation: posters, manga/anime storyboards, multi-panel compositions
  • Broad stylistic range: realistic photography to cinematic film-like aesthetics
  • Compact and deployable: 8B parameters, runs on 24 GB VRAM
  • Built-in Prompt Enhancer: 3B model that expands short inputs into richer prompts
Related links:

ERNIE-Image text-to-image workflow

Make sure your ComfyUI is updated.Workflows in this guide can be found in the Workflow Templates. If you canโ€™t find them in the template, your ComfyUI may be outdated.If nodes are missing when loading a workflow, possible reasons:
  1. You are not using the latest ComfyUI version (Nightly version)
  2. Some nodes failed to import at startup

Ernie Image: Text to Image

Generate images from text prompts using the ERNIE-Image model. Input a text description to produce detailed, structured visuals with a broad stylistic range. ERNIE-Image text-to-image workflow preview

Run on Comfy Cloud

Run this workflow directly on Comfy Cloud

Download Workflow

Download the ERNIE-Image text-to-image workflow JSON file
Example output ERNIE-Image example output

Get started

  1. Update ComfyUI to the latest version or use Comfy Cloud
  2. Go to Template and search for ERNIE-Image
  3. Select the ERNIE-Image workflow
  4. Download any missing models, update the prompt, and click Run

ERNIE-Image model downloads

You can find all repackaged model files at Comfy-Org/ERNIE-Image on Hugging Face.

ernie-image.safetensors

Diffusion model for ERNIE-Image.

ministral-3-3b.safetensors

Text encoder for ERNIE-Image.

ernie-image-prompt-enhancer.safetensors

Prompt Enhancer text encoder for ERNIE-Image.

flux2-vae.safetensors

VAE for ERNIE-Image.
Model storage location

ERNIE-Image-Turbo

ERNIE-Image-Turbo is a faster variant optimized with DMD and RL, generating images in just 8 steps compared to the ~50 steps required by the standard model.

Ernie Image Turbo: Text To Image

Generate images from text prompts using the ERNIE-Image turbo model. Input a text description and receive a high-quality image with precise text rendering. ERNIE-Image-Turbo text-to-image workflow preview

Run on Comfy Cloud

Run this workflow directly on Comfy Cloud

Download Workflow

Download the ERNIE-Image-Turbo text-to-image workflow JSON file
Example output ERNIE-Image-Turbo example output

ERNIE-Image-Turbo model downloads

ernie-image-turbo.safetensors

Diffusion model for ERNIE-Image-Turbo.

ministral-3-3b.safetensors

Text encoder for ERNIE-Image-Turbo.

ernie-image-prompt-enhancer.safetensors

Prompt Enhancer text encoder for ERNIE-Image-Turbo.

flux2-vae.safetensors

VAE for ERNIE-Image-Turbo.
Model storage location

Available models

Examples

Text rendering and design layouts

Coffee making process infographic
Japanese Ukiyo-e language flashcard
Split-screen conceptual poster

Cinematic and stylized aesthetics

Urban night street scene
Editorial fashion photograph
3D illustration of a duck with coffee

Multi-panel compositions

North American native species infographic
6-panel comic page