GitHub
Cosmos-Predict2 source code and documentation
Hugging Face
Cosmos-Predict2 model collection
Cosmos Predict2 Text to Image
Using Cosmos-Predict2 for text-to-image generation
Run on Comfy Cloud
Run Cosmos-Predict2 workflows on Comfy Cloud with powerful GPUs
Cosmos Predict2 Video2World Workflow
When testing the 2B version, it takes around 16GB VRAM.1. Workflow File
Please download the video below and drag it into ComfyUI to load the workflow. The workflow already has embedded model download links.Run on Comfy Cloud
Run this workflow on Comfy Cloud with pre-installed models
Download Workflow File
Download the JSON format workflow file
2. Manual Model Installation
If the model download was not successful, you can try to download them manually by yourself in this section. Diffusion modelDiffusion Model
cosmos_predict2_2B_video2world_480p_16fps.safetensors
Text Encoder
oldt5_xxl_fp8_e4m3fn_scaled.safetensors
VAE
wan_2.1_vae.safetensors
3. Complete Workflow Step by Step

- Ensure the
Load Diffusion Modelnode has loadedcosmos_predict2_2B_video2world_480p_16fps.safetensors - Ensure the
Load CLIPnode has loadedoldt5_xxl_fp8_e4m3fn_scaled.safetensors - Ensure the
Load VAEnode has loadedwan_2.1_vae.safetensors - Upload the provided input image in the
Load Imagenode - (Optional) If you need first and last frame control, use the shortcut
Ctrl(cmd) + Bto enable last frame input - (Optional) You can modify the prompts in the
ClipTextEncodenode - (Optional) Modify the size and frame count in the
CosmosPredict2ImageToVideoLatentnode - Click the
Runbutton or use the shortcutCtrl(cmd) + Enterto run the workflow - Once generation is complete, the video will automatically save to the
ComfyUI/output/directory, you can also preview it in thesave videonode