About Wan2.1-Fun-InP
Wan-Fun InP is an open-source video generation model released by Alibaba, part of the Wan2.1-Fun series, focusing on generating videos from images with first and last frame control. Key features:- First and last frame control: Supports inputting both first and last frame images to generate transitional video between them, enhancing video coherence and creative freedom. Compared to earlier community versions, Alibaba’s official model produces more stable and significantly higher quality results.
- Multi-resolution support: Supports generating videos at 512x512, 768x768, 1024x1024 and other resolutions to accommodate different scenario requirements.
- 1.3B Lightweight: Suitable for local deployment and quick inference with lower VRAM requirements
- 14B High-performance: Model size reaches 32GB+, offering better results but requiring higher VRAM
- Wan2.1-Fun-1.3B-Input
- Wan2.1-Fun-14B-Input
- Code repository: VideoX-Fun
Wan2.1 Fun InP Workflow
1. Download Workflow
Download the image below and drag it into ComfyUI to load the workflow:
Run on Comfy Cloud
Open in Comfy Cloud
Download Workflow
Download JSON or search “Wan 2.1 Inpainting” in Template Library
Start Image
Starting frame for the video generation. Download and use this image, or replace with your own.
End Image
Ending frame for the video generation. Download and use this image, or replace with your own.
2. Manual Model Installation
If automatic model downloading is ineffective, please download the models manually and save them to the corresponding folders. All models involved in this guide can be found at Wan_2.1_ComfyUI_repackaged and Wan2.1-Fun. Diffusion Models : Choose 1.3B or 14B. The 14B version has a larger file size (32GB) and higher VRAM requirements:Diffusion Model: Wan2.1 Fun InP 1.3B
models/diffusion_models/wan2.1_fun_inp_1.3B_bf16.safetensorsDiffusion Model: Wan2.1 Fun InP 14B
models/diffusion_models/Wan2.1-Fun-14B-InP.safetensors (rename after download)Text Encoder: umt5_xxl_fp16
models/text_encoders/umt5_xxl_fp16.safetensorsText Encoder: umt5_xxl_fp8
models/text_encoders/umt5_xxl_fp8_e4m3fn_scaled.safetensorsVAE: wan_2.1_vae
models/vae/wan_2.1_vae.safetensorsCLIP Vision: clip_vision_h
models/clip_vision/clip_vision_h.safetensors3. Complete the Workflow Step by Step

- Ensure the
Load Diffusion Modelnode has loadedwan2.1_fun_inp_1.3B_bf16.safetensors - Ensure the
Load CLIPnode has loadedumt5_xxl_fp8_e4m3fn_scaled.safetensors - Ensure the
Load VAEnode has loadedwan_2.1_vae.safetensors - Ensure the
Load CLIP Visionnode has loadedclip_vision_h.safetensors - Upload the starting frame to the
Load Imagenode (renamed toStart_image) - Upload the ending frame to the second
Load Imagenode - (Optional) Modify the prompt (both English and Chinese are supported)
- (Optional) Adjust the video size in
WanFunInpaintToVideo, avoiding overly large dimensions - Click the
Runbutton or use the shortcutCtrl(cmd) + Enterto execute video generation
4. Workflow Notes
- When using Wan Fun InP, you may need to frequently modify prompts to ensure the accuracy of the corresponding scene transitions.