- Audio-Driven Lip Sync: Generate natural mouth movements that match the input audio
- Full-Body Motion: Preserves identity, background, and camera movement while adding synchronized body motion
- Dual Mode: Supports both single-character and multi-character scenarios
- ComfyUI Native: Built-in
WanInfiniteTalkToVideonode, no custom nodes required
Learn about Subgraph
This workflow uses Subgraph nodes for modular processing. Check out the Subgraph documentation to learn how to customize and extend the workflow.
InfiniteTalk image-to-video workflow
Run in Comfy Cloud
Open in Comfy Cloud
Download Workflow
Download JSON or search “InfiniteTalk” in Template Library
Model Installation
The following models need to be downloaded and placed in the correct directories: diffusion_models text_encoders model_patches- wan2.1_infiniteTalk_single_fp16.safetensors: For single-character scenarios
- wan2.1_infiniteTalk_multi_fp16.safetensors: For multi-character scenarios
Model Storage Location
Sample Input Files
Download these sample input files and drag them into the corresponding nodes:Download Sample Image
Character reference image
Download Speaker 1 Audio
Audio track for character 1
Download Speaker 2 Audio
Audio track for character 2
Workflow Steps
- Load the input image: Drag the character reference image to the
Load Imagenode. For multi-character scenarios, use the Mask Editor to draw masks for each character. - Load audio tracks: Connect audio files to the
Load Audionodes (one per character). - Load diffusion model: Ensure the
Load Diffusion Modelnode is usingWan2_1-I2V-14B-480p_fp8_e4m3fn_scaled_KJ.safetensors. - Load model patches: Load the appropriate InfiniteTalk patch (
singleormultivariant) via theModelPatchLoadernode. - Configure InfiniteTalk: Adjust the
WanInfiniteTalkToVideonode parameters, such aslength(video length in frames),motion_frame_count(previous frames used as motion context), andaudio_scale. - Generate: Run the workflow. The model will produce a full-body talking video synchronized with the input audio.
Extending Video Length
Each Video Extend group extends the video by approximately 3.24 seconds (81 frames at 25 fps). If your audio is longer, you can:- Box-select the “Video Extend” group
- Press
Ctrl-C(copy), thenCtrl-Shift-V(paste with connections) - Connect the previous
InfiniteTalk (Base Generation)node’sIMAGEoutput to the newWanInfiniteTalkToVideonode’sprevious_framesinput - Connect the
Batch Imagesnode’sIMAGEoutput to the newBatch Imagesnode’s image input