Skip to main content
Wan2.1 InfiniteTalk is an open-source audio-driven video generation model built on Wan2.1, developed by MeiGen-AI. It enables you to generate full-body talking videos from a single reference image and an audio input. The character’s mouth movements and body motions are automatically synchronized to match the provided audio. Key Features:
  • Audio-Driven Lip Sync: Generate natural mouth movements that match the input audio
  • Full-Body Motion: Preserves identity, background, and camera movement while adding synchronized body motion
  • Dual Mode: Supports both single-character and multi-character scenarios
  • ComfyUI Native: Built-in WanInfiniteTalkToVideo node, no custom nodes required
Related Links:

Learn about Subgraph

This workflow uses Subgraph nodes for modular processing. Check out the Subgraph documentation to learn how to customize and extend the workflow.

InfiniteTalk image-to-video workflow

Run in Comfy Cloud

Open in Comfy Cloud

Download Workflow

Download JSON or search “InfiniteTalk” in Template Library
Wan2.1 InfiniteTalk Workflow Preview
Make sure your ComfyUI is updated.Workflows in this guide can be found in the Workflow Templates. If you can’t find them in the template, your ComfyUI may be outdated.If nodes are missing when loading a workflow, possible reasons:
  1. You are not using the latest ComfyUI version (Nightly version)
  2. Some nodes failed to import at startup

Model Installation

The following models need to be downloaded and placed in the correct directories: diffusion_models text_encoders model_patches audio_encoders vae loras

Model Storage Location

Sample Input Files

Download these sample input files and drag them into the corresponding nodes:

Download Sample Image

Character reference image

Download Speaker 1 Audio

Audio track for character 1

Download Speaker 2 Audio

Audio track for character 2

Workflow Steps

  1. Load the input image: Drag the character reference image to the Load Image node. For multi-character scenarios, use the Mask Editor to draw masks for each character.
  2. Load audio tracks: Connect audio files to the Load Audio nodes (one per character).
  3. Load diffusion model: Ensure the Load Diffusion Model node is using Wan2_1-I2V-14B-480p_fp8_e4m3fn_scaled_KJ.safetensors.
  4. Load model patches: Load the appropriate InfiniteTalk patch (single or multi variant) via the ModelPatchLoader node.
  5. Configure InfiniteTalk: Adjust the WanInfiniteTalkToVideo node parameters, such as length (video length in frames), motion_frame_count (previous frames used as motion context), and audio_scale.
  6. Generate: Run the workflow. The model will produce a full-body talking video synchronized with the input audio.

Extending Video Length

Each Video Extend group extends the video by approximately 3.24 seconds (81 frames at 25 fps). If your audio is longer, you can:
  1. Box-select the “Video Extend” group
  2. Press Ctrl-C (copy), then Ctrl-Shift-V (paste with connections)
  3. Connect the previous InfiniteTalk (Base Generation) node’s IMAGE output to the new WanInfiniteTalkToVideo node’s previous_frames input
  4. Connect the Batch Images node’s IMAGE output to the new Batch Images node’s image input