Skip to main content
WanInfiniteTalkToVideo generates video sequences from audio input. It uses a video diffusion model, conditioned on audio features extracted from one or two speakers, to produce a latent representation of a talking head video. The node can generate a new sequence or extend an existing one using previous frames for motion context.

Inputs

Common Inputs

Two Speakers Inputs

The inputs in this section are shown when mode is set to "two_speakers". Parameter Constraints:
  • When mode is set to "two_speakers", audio_encoder_output_2, mask_1, and mask_2 are required for the second speaker setup.
  • If audio_encoder_output_2 is provided, both mask_1 and mask_2 must also be provided.
  • If both mask_1 and mask_2 are provided, audio_encoder_output_2 must also be provided.
  • If previous_frames is provided, it must contain at least as many frames as specified by motion_frame_count.

Outputs

This documentation was AI-generated. If you find any errors or have suggestions for improvement, please feel free to contribute! Edit on GitHub

Source fingerprint (SHA-256): b7359490c1de86d9c82122bc227295b3b7f8a3493f629365ae0f22f9f34d9a66