Skip to main content
The WanSoundImageToVideoExtend node extends an existing video latent by generating additional frames, optionally guided by audio, a reference image, and a control video. It takes a starting video latent and produces a longer video sequence, using the provided conditioning and audio cues to influence the new content.

Inputs

Note: The output latent is initialized as zeros with the target dimensions. The input video_latent is not copied into this output; its last 19 frames are used as the reference motion instead.

Outputs

This documentation was AI-generated. If you find any errors or have suggestions for improvement, please feel free to contribute! Edit on GitHub

Source fingerprint (SHA-256): 32b58aaba566f346a0388ba804fc92e7ad426bf2e9e7039e5fdb0bf6a746e972