Skip to main content
This node applies cross-modal (audio-video) guidance to an LTXV-AV model. During sampling, it runs one extra forward pass per step with the audio-to-video and video-to-audio cross-attention connections disabled, then pushes the result toward the coupled prediction. This strengthens audio-visual synchronization, such as lip-sync. The reference default for modality_scale is 3.0; setting it to 1.0 disables the extra pass.

Inputs

Guidance is applied only for sampling steps whose sigma values fall within the range defined by start_percent and end_percent. Outside this range, the node returns the denoised result unchanged. A modality_scale of 1.0 also disables the extra forward pass entirely.

Outputs

This documentation was AI-generated. If you find any errors or have suggestions for improvement, please feel free to contribute! Edit on GitHub

Source fingerprint (SHA-256): 038be607c42e626a8a8f5fe336ee466d0847d43835edb71e20ff38f668069cfb