Skip to main content
The VOIDInpaintConditioning node prepares the conditioning data needed for inpainting with CogVideoX models. It takes a source video and a preprocessed quadmask, encodes them through the VAE, and combines them into a 32-channel conditioning signal (16 channels from the mask + 16 channels from the masked video) that the model uses to fill in the masked areas.

Inputs

Note: Because CogVideoX-Fun-V1.5 uses patch_size_t=2, the encoded latent must have an even temporal dimension. If length would produce an odd latent_t, the node automatically rounds it down to the nearest valid value and logs a warning. Using an odd latent_t corrupts the last frame through circular padding, which can cause visible jitter or disappearing subjects near the end of the decoded video.

Outputs

This documentation was AI-generated. If you find any errors or have suggestions for improvement, please feel free to contribute! Edit on GitHub

Source fingerprint (SHA-256): 885e462c0f17a3e9610146a05ba3b9c879db0112d3961c95a83f63ba2cd511f1