Inputs
Note: When a
start_image is provided, it is automatically resized to match the specified width and height using bilinear interpolation. Only the first length frames of the image batch are used for encoding; any additional frames are ignored. If the image batch has fewer than length frames, only those frames are used. Only the RGB channels of the image are encoded. The encoded latent is then injected into both the positive and negative conditioning to guide the video’s initial appearance, and the clean encoded frames replace the noisy start of the model’s output latents.
Outputs
This documentation was AI-generated. If you find any errors or have suggestions for improvement, please feel free to contribute! Edit on GitHub
Source fingerprint (SHA-256):
7212f0ea912578d3b72dddf1333a20054a881e3f22c2b8abd9645fc21e75a08b