Inputs
Note: The latent video is compressed compared to the input dimensions: the spatial dimensions (width and height) are divided by 32, and the frame count (length) is divided by 8 and rounded up to the nearest whole number. The step values for width, height, and length help keep these divisions even.
Outputs
This documentation was AI-generated. If you find any errors or have suggestions for improvement, please feel free to contribute! Edit on GitHub
Source fingerprint (SHA-256):
0b1e57baf9730d852b03b6bccbb8a033e2be9b9cd2420a0aa3638c31f6d3cd26