Skip to main content
Animate a still portrait into a talking video driven by speech audio, using sync.so’s sync-3 model. The output duration matches the audio duration, and cost scales with output duration.

Inputs

Note: The speaker_x and speaker_y parameters are only used when speaker_selection is set to "coordinates". When auto_downscale is enabled, speaker coordinates are automatically scaled to match the downscaled image dimensions.

Outputs

This documentation was AI-generated. If you find any errors or have suggestions for improvement, please feel free to contribute! Edit on GitHub

Source fingerprint (SHA-256): 21f722cdcc5ff017949887ed2252854feebb9b913034dc6a6c3ce196ad089468