Skip to main content
Generate a video with synchronized dialogue and sound from a text prompt using HeyGen Video 1.0, optionally guided by connected reference material. Images of people, products or places, videos to reuse and audio clips that provide a voice can all be used as references. Mention each reference in the prompt as @Image1, @Video1 or @Audio1, numbered per type in the order the inputs are connected: these tags are rewritten into the labels HeyGen expects, and a tag that points to a reference which is not connected raises an error. Without any reference image or video the node runs as text-to-video generation.

Inputs

Parameter Constraints

  • Reference limit: at most 12 references in total across reference_images, reference_videos and reference_audios; connecting more raises an error.
  • Audio needs image or video: reference audio is rejected when no reference image and no reference video is connected.
  • Reference image rules: each reference_images input must hold exactly one image (a batch is rejected), and each image must be between 1:4 and 4:1 in aspect ratio.
  • Prompt tags: @ImageN, @VideoN and @AudioN are matched case-insensitively. The number must not exceed the count of connected references of that type, and the prompt must be non-empty after trimming whitespace.
  • Mode: with at least one reference image or video the request is a reference-to-video run; with none it is a plain text-to-video run.
  • Seed: the seed only decides whether the node re-runs; results are not reproducible with the same seed.

Outputs

This documentation was AI-generated. If you find any errors or have suggestions for improvement, please feel free to contribute! Edit on GitHub

Source fingerprint (SHA-256): 44de381703821043aa1399e58c9132f9b6a2b0ac5dbf47197dffed0438ea7ad7