Skip to main content
The TextGenerateLTX2Prompt node expands a short user prompt into a detailed, audio-visual description suitable for generating video with the LTX-2 series of video models. It automatically adds task-specific system instructions, sends the formatted prompt to a language model, and returns the enhanced text. When an optional reference image is supplied, the node switches to image-to-video mode and expands the prompt starting from that image’s content.

Inputs

Note: The behavior of the node changes based on its inputs:
  • If an image is provided, the generated prompt is formatted for an image-to-video task using a system prompt that describes how to expand the prompt based on the image’s content. If no image is provided, the formatting is for a text-to-video task using a system prompt that expands the prompt into a detailed video generation description.
  • If the CLIP tokenizer’s name contains “gemma4”, the node uses the LTX-2.4 system prompts and the Gemma 4 chat format. Otherwise, it uses the LTX-2 (Gemma 3) system prompts and chat format.
  • If the language model produces no usable text after removing reasoning blocks, the node returns the original prompt instead.

Outputs

This documentation was AI-generated. If you find any errors or have suggestions for improvement, please feel free to contribute! Edit on GitHub

Source fingerprint (SHA-256): 8f524ea60a247217dde8a1edaf7a689e253ae05acc9eb52ad47b91e879dba1df