- Text-to-Video (T2V): Generate videos from text prompts
- Image-to-Video (I2V): Generate videos from an input image
- FLF2V: Interpolate between a start and end image
- Image-Audio-to-Video (IA2V): Generate lip-synced videos from an image and audio
- IC-LoRA Union Control: Control video generation with depth, pose, or edge guidance
- ID-LoRA: Generate personalized videos with synchronized audio from a reference image and audio clip
Key features
- Finer details: New latent space and updated VAE for sharper textures, cleaner edges, and more precise visuals
- 9:16 Portrait support: Greatly improved quality for vertical portrait videos
- Better audio: Cleaner sound with reduced noise and enhanced dialogue
- Improved image-to-video: More consistent motion and fewer glitches
- Smarter prompt understanding: Improved text encoder for more accurate interpretation
- Native ComfyUI support: All workflows are built-in, no custom nodes required
Getting started
LTX-2.3 is natively supported in ComfyUI. To get started:- Update ComfyUI to the latest version
- Go to Template Library > Video > choose any LTX-2.3 workflow
- Follow the pop-up to download models and run the workflow
ComfyUI Native Workflows
LTX-2.3 Text to Video (T2V)
Generate videos from text prompts with improved prompt understanding and text rendering.Run in Comfy Cloud
Open in Comfy Cloud
Download Workflow
Download JSON or search “LTX-2.3 T2V” in Template Library
Model downloads
Checkpoint: ltx-2.3-22b-dev-fp8
Place in
ComfyUI/models/checkpoints/LoRA: distilled 1.1
Place in
ComfyUI/models/loras/LoRA: gemma abliterated
Place in
ComfyUI/models/loras/Text Encoder: gemma 12B
Place in
ComfyUI/models/text_encoders/Upscaler: spatial x2
Place in
ComfyUI/models/latent_upscale_models/Model storage
Prompting tips
- Core Actions: Describe events and actions as they occur over time
- Visual Details: Describe all visual details you want to appear in the video
- Audio: Describe sounds and dialogue needed for the scene
LTX-2.3 Image to Video (I2V)
Generate videos from an input image with more consistent motion and smoother animations.Run in Comfy Cloud
Open in Comfy Cloud
Download Workflow
Download JSON or search “LTX-2.3 I2V” in Template Library
Input Image: egyptian_queen.png
Download the default input image, or use your own image.
Model downloads
The I2V workflow uses the same model set as Text-to-Video.Checkpoint: ltx-2.3-22b-dev-fp8
Place in
ComfyUI/models/checkpoints/LoRA: distilled 1.1
Place in
ComfyUI/models/loras/LoRA: gemma abliterated
Place in
ComfyUI/models/loras/Text Encoder: gemma 12B
Place in
ComfyUI/models/text_encoders/Upscaler: spatial x2
Place in
ComfyUI/models/latent_upscale_models/Model storage
LTX-2.3 FLF2V
Interpolate between a start image and an end image to generate a smooth video transition.Run in Comfy Cloud
Open in Comfy Cloud
Download Workflow
Download JSON or search “LTX-2.3 FLF2V” in Template Library
Start Image: high_view_classic_car.png
First frame of the video.
End Image: low_view_classic_car.png
Last frame of the video.
Model downloads
The FLF2V workflow uses the distilled checkpoint instead of the dev checkpoint.Checkpoint: ltx-2.3-22b-distilled-fp8
Place in
ComfyUI/models/checkpoints/Text Encoder: gemma 12B
Place in
ComfyUI/models/text_encoders/Model storage
LTX-2.3 Image Audio to Video (IA2V)
Upload an image and an audio file to generate a high-quality video with synchronized lip movements.Run in Comfy Cloud
Open in Comfy Cloud
Download Workflow
Download JSON or search “LTX-2.3 IA2V” in Template Library
Input Image: cactus_man.png
Reference image for the speaking character.
Reference Audio: ltx_23_audio.mp3
Audio track for lip synchronization.
Model downloads
The IA2V workflow uses the same model set as Text-to-Video.Checkpoint: ltx-2.3-22b-dev-fp8
Place in
ComfyUI/models/checkpoints/LoRA: distilled 1.1
Place in
ComfyUI/models/loras/LoRA: gemma abliterated
Place in
ComfyUI/models/loras/Text Encoder: gemma 12B
Place in
ComfyUI/models/text_encoders/Upscaler: spatial x2
Place in
ComfyUI/models/latent_upscale_models/Model storage
LTX-2.3 IC-LoRA Union Control
Generate LTX-2.3 videos with IC-LoRA using aligned control inputs like depth, pose, or edges. IC-LoRA is an in-context LoRA trained on top of LTX-2.3 that applies structural guidance from a reference video. It accepts control signals from various preprocessors including depth maps, Canny edges, and pose skeletons. This workflow uses a subgraph for modular processing.
Run in Comfy Cloud
Open in Comfy Cloud
Download Workflow
Download JSON or search “LTX-2.3 IC-LoRA” in Template Library
Learn about Subgraph
This workflow uses a Subgraph node for modular processing. Check out the Subgraph documentation to learn how to customize and extend the workflow.
Input materials
Control Video: stone_ruins.mp4
Driving video for control signal extraction.
Reference Image: the_forgotten_gate.png
Reference image for style and content.
Model downloads
Checkpoint: ltx-2.3-22b-distilled-fp8
Place in
ComfyUI/models/checkpoints/LoRA: IC-LoRA Union Control
Place in
ComfyUI/models/loras/LoRA: gemma abliterated
Place in
ComfyUI/models/loras/Text Encoder: gemma 12B
Place in
ComfyUI/models/text_encoders/Geometry: MoGe Normal
Place in
ComfyUI/models/geometry_estimation/Model storage
LTX-2.3 ID-LoRA
Generate personalized videos with synchronized audio from a text prompt, reference image, and short audio clip. Uses ID-LoRA to adapt a person’s appearance and voice in a single generative model.Run in Comfy Cloud
Open in Comfy Cloud
Download Workflow
Download JSON or search “LTX-2.3 ID-LoRA” in Template Library
Reference Image: vintage_thinker.png
Reference image for the character appearance.
Reference Audio: ltx23_reference_audio.mp3
Audio clip for voice cloning and lip sync.
Model downloads
Checkpoint: ltx-2.3-22b-dev-fp8
Place in
ComfyUI/models/checkpoints/LoRA: distilled 1.1
Place in
ComfyUI/models/loras/LoRA: ID-LoRA TalkVid
Place in
ComfyUI/models/loras/Text Encoder: gemma 12B
Place in
ComfyUI/models/text_encoders/Upscaler: spatial x2
Place in
ComfyUI/models/latent_upscale_models/Model storage
Prompting tips
When writing prompts for LTX-2.3, focus on detailed, chronological descriptions of actions and scenes:- Core Actions: Describe events and actions as they occur over time
- Visual Details: Describe all visual details you want to appear in the video
- Audio: Describe sounds and dialogue needed for the scene