Skip to main content
FLUX 3 is Black Forest Labs’ multimodal foundation model, announced July 23, 2026 and currently in Early Access. It jointly learns from images, videos, and audio within a unified architecture built on Self-Flow, their approach for aligning multimodal generation and understanding in the same model. Instead of treating each modality in isolation, FLUX 3 learns a shared representation of the world: how objects hold together, how things move, and how events sound. Capabilities and limits may change during the Early Access rollout. For video, FLUX 3 creates highly diverse clips with native audio up to 20 seconds long in a single generation. All outputs come with synchronized audio generation, including ambient sound, speech, and effects. It supports text-to-video and image-to-video generation, with multi-shot output that chains individual clips into longer sequences and flexible aspect ratios.
To use the Partner Nodes, you need to ensure that you are logged in properly and using a permitted network environment. Please refer to the Partner Nodes Overview section of the documentation to understand the specific requirements for using the Partner Nodes.
Make sure your ComfyUI is updated.Workflows in this guide can be found in the Workflow Templates. If you can’t find them in the template, your ComfyUI may be outdated.If nodes are missing when loading a workflow, possible reasons:
  1. You are not using the latest ComfyUI version (Nightly version)
  2. Some nodes failed to import at startup

Key capabilities

  • Native audio: Every video comes with synchronized audio, including ambient sound, speech, and effects
  • Up to 20 seconds: Generates long clips in a single pass
  • Text and image input: Generates video from a text prompt, or continues from a starting frame image
  • Flexible formats: 720p or 1080p resolution, 24fps, aspect ratios from 9:16 to 21:9
  • Multi-shot output: Chains individual clips into longer sequences for extended storytelling

FLUX 3 text-to-video workflow

Generate a full video from a single text prompt. FLUX 3 expands natural language into a complete scene with synchronized audio, motion, and physics, without needing reference images or complex settings.

Run on Comfy Cloud

Open in Comfy Cloud

Download Workflow

Download JSON or search “FLUX 3 Video” in Template Library

FLUX 3 image-to-video workflow

Generate a full video from a text prompt and a starting frame image. FLUX 3 automatically expands your description into a cohesive scene with synchronized audio, motion, and physics.

Run on Comfy Cloud

Open in Comfy Cloud

Download Workflow

Download JSON or search “FLUX 3 Video” in Template Library
FLUX 3 image to video input

Get started

  1. Update ComfyUI to the latest version
  2. Go to Template Library, search for FLUX 3 Video