FLUX 3: Multimodal Flow Models for Visual Intelligence

FLUX 3: Multimodal Flow Models for Visual Intelligence

FLUX 3 integrates images, video, and audio into a single architecture to build a world model capable of perceiving and predicting physical reality. By training on multiple modalities simultaneously, the model learns the mutual constraints between them—such as the relationship between a physical impact and its corresponding sound—rather than treating each medium as an isolated data stream.

Multimodal Capabilities and Generation

FLUX 3 enables the joint generation of video and audio from text prompts or reference inputs. The model is designed to move beyond simple projection and instead capture the temporal dynamics and physical laws inherent in the real world.

Video and Audio Generation

FLUX 3 can produce videos up to 20 seconds in length with native audio. Key capabilities include:

  • Text-to-Video: Generating cinematic or candid footage from text descriptions.
  • Image-to-Video: Animating a starting frame or using images as visual references.
  • Video-to-Video: Transferring central elements (such as a specific character) from a source clip into a new scene.
  • Controlled Transitions: Using keyframes to generate specific transitions between moments.
  • Advanced Features: Support for multilingual dialogue, diverse aspect ratios, and the agentic chaining of clips into longer sequences.

Image Synthesis

The model supports high-resolution image synthesis and editing across various styles. Preliminary mid-training evaluations indicate significant improvements over previous FLUX versions, specifically in handling complex prompts and rendering high-accuracy text in multiple languages.

Physical AI and Action Prediction

FLUX 3 extends its world understanding into the domain of robotics and action prediction. This is achieved through two primary methods:

  1. Native Integration: Scaling the "Self-Flow" approach to integrate action prediction directly into the model.
  2. Specialized Fine-tuning: Using the pretrained video backbone as a dynamics-aware foundation for specialized action models.

In collaboration with mimic robotics, this has resulted in FLUX-mimic, a video-action model optimized for dexterous manipulation in production environments, currently being tested at Audi.

Performance Benchmarks

Early preliminary evaluations of FLUX 3 Video show a strong preference over several competing models in head-to-head comparisons:

  • Runway Gen-4.5: Preferred in 77% of comparisons.
  • Luma Ray 3.2: Preferred in 93% of comparisons.
  • Grok Imagine Video: Preferred in 69% of comparisons.
  • Kling v3 Pro: Preferred in 60% of comparisons.
  • Gemini Omni Flash and Seedance 2.0: Preferred in 52% of comparisons.

Rollout and Access Plan

BFL is implementing a phased rollout involving early access periods for safety testing and feedback:

  • FLUX 3 Video: Video and audio generation/editing via APIs and private weight access.
  • FLUX 3 Action / FLUX-mimic: Action prediction for selected research and commercial partners.
  • FLUX 3 Image: Image synthesis and editing via APIs and private weight access.
  • FLUX 3 Dev: Open-weight access to the multimodal backbone for content creation and action prediction.

Community Perspectives and Critiques

Discussion among technical users highlights a mix of anticipation for the open-weight release and skepticism regarding the "world model" claims.

Technical Skepticism

Some users questioned the evidence provided in the announcement, noting a lack of diverse human examples and the use of jump-cuts in demonstration clips. One critic argued that the concept of a world model is overused, stating:

"Claims 20 seconds of video, shows only jumpcuts. Coming soon!"

Open-Weight Expectations

There is significant interest in the "FLUX 3 Dev" version, with users hoping it maintains the accessibility and performance of previous open-weight releases for hobbyists and those with limited VRAM.

Philosophical and Practical Concerns

Critics have pointed out the gap between multimodal training and true physical interaction, noting that models trained only on audio/video/images lack tactile data, which may explain the "hesitant" movements often seen in AI-driven robotics.

Sources