Black Forest Labs has unveiled Flux 3, a new multimodal foundation model capable of generating video with native sound for durations up to 20 seconds. This release marks a significant milestone for the company, as it represents their first model to integrate native audio directly into generated video content. The model learns from a diverse range of data, including images, video, and audio, positioning it as a comprehensive tool in the evolving landscape of generative AI. Internal testing by Black Forest Labs indicates Flux 3 performs just ahead of Seedance 2.0, a current market leader, signaling a potential shift in competitive dynamics for video generation.

Key Developments

  • Black Forest Labs launched Flux 3, a multimodal foundation model that generates video with native audio.
  • Flux 3 is the first model from Black Forest Labs to produce video content with integrated sound.
  • The model can generate videos up to 20 seconds in length, a new capability for the company.
  • Internal benchmarks suggest Flux 3 surpasses Seedance 2.0 in performance, though independent validation is pending.
  • Black Forest Labs is already exploring Flux 3’s application in robotics tasks, aligning with its long-term goal of building a world model.

What Happened

Black Forest Labs recently introduced Flux 3, a sophisticated foundation model designed to process and generate content across multiple modalities. This new iteration distinguishes itself by its ability to create video sequences that include native, synchronized audio, a feature previously unavailable in the company’s offerings. The model’s training regimen incorporates a wide array of visual and auditory data, enabling it to synthesize coherent video and sound experiences.

Initial assessments conducted by Black Forest Labs suggest Flux 3 demonstrates superior performance compared to Seedance 2.0, a recognized leader in the generative video space. While these findings await independent verification, they highlight Flux 3’s competitive potential. The company’s ambition extends beyond media generation, as it is actively deploying Flux 3 in preliminary robotics applications, indicating a strategic move towards developing a comprehensive “world model.”

Analysis

The introduction of Flux 3 with native audio generation capabilities represents a notable technical advancement in the field of generative AI. Integrating synchronized audio directly into video output addresses a significant challenge for creators, who often rely on separate tools or post-production processes to add sound to AI-generated visuals. This unified approach streamlines content creation and enhances the realism and immersion of synthetic media.

Black Forest Labs’ internal performance claims, placing Flux 3 ahead of Seedance 2.0, position the company as a serious contender in the increasingly competitive market for AI video generation. While external validation is essential to confirm these benchmarks, the assertion itself signals a rising bar for quality and capability. The strategic decision to test Flux 3 on robotics tasks also underscores a broader vision, suggesting that the underlying multimodal architecture is being developed with general intelligence and real-world interaction in mind, moving beyond purely creative applications.