FLUX 3 Video
| Field | Value |
|---|---|
| Developer | Black Forest Labs |
| Model family | FLUX 3 (multimodal foundation model) |
| Type | Video generation with native audio |
| Announced | July 23, 2026 (FLUX 3 early access) |
| General availability | August 4, 2026 |
| Modes | Text-to-video, image-to-video (keyframes), video continuation |
| Clip length | 5 to 20 seconds (continuation 5 to 15 seconds) |
| Resolution | HD (720p class) or Full HD via upscaler, 24 fps |
| Audio | Native dialogue, sound effects, and ambience, generated with the frames |
| Access | BFL API and dashboard, pay per second of output |
| Website | https://bfl.ai/models/flux-3 |
FLUX 3 Video is a video generation model by Black Forest Labs (BFL), released into general availability on August 4, 2026. It is the video-generation product of FLUX 3, a multimodal foundation model announced on July 23, 2026 that is trained jointly on images, video, and audio, and it is the first video generator BFL has shipped: the company's earlier FLUX.1 and FLUX.2 families generate images only.[1][2][3][6] The model produces clips of up to 20 seconds with natively generated audio, supports text-to-video, image-to-video, and video continuation, and is offered through BFL's API on per-second pricing rather than as open weights.[3][5] In the first week after general availability it placed second overall on Design Arena's crowdsourced video leaderboard, behind Google's Gemini Omni Flash and ahead of MiniMax H3.[10][12]
Background and release
Black Forest Labs, the Freiburg-based company founded by Robin Rombach and other authors of the latent diffusion paper, built its first two model generations on text-to-image and image editing. FLUX 3, announced on July 23, 2026, is the company's move into multimodality. BFL describes it as one model that "jointly learns from images, videos, and audio within a unified architecture," on the argument that each modality is a partial projection of the same physical reality and that training on them together teaches the model constraints no single modality contains: the sound has to match the impact, the motion has to obey the mass.[1][2]
The July 23 announcement opened a gated early access phase and set out a staged rollout under four product names: FLUX 3 Video (video and audio generation), FLUX 3 Image (image synthesis and editing), FLUX 3 Action (action prediction for robotics, rolling out through partners beginning with mimic robotics), and FLUX 3 Dev, a planned open-weight release of the multimodal backbone.[1][2][6] On August 4, 2026, BFL made an initial version of FLUX 3 Video generally available through its API and selected partners, the first of the four to exit early access.[3][7]
The underlying model builds on Self-Flow, BFL's method for aligning multimodal generation and understanding in one architecture, which the company scaled up with more compute and data for FLUX 3.[1][9] According to BFL, video prediction accounts for over 95 percent of FLUX 3's training compute, while audio makes up less than 0.5 percent of the tokens in a 720p video; the company argues that a model forced to render motion, contact, and cause and effect accurately ends up learning a usable model of how the world behaves.[8] BFL says the model was trained on tens of millions of hours of general video plus hundreds of thousands of hours of footage focused on human and robot manipulation.[8]
That framing connects the video product to BFL's physical AI work. Alongside the FLUX 3 announcement, BFL and mimic robotics introduced FLUX-mimic, a video-action model that decodes robot actions from the FLUX 3 backbone's internal representations; it has been tested and deployed with Audi on production tasks such as kitting parts and handling cables and seals, and BFL claims the backbone runs in under 80 milliseconds on a single RTX 5090.[2][8] The press release claims FLUX-mimic can be fine-tuned for a new manipulation task with as little as 30 minutes of robot data where earlier approaches needed 30 or more hours.[2]
Capabilities
FLUX 3 Video exposes three request modes through a single API endpoint, all with optional native audio generation:[3][5]
| Mode | Input | Output length | Notes |
|---|---|---|---|
| Text-to-video (t2v) | A prompt | 5 to 20 s | Scene logic, motion, and audio from the prompt alone |
| Image-to-video (i2v) | A prompt plus 1 to 10 images | 5 to 20 s | Images pin the start frame, end frame, or timed keyframes |
| Video continuation (v2v) | A prompt plus an existing clip | 5 to 15 s | Continues motion, camera behavior, dialogue, and audio across the seam |
The keyframe system doubles as storyboarding: one image sets a start frame, two pin the start and end, and up to ten can be placed at exact timestamps using second-and-image pairs.[5] Video continuation accepts up to four seconds of existing video and audio as context.[3] Other documented capabilities include multiple scenes and camera angles within a single generation, typography rendered as part of the scene, dialogue with lip-syncing in languages BFL lists as including English, Chinese, Spanish, French, German, Japanese, Portuguese, Russian, Italian, Indonesian, Turkish, Hindi, and Punjabi, and agentic chaining of clips into longer multi-shot sequences.[1][3][4] BFL positions the output range as deliberately wider than a cinematic default, extending to camcorder-style, animated, and stylized looks.[3][4]
A draft mode returns a fast preview at roughly a third of the cost of a full render; a chosen draft can then be re-rendered at full quality through a "draft_enhance" request that reproduces the same subjects, composition, and motion rather than reinterpreting the prompt.[3][5]
Specifications and pricing
Per BFL's API documentation as of August 8, 2026:[5]
| Item | Detail |
|---|---|
| Output | 24 fps, HD (up to 1 megapixel per frame) or FHD (up to 2 megapixels, finished by a video upsampler; 1920 x 1088 for 16:9) |
| Aspect ratios | 21:9, 2:1, 16:9, 4:3, 1:1, 3:4, 9:16 |
| Text/image-to-video price | $0.17 per second (HD), $0.29 per second (FHD), $0.06 per second draft |
| Video continuation price | $0.43 per second (HD), $0.54 per second (FHD), $0.12 per second draft |
| Audio | Included at no extra charge, on by default |
| Delivery | Asynchronous submit-and-poll API; result URLs expire about two hours after completion |
BFL's own pages disagree slightly on continuation pricing: the model page FAQ and press coverage from early August list video-to-video at $0.41 per second HD and $0.53 FHD, while the API documentation's specification table lists $0.43 and $0.54.[4][5][7] The documentation labels FLUX 3 a preview model, with video editing and reference-based generation from images and videos named as upcoming additions.[5]
Leaderboard performance
BFL's own evaluations are company claims. At the July 23 announcement, using 10-second 720p clips, BFL reported that human raters preferred FLUX 3 over Luma Ray 3.2 in 93 percent of comparisons, Runway Gen-4.5 in 77 percent, Grok Imagine Video in up to 69 percent, Kling v3 Pro in 60 percent, and Seedance 2.0 and Gemini Omni Flash in 52 percent each.[1][6] At general availability BFL said the released model beats existing state-of-the-art models in text-to-video "by a solid margin" and ties Seedance 2.0 in image-to-video.[3][7]
Independent crowdsourced data comes so far mainly from Design Arena, whose leaderboards rank models by Elo from blind pairwise human votes. On its main Video leaderboard, as of the August 7, 2026 update (accessed August 8), FLUX 3 Video ranked second of 39 listed models with an Elo of 1,326 from 5,900 battles, behind Google's Gemini Omni Flash (1,385) and just ahead of MiniMax H3 (1,318), on a board totaling about 391,000 votes; Design Arena had announced the model's second-place debut, then at an Elo of 1,325, on August 6.[10][12] On the same site's Image to Video board it ranked third at 1,301, behind MiniMax H3 (1,352) and Grok Imagine Video 1.5 Preview (1,324), and level on Elo with Seedance 2.0.[11] The board's metadata also records an average generation time of about 220 seconds for FLUX 3 Video, against roughly 51 seconds for Gemini Omni Flash and 34 seconds for Grok Imagine Video.[10]
The other two widely cited video leaderboards had not listed the model as of August 8, 2026: FLUX 3 Video appeared on neither LMArena's text-to-video and image-to-video boards nor Artificial Analysis's video leaderboards on that date.[13][14][15][16] These are three separate ranking systems with different voter pools, and their numbers are not comparable across sites.
Availability, partners, and safety
FLUX 3 Video is sold through the Black Forest Labs dashboard and API on pay-as-you-go terms, with no open weights for the released model; open access is planned only for the separate FLUX 3 Dev backbone.[1][4][5] The July 23 press release named Canva, Burda, Magnific (formerly Freepik), Krea, and Picsart as companies already testing FLUX 3.[2] BFL's model page carries deployment statements from Picsart, Burda, Magnific, Envato, and Nous Research, whose Hermes Agent chains FLUX 3 clips into longer pieces by generating a shot, checking it, and continuing from it.[4]
Before release BFL had the model evaluated by the third-party trust and safety firm Cinder across its supported modalities, including for non-consensual intimate imagery and child sexual abuse material, alongside the company's own mitigations.[3]
See also
- Black Forest Labs
- FLUX.2
- MiniMax H3
- Seedance 2
- Gemini Omni Flash
- Design Arena
- AI video generation
- mimic robotics
References
- ^FLUX 3 - Real World Models: Towards Multimodal Flow Models as the Backbone of Visual Intelligence - Black Forest Labs, July 23, 2026
- ^Black Forest Labs Unveils FLUX 3, A New Multimodal Frontier Model For Visual Intelligence - GlobeNewswire (Black Forest Labs press release), July 23, 2026
- ^FLUX 3 Video, Part 1: Generation - Black Forest Labs, August 4, 2026
- ^FLUX 3: One Multimodal Model - Black Forest Labs, accessed August 8, 2026
- ^FLUX 3 Overview - Black Forest Labs documentation, accessed August 8, 2026
- ^Flux 3 generates videos with native audio up to 20 seconds long, a first for Black Forest Labs - The Decoder, July 23, 2026
- ^Black Forest Labs makes FLUX 3 Video generally available and claims it beats Seedance 2.0 - The Decoder, August 5, 2026
- ^FLUX 3 x mimic: The Next Generation of Video-Action Models - Black Forest Labs, July 23, 2026
- ^Black Forest Labs Releases FLUX 3: A Multimodal Flow Model for Image, Video, Audio and Robot Action Prediction - MarkTechPost, July 26, 2026
- ^Video Leaderboard - Design Arena, accessed August 8, 2026
- ^Image to Video Leaderboard - Design Arena, accessed August 8, 2026
- ^FLUX 3 Video by @bfl_ai is 2nd overall on Video Arena with an Elo of 1325 - Design Arena on X, August 6, 2026
- ^Text to Video Leaderboard - LMArena, accessed August 8, 2026
- ^Image to Video Leaderboard - LMArena, accessed August 8, 2026
- ^Text to Video Leaderboard - Artificial Analysis, accessed August 8, 2026
- ^Image to Video Leaderboard - Artificial Analysis, accessed August 8, 2026
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
v1 · 1,842 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Reviewer note: New article verified against Black Forest Labs posts and docs, the GlobeNewswire release, and live leaderboard APIs on August 8, 2026, including the documented pricing discrepancy between BFL's own pages.
Cite this page: AI Wiki. "FLUX 3 Video." aiwiki.ai, updated 7 Aug 2026, fact-checked 7 Aug 2026. CC BY 4.0. https://aiwiki.ai/wiki/flux_3_video