WritingReplicateReplicatepublished Aug 4, 2026seen 2w

The Ultimate Guide to FLUX 3

Open original ↗

Captured source

source ↗
published Aug 4, 2026seen 2wcaptured 2whttp 200method plain

The Ultimate Guide to FLUX 3 – Replicate blog

Replicate Blog

The Ultimate Guide to FLUX 3

Posted August 4, 2026 by shridharathi

Try FLUX 3 on Replicate Run FLUX 3

Black Forest Labs made history when they released their FLUX.1 image model. They were one of the first independent research labs to pioneer fast, high-quality, open source image generation. Now, they are taking a swing at video, and we are intrigued.

FLUX 3 is the lab’s new multimodal foundation model. Compared to other labs, they are taking a new approach. BFL has developed a model with a consolidated architecture, learning from images, audio, and video to generate multimodal outputs. Taken straight from their release blog:

“Images capture spatial structures and relationships at a specific point in time. Videos restore the dimension of time and reveal temporal dynamics and physical laws. Audio reveals causal relationships between mechanical phenomena and acoustics that vision alone cannot detect. Language links these perceptions to goals, abstractions, and instructions.

Learn from one and you get a good model of that projection. Learn from all of them at once and their mutual constraints tell you more: the sound has to match the impact, the motion has to obey the mass, the future has to follow from the past. The modalities stop being separate and start being evidence about one underlying reality.”

Seemingly, this training approach means that this model has more laws of physics baked in. Audio informs motion and vice versa. Video introduces temporality to images which already encode spatial relationships. FLUX 3 is an attempt at creating a net of weights that better encapsulates reality.

This post walks through some things FLUX 3 is best at, with real examples for each.

Text to video

FLUX 3 doesn’t need much to work with. Feed it a plain sentence or a dense, specific one, and it holds up either way. That includes things that require actually knowing how something works, not just what it looks like.

A desert off-road trophy truck racing across open dunes at full speed, kicking up a huge rooster tail of sand and dust behind it, suspension soaking up jumps as it crests a ridge and goes briefly airborne. Camera tracks alongside at high speed, low to the ground. Hyper-realistic, 8k, harsh midday desert light, the roar of the engine and the crunch of sand under tires.

Cinematic underwater wildlife video of a large octopus slowly crawling across a dark rocky seafloor. Low, close tracking angle at eye level, showing its textured mantle, expressive eyes, and eight arms moving independently across wet stone. Suction cups attach and release naturally as the arms pull the body forward; subtle skin ripples and realistic muscular motion. Deep teal water, drifting particles, soft shafts of filtered light, muted rust-red and brown octopus coloration, highly detailed photorealistic skin texture, shallow depth of field, natural documentary cinematography.

I like that FLUX 3 doesn’t necessarily need your prompts to be dense or overwrought with details or keywords. Sometimes, you can offload the details of motion and aesthetics to the model itself, and you’ll probably end up with a satisfying result.

A young man staring out the window of a train rumbling along in rural Switzerland, beautiful mountains in the background whooshing by. Handheld camera feel, film shot.

A diver descends through crystal-clear turquoise water into a jungle cenote, sunbeams piercing down from the opening far above. Beneath the surface, the walls of a submerged Mayan temple come into view, ancient glyphs carved into the stone, fallen pillars scattered across the floor. Fish dart through the ruins as the diver’s flashlight beam sweeps across the carvings. Hyper-realistic, 8k, National Geographic underwater cinematography, muffled bubbles and the diver’s steady breathing.

You’ll notice that FLUX 3 will default to multiple cuts of scenes unless otherwise specified.

A high-speed motorcycle chase through a rain-soaked underground parking garage at night. The rider leans hard through tight turns, headlights strobing past concrete pillars, tires screeching on wet pavement. Sparks fly as the bike clips a support beam and rights itself. Camera cuts between a low chase angle and a handlebar POV. Hyper-realistic, 8k, neon-lit puddles reflecting the headlights, screaming engine and echoing tires.

A hiker moves through a dense Costa Rican rainforest canopy, shafts of golden light breaking through layers of green leaves. A toucan takes flight overhead, a sloth hangs motionless in the branches above, and mist rises off the forest floor. The camera follows from behind at a steady walking pace, then rises to a wide canopy shot. Hyper-realistic, 8k, lush biodiversity, howler monkey calls echoing in the distance and leaves rustling underfoot.

Image to video

FLUX 3 can use a start frame and an end frame to guide a transformation. A useful test is to keep the same car and scene while changing its condition from abandoned to roadworthy.

Start and end frame. Give it two images and describe the physical change you want to see between them.

Start

End

Single continuous shot of the same rusted junkyard car transforming naturally into the restored red car. Keep the exact same vehicle identity, camera angle, wheelbase, body shape, and background throughout. Rust flakes dissolve into clean red paint, dents slowly pull themselves smooth, the cracked windshield becomes clear, missing trim reappears, flat tires inflate, and weeds pull back from around the tires. Metal panels reform in place with continuous physical motion. The car then starts and drives forward onto the highway. No cuts, no teleporting, no sudden jumps, no change of vehicle identity, realistic mechanical transformation, documentary camera, natural light.

Video continuation

The start_video input is a new parameter. You give FLUX 3 an existing clip and it keeps going from the final frame. It preserves the same momentum, framing logic, and audio. It is surprisingly fluid.

Base clip:

A skateboarder rolls down a sunlit street and approaches a quarter-pipe ramp at the end of a skatepark, picking up speed, wheels clicking over pavement cracks.

Continuation:

The skateboarder launches off the ramp, catching air, grabbing the board mid-flight, then lands cleanly and rolls away, wheels landing hard on the ramp’s transition.

Same idea, a different sport:

Base clip:

A surfer paddles out through rolling...

Excerpt shown — open the source for the full document.

Notability

notability 5.0/10

Tutorial for model FLUX 3, no traction info.