Latest
September 2026 ◆ Black Forest Labs introduces FLUX 3 Action ◆ From image generation to models of motion and action

Person / Visual intelligence

Andreas Blattmann and the Art of Making Images Move

He helped make image generation practical, then built a visual AI company in Freiburg. For Andreas Blattmann, the next question is what happens when pictures begin to move - and machines begin to act.

inXf

The story begins with a problem of scale. An image contains an extravagant number of pixels, and early diffusion models had to do their work across that sprawling surface. Each attempt at a better picture demanded more computation. Andreas Blattmann and his collaborators found a way to make the canvas smaller before drawing on it. They compressed the image into a more manageable representation, generated there, and decoded the result back into something people could see. It was an economical idea with unusually large consequences.

The method was called latent diffusion. Its 2021 research paper named Robin Rombach, Blattmann, Dominik Lorenz, Patrick Esser and Björn Ommer. The same line of work helped lead to Stable Diffusion, released in 2022, and brought image generation within reach of a much broader community. The achievement made the researchers visible, but the more interesting part of Blattmann's story is what he did after a still image became easy to make.

2022Latent diffusion paper at CVPR
2024Black Forest Labs launches FLUX.1
2026FLUX 3 moves into video and action

The useful trick was subtraction

Blattmann's work belongs to a long tradition in engineering: find the part of a difficult task that does not need to be done at full size. In image synthesis, the original task involved repeated calculations over pixels. Latent diffusion shifted those calculations into a compressed space. The final picture still needed detail, but the model could reach it by a more efficient route. Blattmann later said the technique let his team produce powerful models with far fewer resources than competing approaches.

That is a technical explanation with a human consequence. Smaller computing demands made experimentation possible for more researchers and developers. People could download a model, alter it and build tools around it. The open release of Stable Diffusion turned a specialist result into a lively public ecosystem. Its effects appeared in software, design and the everyday vocabulary of the internet. A generation of users learned that a sentence could become an image, and a generation of builders learned to ask what else those models might do.

Blattmann's academic route crossed Heidelberg University and LMU Munich; his research work also took him to NVIDIA and Stability AI. Earlier papers had already shown an interest in motion rather than only isolated frames. One explored how a small "poke" at a still image might become a plausible sequence of movement. Later work on latent video diffusion and Stable Video Diffusion made the step into moving images explicit. In retrospect, the transition from photographs to video looks less like a turn than a continuation of the same curiosity.

“Visual intelligence is so much more than content creation.”Andreas Blattmann, 2026

A company with a Freiburg address

In August 2024, Blattmann, Rombach and Esser founded Black Forest Labs. Their first release was FLUX.1, a family of text-to-image models. It came in three versions whose names suggested different audiences: pro for a commercial service, dev with downloadable weights for noncommercial work, and schnell - German for fast - with an open license and a practical invitation to experiment. It was a product launch, but it also carried forward a research habit: make the machinery visible enough that other people can work with it.

The company chose Freiburg as its headquarters. That decision gives the story a geography. Freiburg is close to the Black Forest and far from the daily traffic of Silicon Valley, though the lab also has a San Francisco presence. Blattmann has described the distance as an advantage for concentration. He likes visiting San Francisco, he said, but its constant activity can make sustained focus difficult. For a lab trying to improve models one difficult release at a time, quiet can be productive equipment.

The choice did not limit the audience. FLUX models found their way into tools and workflows used internationally. The company has named partners and customers including Adobe, Canva, Meta and Microsoft. Its Freiburg address became a distinction, not a boundary. There is a modest joke hidden in the arrangement: a company named for a forest spent its first years persuading computers to invent things no camera could find there.

Andreas Blattmann, left, and Robin Rombach together in Freiburg
Andreas Blattmann, left, with co-founder Robin Rombach in Freiburg. Two colleagues, one lab, and a great deal of unfinished visual research. Photo: Black Forest Labs.

The founders brought different parts of the same research history into the new company. Rombach and Blattmann had collaborated on latent diffusion; Esser was another co-author of that work. A startup can describe itself as a fresh beginning, but Black Forest Labs also carried an established set of working relationships. That matters in a field where progress depends on experiments that fail quietly before a public model appears. The launch announcement cited years of work on image, video and faster sampling methods, making the new name a continuation of old questions.

From prompt to production

The first FLUX release answered the question that dominated 2024: how good an image can a model make from words? The next releases asked how useful those images could be in actual work. FLUX.1 Kontext expanded editing. FLUX.2, released in November 2025, let users guide an image with several references, maintain a person or product's appearance across changes, and handle layouts and typography more reliably. These are less theatrical tasks than asking a model for a fantastical scene. They are also closer to what designers, marketers and software teams need on a Tuesday afternoon.

That progression shaped the company's business as well as its research. Open weights gave independent developers something to inspect and extend. Managed APIs offered teams a way to use the models at scale. The two paths were designed to coexist. A researcher can test an idea locally; a creative platform can integrate a service into a product. Blattmann had seen what happened when latent diffusion entered the wider community. At Black Forest Labs, access remained part of the plan, even as the company served large customers.

2021Latent diffusion research is first posted.
2022Stable Diffusion brings the research to a broad audience.
2023Blattmann co-authors work on latent video diffusion and Stable Video Diffusion.
2024Black Forest Labs launches in Freiburg with FLUX.1.
2025FLUX.2 expands editing and multi-reference control.
2026FLUX 3 extends the lab's work into video and action prediction.

There was public recognition along the way. Blattmann was included in a 2024 German Top 40 Under 40 list. In November 2025, he joined the board of the German AI Association, KI Bundesverband, responsible for technology and innovation. The association describes his aim there as strengthening European development of foundation models and an open AI ecosystem. It is a fitting second role for someone whose work grew from German university research into products used across borders.

The lab's funding growth brought another change of scale. A $31 million seed round accompanied the 2024 launch. A later $300 million Series B valued the company at $3.25 billion. Numbers of that size can make a story feel abstract. The more concrete measure is how the company's research questions widened. The lab moved from composing still scenes to preserving their identity through edits, and then to generating what happens over time.

The frame after the frame

In 2026, Black Forest Labs released FLUX 3 Video, generating video with synchronized audio. The step from image to video is easy to underestimate. A single frame can be convincing while the next contradicts it. A person may change appearance, an object may jump across a room, or a movement may ignore the physics the eye expects. Video asks the model to keep a promise from one instant to the next. Audio asks it to keep another promise at the same time.

The lab's September 2026 release, FLUX 3 Action, took a further step. Its open-weights world action model links video and action prediction. The wording sounds remote from the friendly text prompt that made image generation famous, but the connection is direct. If a system can predict what a scene will look like after an action, it begins to model consequence. That is useful to research on robotics and other systems that must work in a physical world rather than merely describe one.

Blattmann has spoken of visual intelligence as a field much wider than content creation. At HumanX in 2026, he pointed toward physical AI as a personal interest and described content creation as an entrance to the technology. At a Stanford class session, he discussed the possibility of combining image, video, audio and action in models that learn from more than a frozen frame. These are ambitions, not completed claims. What exists now is a visible sequence of research and releases that makes the ambition intelligible.

“It can be a huge asset to not be where everyone else is.”Andreas Blattmann on building from Freiburg

His story remains unusual because the grandest claims in AI often come from the loudest addresses. Blattmann's work began with a quieter proposition: make an expensive process more efficient. That proposition changed who could make images. His company stayed in Freiburg, using the same economy of attention to tackle larger problems. The distance has not prevented global partnerships, conference stages or a place in Europe's AI debate. It has given the work a recognizable home.

A good portrait of Blattmann, then, does not end with an immaculate generated image. It ends with the next frame. The researcher who helped a model learn to paint in a smaller space now works with a team asking how machines can follow a scene through time and anticipate what an action will do. The forest in the company name is real. The world these models are being asked to understand is larger, noisier and harder to compress. That is what makes the next picture interesting.