Profile
Patrick Esser✦VQGAN → latent diffusion → FLUX✦Freiburg, Germany✦Visual intelligence

People / Research / Image generation

Patrick Esser and the Art of Making Images Smaller

Before an image could be conjured from a sentence, Patrick Esser helped solve a less glamorous problem: how to make the image small enough for a machine to learn. The route from a Heidelberg lab to Black Forest Labs runs through that deceptively simple idea.

An image arrives on a screen all at once. A computer meets it as a punishingly large arrangement of numbers. For anyone trying to teach a machine to make pictures, that gap between what a person sees and what a computer must process is more than a philosophical curiosity. It is a bill for memory, compute and time. Patrick Esser built much of his research career around a question hidden inside that bill: what can be set aside while keeping the parts of a picture that matter?

The question carried him from the computer vision group around Björn Ommer at Heidelberg University into a string of papers that shaped image generation. He worked with Robin Rombach on VQGAN, then with Rombach, Andreas Blattmann, Dominik Lorenz and Ommer on latent diffusion. The latter became part of the technical foundation for Stable Diffusion. Esser later worked at Runway, where his research moved toward video, and in 2024 he co-founded Black Forest Labs in Freiburg. Its FLUX models have their own architecture and product lives, but the path to them still passes through that early concern with how an image is represented.

A vocabulary for pictures

At Heidelberg, Esser and his collaborators looked for a way to give a generative model something more useful than raw pixels. Their 2020 paper, Taming Transformers for High-Resolution Image Synthesis, combined convolutional networks, good at noticing local visual structure, with transformers, good at modeling how pieces relate across a whole scene. The resulting system, often called VQGAN, learned a compact vocabulary of visual parts. A transformer could then compose those parts into an image.

It was a tidy division of labor. A model need not attend to every pixel as though each carried a separate thought. It could work with learned pieces and their relationships. In the paper, the team reported high-resolution image synthesis, including guided generation at the megapixel scale. In a later interview, Esser described the work as part of a wider search for image representations useful for making pictures, rather than merely identifying their contents.

The pleasure of such a solution is that it sounds obvious only after it works. A newspaper can describe a scene without listing the color of every dot on the page. A computer needs a learned equivalent of that shorthand. Esser's work helped make such shorthand useful for synthesis, not merely for storage.

The small room behind the big picture

By the time the group turned to diffusion models, the problem of cost had become unavoidable. Diffusion begins with noise and learns to reverse a process that destroys an image. The method can produce striking results, but doing the work directly on full-size pixel grids demands substantial computation. The latent diffusion paper proposed performing the costly generative steps in a compressed space learned by an autoencoder, then translating the result back into an image.

The paper's authors were careful about the tradeoff. Compress too hard and visual detail disappears. Compress too little and the computational bill returns. They looked for a useful middle point, with cross-attention layers that allowed guidance from text and other inputs. That combination supported text-to-image generation, editing, inpainting and other tasks while using far less computation than a comparable pixel-space approach.

Esser remembered promising smaller-scale results and then the obvious, nerve-racking next question: could they scale? Stable Diffusion was the answer that reached a public audience in 2022. The names associated with it span a university group, Runway and Stability AI; the research story is collaborative. Esser's own account emphasizes continuity between the compact-model work and the larger release. The breakthrough looked sudden to people encountering a text box that could make pictures. In the lab, it had taken several rounds of deciding what a model should actually see.

“You have to be okay with uncertainty.”Patrick Esser, on doing research

That attitude mattered because there were too many interacting design choices to settle in advance. In 2022, while describing how he built a research team at Runway, Esser argued for people willing to run experiments before they had a complete map of the terrain. It is a practical temperament, less interested in pretending to know the answer than in making the next test possible. The work's technical elegance did not eliminate the mess of finding it.

Patrick Esser in a portrait published by the Computer Vision and Learning Group at Heidelberg
A portrait from the Computer Vision and Learning Group. The picture is small; the ideas associated with that lab traveled considerably farther.

A picture becomes a tool

Esser's account of why the research mattered was unusually direct. In his Runway interview, he spoke of the desire to express ideas visually and of how much practice drawing can require. A model that turns words into pictures changes the first step. It gives someone with a visual idea another way to get started. He was equally clear that the interface had limits. Text could be imprecise. For some editing tasks, direct visual gestures could be more useful than another sentence in a prompt box.

Those reservations give the story a useful edge. A successful text-to-image model does not end the design problem. It exposes new questions about control: which part of the scene should change, what should stay consistent, and how can a person express that intention without learning a private language of prompts? The questions became more pointed as Esser's work moved from still images to motion.

At Runway, he was among the authors of Gen-1 research, which explored generating new video from an existing clip using image or text guidance. The premise joined two constraints that image models could avoid: a scene has to look plausible from frame to frame, and a transformation has to respect the structure of the original video. Runway presented uses from stylization to storyboarding. It was a sign that the research had begun to leave the laboratory paper and inhabit a creative workflow.

The lesson of that period was not that every task needed more automation. Esser had argued that a visual interface might sometimes beat text. Once people could generate a picture, they naturally wanted to revise it, place it in a sequence, or keep a character intact between scenes. The maker's intention became as important as the model's ability to surprise.

2020VQGAN paper submitted
2021Latent diffusion paper submitted
2024Black Forest Labs launched

The company in the forest

Black Forest Labs was founded in 2024 by a team that included Esser, Rombach and Blattmann. The location, Freiburg, gave the company its name and a home outside the usual American technology map. On launch day, it released FLUX.1 in three versions: a commercial model, an open-weight development model and a faster model with a permissive license. The release made the founders' research ambitions legible as a product strategy: let people experiment, while offering a route for more demanding commercial use.

FLUX.1 was built around flow matching and diffusion transformer blocks. Esser had been lead author on a 2024 paper studying rectified flow transformers for high-resolution synthesis, work associated with Stable Diffusion 3. The architecture was a departure from the earlier latent diffusion recipe, but its governing concern felt familiar. Better generation needed a workable way for text and image information to meet, and a model efficient enough to be put to use.

The company's later releases made the control question increasingly explicit. FLUX.1 Kontext, introduced in 2025, combined generation and editing with text and image inputs. A person could ask for a local change while preserving other visual elements, or iterate through a sequence of edits. FLUX.2 extended the idea with multiple reference images, structured prompts and higher-resolution editing. These are company products developed by a broader team; they are also close relatives of the questions Esser was posing publicly years earlier about direct interaction, consistency and what it means to make an image model useful.

In December 2025, Black Forest Labs announced a $300 million Series B. Funding figures can make an easy ending to a founder story, but money is a resource, not a result. The more telling development is that the lab is trying to make visual generation part of ordinary creative and software work. Its stated direction now includes systems that can perceive, generate and reason about visual material. That ambition will be judged in the tools people can actually use, and in the work those tools allow them to make.

What other people make

Esser has offered a more personal measure of success. He said that seeing someone take his research and build something else with it was a good sign that the work mattered. He also described the shortage of time to chase every promising idea as a reason to publish models openly. More hands could test more possibilities. It is a generous view of research, and a demanding one: once other people have the tools, their unexpected uses become part of the evidence.

That view helps connect the apparently different chapters of his career. The Heidelberg papers supplied methods. Stable Diffusion put them before a large public. Runway tested what the methods could do in video and creative tools. Black Forest Labs turned a research group into a company that has to balance openness, performance and a sustainable business. Through it all, Esser has returned to the useful experiment: make something, release it where possible, and notice what happens next.

There is a quiet joke in the underlying technique. To let people imagine bigger, the researchers first made the image smaller. A dense rectangle of pixels became a representation a machine could handle, and then an invitation: describe a scene, change a detail, build another picture. The final image may fill a screen. Its history includes the patient work of deciding which parts of a picture the machine had to carry in the first place.