The photograph offered a familiar bargain: look closely enough, and it might tell you who was there. Christoph Lassner wanted more from it. At the Max Planck Institute for Intelligent Systems in Tübingen, he was teaching computers to recognize a whole human body in an ordinary digital image. Finding a few points for elbows and knees was useful, he explained in a 2016 interview, but his question was bigger. Could the machine form something closer to a three-dimensional person?
It sounds like a small shift in vocabulary, from points to a body. In practice it is the difference between marking a map and understanding the landscape. A shoulder moves with a torso; clothing changes the outline; a camera compresses depth into a flat rectangle. The computer has to work backward from the rectangle, supplying structure it cannot directly see. Lassner described a system that could, in effect, imagine a body while reasoning about the picture.
That early question has proved remarkably durable. It has followed him through a doctoral lab, a body-modeling startup, Amazon, Meta Reality Labs, Epic Games and the founding of World Labs. The subject grew from a person to a scene and then to a world one might enter. The habit stayed the same: look at the flat evidence and ask what space must lie behind it.
A teenager building places
Lassner has offered one origin for that habit. Speaking at TEDAI Vienna in 2025, he recalled playing 3D games as a teenager and making mods, levels and characters. A game level is a peculiar kind of artwork. It has to keep making sense after the visitor turns a corner. There is no single privileged camera angle to conceal the untidy parts. A player will, sooner or later, open the door the designer hoped they would ignore.
That childhood work was play; later, the problems acquired the formal language of computer vision and graphics. He studied computer science at the University of Augsburg, then pursued doctoral research in Tübingen, where the Max Planck Institute’s Perceiving Systems group worked on ways to describe bodies in three dimensions. His 2016 interview lists Augsburg as his hometown. It also records a young scientist who called himself “passionately curious” and was interested in what studying perception might reveal about how people learn.
Research papers from this period show the question becoming more demanding. Keep it SMPL, which he co-authored, addressed the estimation of human pose and shape from a single image. Unite the People connected 2D observations with 3D human representations. In 2017, with Gerard Pons-Moll and Peter Gehler, he published a generative model for people in clothing. The clothes matter. A person does not arrive conveniently dressed as a geometry diagram, and fabric is excellent at disguising the shape beneath it.
He was also drawn to the machinery that makes research repeatable. Decision forests, an earlier strand of his software work, earned an honorable mention in the 2016 ACM Multimedia open-source software competition. For someone trying to build convincing virtual people, a clever result on a slide is only part of the job. Another researcher has to be able to run it, inspect it and build on it.
The body, the camera and the renderer
After Tübingen, Lassner worked at Body Labs, a startup acquired by Amazon. On his own site, he describes creating the human pose estimation system used in Amazon Halo’s smartphone-based 3D body-modeling feature. The assignment joined the abstract to the everyday. Instead of a carefully arranged research capture, the input was the sort of camera a person already carried. To make that useful, the model had to cope with the errors, angles and ordinary mess of consumer photographs.
At Meta Reality Labs Research, where he later led a team, the questions spread from capture to interactive rendering. If a computer can reconstruct a person or object, can it show the result quickly enough for someone to move around it? His site points to the team’s work on reconstructing and rendering radiance fields, presented at Meta Connect in 2022. He later led research at Epic Games, another place where the requirement for immediacy is hard to avoid. In an interactive environment, “almost real time” is a polite way of saying the visitor has gone home.
Pulsar, a renderer Lassner wrote and published with Michael Zollhöfer, makes the engineering side of this story unusually vivid. It represents a scene using spheres, then renders and optimizes them efficiently. The paper reported real-time work with millions of spheres, and Lassner says the system became the sphere-based rendering backend for PyTorch3D. This does not make a virtual world out of marbles. The spheres are a computational choice: simple pieces arranged to express complex surfaces, with a way to adjust them from image evidence.
The name has a little theatrical flair. The work itself is specific about costs. Memory, speed and the ability to change the scene while learning all decide whether a method leaves the paper. Lassner’s career repeatedly returns to that unromantic threshold. A virtual person or room becomes interesting when it survives interaction.

“I am deeply curious about how we can build interactive virtual representations for the physical world.”Christoph Lassner, personal website
When the canvas becomes a world
In 2024, Lassner co-founded World Labs with Fei-Fei Li, Justin Johnson and Ben Mildenhall. Their stated project was spatial intelligence: AI systems that can perceive, generate, reason about and interact with three-dimensional worlds. The phrase can seem broad enough to swallow the furniture. The group’s first product, Marble, makes it concrete. It generates persistent, explorable 3D spaces from images, video, text or layouts. A visitor can move through the result rather than merely watch a clip of it.
Lassner’s arrival at that point did not require a sudden conversion. His personal biography had long emphasized interactive virtual representations of the physical world, and his publication list moved across people, hair, objects, video and dynamic scenes. Neural Assets, a paper with Aljaž Božič, Denis Gladkov and Luke Doukakis, explored capturing an object from smartphone video and rendering it inside an interactive game environment. Each project widened the frame. The person in the photograph had become an object that could be revisited from another angle; the object became part of a place.
World Labs also brought together complementary histories. Li’s work helped define modern computer vision. Mildenhall is associated with neural radiance fields, a way of rendering scenes from learned representations. Johnson’s research spans visual understanding and generation. Lassner brought a practiced concern for the junction where learned geometry has to behave like a usable world. Founding a company did not settle all of the research questions. It gave the questions a shared product and a deadline.
The company’s public demonstrations expanded through 2025, and Marble became available in November of that year. Its promise is easy to grasp if you have ever wanted to walk into an image. Its difficulty lies in the details: what happens around the corner, whether the same chair remains where it was, whether changes leave the space coherent. The old game-level problem returns, now with a model doing some of the building.
The audience walks onstage
At TEDAI Vienna, Lassner chose a less technical doorway into the same ideas: a movie character who could talk back. His 2025 talk asked viewers to imagine a story whose world and characters could respond as the audience explored. He called the possibility “Content 3.0,” a next step after fixed media and on-demand distribution. In his telling, artists and storytellers would establish the conditions, while the experience could develop differently for each person.
It is an appealing proposition and an awkward one. Stories gain force from choices their creators make: what to reveal, what to leave unsaid, where to stop. An endlessly responsive world needs those choices too, even if they are expressed as rules, characters and limits rather than a single finished sequence. Lassner’s point was that the medium may allow creators to hand an audience a stage instead of a path. The teenager making levels would recognize the pleasure. The researcher knows how much work is hidden beneath a door that can actually open.
His talk moves between the fanciful and the practical, from chatting with a familiar fictional character to the stubborn requirement that AI understand spaces. The examples are playful because play is a demanding test. A viewer can forgive a painted background for staying still. A participant who is invited to move will immediately discover whether it has depth.

The next corner
In June 2026, Lassner announced that he was stepping away from his operating role at World Labs while continuing to support the company as an advisor. He expressed gratitude to Li, Johnson, Mildenhall and their colleagues, and said he had not yet decided what his next chapter would be. It was a personal update, not a manifesto. The work he described, however, had the familiar vocabulary: systems, models and the foundations for spatial intelligence.
Across his public career, the scale has changed more than the underlying question. First, make a computer infer a body from an image. Then make the inferred thing render quickly and convincingly. Then try to keep a whole environment coherent when someone walks into it. Each answer exposes another corner. That is why the earlier research belongs in the story of World Labs, and why the TED stage does too. The same curiosity can be heard in a laboratory explanation, seen in a renderer built from spheres and felt in the wish to ask a character a question.
There is an old distinction between looking at a picture and entering a room. Lassner has spent much of his career narrowing the distance between them. If that work succeeds, the reward will be more than a new kind of image. It will be the ability to turn, take a step and find that the rest of the place is still there.