A camera has a stubborn habit: it only records the view from the place where it stood. Walk three feet to the left, and the photograph can tell you very little about what comes into sight. For most people, this is a mild inconvenience. For Ben Mildenhall, it became a research problem, then a career, and eventually part of the reason to start a company.
In 2020, while finishing a PhD at the University of California, Berkeley, Mildenhall led a paper with Pratul Srinivasan, Matthew Tancik, Jonathan Barron, Ravi Ramamoorthi, and Ren Ng. Its name was NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis. The title is a mouthful. The idea is easy to feel: show a computer several photographs of a place, then ask it to make a convincing picture from a position none of the photographers occupied.
A staircase glimpsed from two angles should remain a staircase from the third. A shiny surface should change as the viewer moves. An object should still belong in the room when the virtual camera glides past it. NeRF offered a way to represent a scene with a neural network so that new views could be rendered. Its public demos made an abstract problem unusually legible. You did not have to know the mathematics to notice the camera moving through a place assembled from still images.
The computer learns a continuous representation, then renders a requested view through it.
A research life, in camera moves
The route to that paper passed through several places where pictures are taken seriously. Mildenhall received a B.S. from Stanford in 2015. The summer before, he worked at Pixar Research. At Berkeley, Ren Ng advised his PhD and the Hertz Foundation supported it. He spent a summer at Google Research and another at Fyusion, a company working with 3D imaging. These entries read like a tidy academic itinerary. Their common theme is less tidy: how to make a computer understand and render the visual world.
His own website is matter of fact about the work. A short highlight describes the original project as “Many images to 3D.” That phrase omits the equations, as every useful short phrase must, but it gets the intellectual move right. Rather than begin with a carefully constructed digital model of a chair or a room, begin with the images people already know how to make. The challenge is to recover enough structure that the image becomes a place.

In a 2023 interview, Mildenhall explained why this route appealed to him. High quality 3D training data is difficult to gather; ordinary 2D images are plentiful. His concise version was: “I’m focused on extracting 3D information from 2D.” The sentence has the plain confidence of someone who knows exactly which hard problem he has chosen. It also points to the constraint that runs through his work: the world is three dimensional, but most of our records of it are flat.
NeRF was not a solo act. Its list of authors matters because the method emerged from a network of researchers in vision, graphics, and computational imaging. Srinivasan and Mildenhall later shared an ACM Doctoral Dissertation Award honorable mention. They also shared a SIGGRAPH Significant New Researcher Award and the ACM Grace Murray Hopper Award. The names on the paper mark a collaboration, not a single flash of inspiration descending into one notebook.
“I’m focused on extracting 3D information from 2D.”Ben Mildenhall, 2023 interview
The scene gets larger
A method that renders one set of scenes invites another question: what happens when the scene is bigger, messier, or captured in less convenient conditions? At Google Research from 2021 to 2023, Mildenhall worked among colleagues who pursued that question from several directions. The names of the follow-on projects sound like family members with increasingly elaborate hobbies: Mip-NeRF, Mip-NeRF 360, Block-NeRF, Zip-NeRF. The work tackled image quality, scale, and the familiar problem of aliasing, the visual shimmer that appears when detail and pixels fail to agree.
Other projects widened the frame further. DreamFusion, co-authored by Mildenhall and colleagues, explored creating 3D content from text with the help of a pretrained 2D image model. It received an outstanding paper award at ICLR in 2023. The line from NeRF to DreamFusion is clear: once a computer can represent a scene, can it also help invent one? The work moved from reconstructing what a camera had seen toward creating an object no camera had yet photographed.
This progression can sound inevitable when read backward, but the engineering questions were persistent. A good rendering must keep its details sharp as the camera changes position. A generated object must remain the same object when viewed from another side. A large environment cannot collapse into a collection of attractive but unrelated frames. Every new capability raises the standard for spatial coherence. The machine must remember where things are.
A company built around the missing view
In 2024, Mildenhall co-founded World Labs with Fei-Fei Li, Justin Johnson, and Christoph Lassner. Li brought a long record in computer vision; Johnson and Lassner brought their own work in machine learning and graphics. Mildenhall’s research history made the company’s central question familiar. World Labs calls its field spatial intelligence: AI that can generate, reconstruct, reason about, and interact with worlds rather than only produce a single image or line of text.
The first product, Marble, lets people create explorable 3D environments from text, images, video, and 3D layouts. It is possible to describe that in one sentence, then spend an afternoon discovering how much is packed into the word explorable. An image can be composed to impress from one angle. A place has to hold together while someone wanders through it, looks back, and asks what is behind the door. The door is where a handsome demo becomes a difficult engineering problem.
Mildenhall has spoken about the practical value of being able to try many alternatives in architecture and engineering at lower cost. The point is less theatrical than conjuring a city from a sentence, and perhaps more useful. A design team can move through options before it commits materials, time, or money. For filmmakers and game makers, the same family of tools could make early environments quicker to draft and revise. In his 2023 interview, Mildenhall anticipated those creative fields as early places where neural rendering would find use, because making 3D assets is already part of their daily work.
“You can try a thousand times, exploring many potential alternatives at a much lower cost.”Ben Mildenhall, on architecture and engineering
Then came Atlas. World Labs introduced the model in September 2026 as a system built to work with text, images, video, and 3D information in a shared spatial context. Its demonstrations include generating a new camera view from an image, reconstructing spaces from a few images, and reframing video. The company says Atlas can use two or three images for a faithful reconstruction in some cases, while accepting many more when a scene needs them. These are company claims about a new model, and the wider test will be how well it behaves in hands beyond its builders.
The conceptual step is striking even without a benchmark. NeRF asked a learned scene to render a view. Atlas is designed to work when parts of a scene have to be reconstructed and other parts have to be imagined. That is a delicate assignment. The back of a chair may be hidden in a photograph; a useful model must decide what to preserve, what to infer, and what to leave uncertain. In a world you can walk around, a convenient mistake becomes a visible one as soon as you turn the corner.
On September 28, 2026, a new corporate chapter appeared. AMD announced an agreement to acquire World Labs in an all-stock transaction valued at approximately $8.2 billion. The deal is expected to close by the end of the year, subject to approvals and other conditions. AMD said the World Labs team would continue its model research after closing, with Li set to become AMD’s executive vice president and chief scientist. The announcement says nothing specific about a new title for Mildenhall. For now, he remains a World Labs co-founder working in the field the company was built to pursue.
What the next camera sees
Mildenhall tends to speak about this work through its uses and constraints. In a 2025 conversation about World Labs, he described a path familiar from other generated media: people first find an output cute, then interesting, then ordinary enough that they stop announcing it was made by AI. He also imagined models in which a person could experience and transform places with much greater freedom. Those aspirations are expansive, yet the work remains tied to a simple test: does the world still make sense when you move?
The awards provide one account of his career. The SIGGRAPH recognition, the ACM Grace Murray Hopper Award, and the paper honors mark substantial contributions to computer graphics. His personal site offers another account, less ceremonial and more revealing in its structure. Project after project sits beside a small visual demonstration. Each asks the viewer to look again, or from somewhere else. The format suits a researcher whose subject is the space between the pictures we have and the view we want next.
It is tempting to treat the new World Labs models as a sharp break from an academic paper published in 2020. They are better understood as a long continuation of its question. What can a computer learn about a place from the evidence a camera gives it? How much more can it do if it learns from many places? And when the picture becomes a world, what sort of tools let a person use it? Mildenhall has spent his career moving the camera one position farther. The remarkable part is how much room there was beyond the original frame.