Profile / Alex Yu●From one photograph to a world you can move through●Berkeley research / Luma AI / image generation●Profile / Alex Yu●From one photograph to a world you can move through●Berkeley research / Luma AI / image generation●

The People Issue / Visual AI

Alex Yu and the Art of Making a Picture Move

He helped make a 3D world from a few photographs, then helped put that idea in people’s hands. After Luma AI, Alex Yu has turned his attention to image generation.

The photograph, for most of us, is where a moment stops. Alex Yu spent his early research years asking if it might also be where a world begins. Give a machine one view of an object or a room. Can it imagine what lies around the corner? Can it let a person move through the scene without first demanding a studio full of cameras, a patient specialist and a long wait?

This is an unusually practical question for a researcher of abstract visual mathematics. It led Yu from the University of California, Berkeley to papers that made 3D scenes easier to reconstruct and faster to display. It led him, too, to Luma AI, the company he cofounded with Amit Jain. Luma began with a phone-based way to capture the world in 3D and moved into making new worlds from text and images. Yu's public site now marks Luma as a previous chapter. He says he is working on image generation in San Francisco.

His path has a useful shape: less input, less waiting, more room to make something. Each change sounds technical. Each changes who gets to try.

The boy who liked a difficult problem

Yu has described an early journey from Hangzhou, China, to Vancouver, Canada. He began programming at nine and won a Vancouver entrepreneurship competition while in high school. The contest matters less as a trophy than as a glimpse of a preference. He liked building an answer people could use. Years later, he was discussing not only algorithms but data quality, user experience and what keeps a person coming back to a product.

At Berkeley, he studied computer science and applied mathematics. As a teaching assistant for the introductory CS 61A course, he described enjoying problem solving, coffee and dark chocolate. It also records an interest in virtual and augmented reality. There is a pleasing lack of executive polish to the list. The person who would help create a video model was, for a while, the student on a Tuesday evening teaching people how to program.

His undergraduate work ranged beyond graphics. In 2019, he joined Berkeley's Breakthrough Listen effort, looking for alien technosignatures. He also worked on human and hand projects at the FHL Vive Center and spent a summer at Google building features for Assistant. Those were separate stops, not a tidy advance plan. They put him near an enticing common problem, though: how a computer understands a world that is more complicated than a line of text.

Alex Yu outdoors near the Golden Gate Bridge
Alex Yu in San Francisco. A picture of someone whose research kept asking what a picture could become.

One image, many angles

With professor Angjoo Kanazawa and collaborators at Berkeley AI Research, Yu worked on neural radiance fields, commonly called NeRFs. A radiance field is a way to describe a scene so that a computer can render it from a viewpoint the original camera never occupied. The effect can feel like walking into a photograph. The early methods, however, often needed many carefully positioned images and substantial computation for each new scene. That made for absorbing demonstrations and a stubbornly inconvenient everyday tool.

Yu's 2021 paper pixelNeRF, with Vickie Ye, Matthew Tancik and Kanazawa, tried a leaner proposition: learn a scene representation from one or a few input images. It did not make missing information magically appear. It gave the model a way to draw on patterns learned across scenes, then synthesize new views from sparse evidence. For a creator with a phone and no scanning rig, the distinction between needing a handful of images and needing a carefully assembled image collection is considerable.

Another paper addressed the wait after reconstruction. PlenOctrees, written with Ruilong Li, Tancik, Hao Li, Ren Ng and Kanazawa, represented a scene so it could be rendered quickly. Their reported benchmark ran images of 800 by 800 pixels at more than 150 frames per second, over 3,000 times faster than the conventional NeRF comparison in that test. Benchmark numbers do not promise every scene will behave the same way. They do show what the team was trying to remove: the pause between a user's movement and the scene responding.

PLENoctrees / reported test150+frames per second at 800 × 800 pixels
PLENoctrees / comparison3,000×faster rendering than conventional NeRF
PLENoxels / benchmarks100×faster optimization than NeRF

Then came Plenoxels, with Sara Fridovich-Keil and other colleagues. It offered a cheeky solution to a field named for its neural networks: reconstruct the radiance field without a neural network at all. The paper reported comparable benchmark quality while optimizing about 100 times faster than NeRF. It was presented orally at CVPR in 2022. The result was a reminder that progress does not always mean piling more machinery onto a method. Sometimes it means discovering which machinery can be removed.

The paper becomes a phone app

Amit Jain, who had worked at Apple on 3D computer vision and related systems, encountered Yu's research and connected with him through mutual acquaintances. Yu has said they quickly found common ground. The timing was unusually direct. Yu graduated in 2021 and said he turned down PhD offers from Stanford and MIT to start a company. His own career page dates his Luma work from 2022 through 2024; accounts of the company's founding place it in 2021.

The partnership joined two views of the same bottleneck. Jain knew the demands of cameras, phones and the product experience. Yu had been working on the representations that could turn ordinary images into navigable scenes. Their early ambition was to make 3D capture accessible to people with a phone. A creator could record an object or place, upload the material and get a view that could be explored. The appeal was clear to people who had never typed “radiance field” in their lives.

One creator, Michael Rubloff, later recalled a message from Jain about the company he and Yu were building while it was still in pre-alpha. Rubloff's account of borrowing a friend's computer for months to run an earlier 3D tool gives the context: the technical barrier was real enough to require a loaned machine and hours of setup. A phone app changed the invitation. Yu later said Luma's early growth came largely from users sharing what they had made, with little deliberate marketing.

“The moat for consumer software is the product itself.”Alex Yu

That sentence lands differently coming from someone with three prominent research papers. Yu did not dismiss models or data; he also called both crucial in AI. He was saying that a clever technique had to survive contact with a person's purpose. If the capture failed, the wait was long, or the result was hard to share, the mathematical achievement would remain a curiosity. Luma's early users supplied a kind of field test no conference paper could.

From keeping the world to inventing one

The work then shifted. Capturing a real object in 3D asks a machine to preserve what exists. Generating a 3D object from words asks it to propose something new. Luma's Genie, introduced in November 2023, took the latter route. Yu's site describes it as a model that could make a 3D asset from a text prompt in about ten seconds, with a refinement step for quality. The prompt replaced the walk around an object; the output still had to be useful as an object, with shape and surfaces a person could inspect.

In June 2024 came Dream Machine, the video model Yu names alongside Genie on his site. It generated moving images from text or a still image. Yu describes an initial system capable of making 120 frames in under 120 seconds, followed by controls for keyframes and loops. Those figures belong to that release and description, not to every later version. The more interesting thread is creative control. A result is easier to make than a result that can be shaped, repeated or brought into a larger piece of work.

01 / RECONSTRUCT

From photographs

pixelNeRF and Luma's early capture work explored how images could become explorable 3D scenes.

02 / GENERATE

From words

Genie made a 3D asset from a text prompt, with a separate step to refine it.

03 / SET IN MOTION

From prompts or images

Dream Machine carried Luma's visual generation work into video.

Yu was frank about the hard parts. He pointed to the quality and quantity of 3D training data as major challenges. A model trained on inconsistent or unconvincing objects would learn those flaws too. He also discussed Luma's mix of NeRF and Gaussian splatting methods, each useful in different circumstances. There is a grounded engineer's habit here: name the constraints before promising the future. The choice of representation, the available data and the person's experience all belong in the same conversation.

The team was also growing beyond its first audience. Luma's head of growth described a user base reaching from professional photography studios to social-media creators and ordinary people recording daily life. The sight of someone capturing a motorcycle, a room or a small object carries more information about a product than a perfectly edited demo reel. It says what people actually want to keep, show and alter.

A new blank canvas

Yu appeared on Forbes' 2025 30 Under 30 AI list. Recognition arrived after the series of research and product moves that had already defined his work. His current description is more interesting than the award line because it changes the tense. Luma AI is now “previously.” He describes himself as an AI researcher and founder based in San Francisco and says he is working on image generation.

The shift from 3D to video to images could look like a circle back to the first photograph. It is better understood through the recurring question: what can a person make from what they have? At Berkeley, the available material might be one image. At Luma, it might be a phone video, a sentence or a still frame. The work changes when the input, the wait and the controls change. Yu's career so far has been a series of attempts to make those constraints less severe.

A photograph still stops a moment. It can also begin an experiment. That is the small, durable idea beneath the fast frame rates and the model names. Someone sees an object, a place or a picture and wonders what else might be there. Yu has spent years reducing the distance between that thought and the chance to find out.