The title of Abhinav Gupta's doctoral thesis was both academic and mildly rebellious: Beyond Nouns and Verbs. Computer vision had become rather good at naming things. A chair was a chair. A person was a person. Gupta was interested in the connective tissue that labels leave out - where objects sit, what people do with them, what the scene permits, and what might happen next.
That was 2009, before a robot could become an internet celebrity by folding a shirt badly. Gupta completed the thesis at the University of Maryland after studying computer science and engineering at IIT Kanpur. He then went to Carnegie Mellon's Robotics Institute as a postdoctoral researcher, working with Alyosha Efros and Martial Hebert. By 2011 he was on the faculty. The titles changed. The question did not.
How does a machine acquire common sense about a world that is physical, relational and inconveniently unwilling to arrive with labels attached?
First, let the machine keep looking
In the early version of Gupta's quest, the machine was called NEIL: the Never Ending Image Learner. Introduced in 2013 with Xinlei Chen and Abhinav Shrivastava, NEIL continuously inspected images on the web, trying to discover objects, scenes, attributes and the relationships among them. It might learn that a car can appear on a road, or that a zebra is striped. The point was not merely to fill a database. It was to see whether accumulated visual regularities could become a rough form of common sense.
NEIL belonged to an era when teaching a computer to recognize a scene still felt like explaining a joke to a very diligent customs officer. Yet the project made Gupta's preference visible: scale the experience, loosen the dependence on human annotation, and let structure emerge from the data. His later work on learning visual representations from video used time itself as supervision. Consecutive frames supplied information for free. The world, if watched carefully, wrote some of its own labels.
“Our agents live in the physical world and need the ability to interact in the physical world.”Abhinav Gupta, 2021 award lecture abstract
Looking, however, has limits. A camera can observe a mug. It cannot discover the weight of the mug, the slipperiness of its handle, or the small indignity of missing it by half an inch. For that, the machine needs a body.
The curriculum of 50,000 mistakes
In a 2016 project with Lerrel Pinto, Gupta's lab turned robotic grasping into a long, unsupervised apprenticeship. The system made 50,000 grasp attempts over 700 robot-hours. Each success and failure became another lesson. No exhausted annotator had to describe every angle of every object. The robot generated its own training data by trying to pick things up.
The numbers were theatrical; the method was stubbornly practical. A failed grasp was not an embarrassment to edit out of a video. It was supervision. The robot learned because the world answered every action with a consequence. This is the appeal of self-supervision in Gupta's work: it replaces a scarce human instruction with a signal already present in experience.
Recognition followed. In 2016 he received a Sloan Research Fellowship and the IEEE Pattern Analysis and Machine Intelligence Young Researcher Award. The Office of Naval Research named him a Young Investigator in 2018. The International Association for Pattern Recognition awarded him its 2020 J.K. Aggarwal Prize for contributions to unsupervised and self-supervised learning in computer vision and robotics. The honors describe several fields. His work kept dissolving the borders among them.
One question, expanding contact with the world
A lab, a co-founder, and a larger wager
Gupta joined Facebook AI Research in 2018 and helped build its robotics effort while retaining his Carnegie Mellon role. The work broadened from perception to embodied learning: agents that explored, adapted and learned from interaction. One of his close collaborators was Deepak Pathak, another Carnegie Mellon professor whose research treated curiosity as a useful machine instinct.
Gupta returned to CMU full-time in May 2022. In 2023, he and Pathak founded Skild AI. Pathak became chief executive; Gupta became president. Their company would not build its fortunes around one gleaming robot body. It would try to build the intelligence layer that could inhabit many of them.
Skild calls the idea “omni-bodied intelligence.” The model is meant to control quadrupeds, tabletop arms, mobile manipulators, humanoids and machines not yet fashionable enough to have a name. Instead of writing bespoke software for every task and body, the company pools experience across forms. A shared model can learn something more general about motion, objects and consequences.
The commercial pitch and the scientific thesis are unusually well matched. Robot data is scarce. A single task produces too little of it, as does a single hardware platform. Skild therefore gathers experience from several sources: human videos for broad demonstrations, simulation for cheap and varied practice, teleoperation for precise examples, and deployments for the messy truth.
“If videos were sufficient, all of us can watch Roger Federer videos and become Roger Federer.”Abhinav Gupta on why robots must practice
It is a good line because it rescues an important distinction from jargon. Video can show a forehand. It cannot put the racquet in your hand. Simulation lets a machine rehearse under odd conditions - wind, slippery surfaces, shifted loads - without breaking expensive equipment every afternoon. The physical deployment then reveals all the ways the rehearsal was insufficient.
The demo ends; the shift begins
Gupta has said that one reason he moved into industry was fatigue with robot demonstrations presented as breakthroughs. A video can hide resets, slow speeds and the patient graduate student just outside the frame. A factory cannot. It wants the system to recover when an object moves, a motor weakens or Tuesday is unlike Monday for reasons no benchmark anticipated.
This is why Skild treats deployment as part of the technology, not a ceremonial step after invention. A deployed robot generates the data needed to improve the next deployment. By September 2026, the company said it had more than 60 paying customers across warehouses, factories, data centers, kitchens, delivery and inspection, and had crossed $100 million in annual recurring revenue ten months after its first commercial deployment. Those are company-reported figures, but they make the priority unmistakable: the robot must work where invoices are issued.
Factories and warehouses first. Public and service spaces next. Homes later. Gupta's forecast follows the amount of disorder a robot must survive, not the amount of attention a demo can attract.
The company has also attracted capital at a scale that tends to make a careful person blink twice. It emerged from stealth in July 2024 with a $300 million Series A and a $1.5 billion valuation. In January 2026 it announced a $1.4 billion Series C, led by SoftBank, at a valuation above $14 billion. Its partners and investors have included NVIDIA, ABB Robotics, Universal Robots, Jeff Bezos's investment vehicle, and Carnegie Mellon.
Money does not solve physics. It purchases more encounters with it.
Learning in context, recovering in public
In August 2026, Skild introduced S1, a model designed to perform a new task after seeing a single video demonstration, without updating its weights. The ambition resembles in-context learning in language models, except the next token may involve a frying pan. Gupta has described systems composing familiar manipulation skills into longer, unfamiliar jobs and recovering when an attempt goes wrong.
One incident stays with him. A quadruped robot suffered failed ankle motors yet adapted its gait in real time. The machine had not been trained for that exact failure. To Gupta, its recovery illustrated the value of pooling knowledge across varied bodies and conditions. Resilience had appeared not as a separately programmed trick but as a consequence of breadth.
Skild followed S1 with work on physical self-play, using competition in simulation to post-train robots for dynamic tasks. It is a return to an old learning idea inside a newer base model: let an agent create increasingly difficult experience for itself. The motif is familiar from Gupta's career. First, reduce the labels. Then increase the encounters.
Beyond the clever video
Gupta does not expect a single morning when everyone awakens to find the robot age parked politely outside. He has described the “GPT moment” for robotics as gradual. Factories and warehouses offer semi-structured environments: less predictable than traditional automation prefers, but more bounded than a family kitchen at breakfast. Public spaces may follow. Homes, with their stairs, pets, loose socks and emotionally significant mugs, remain the final examination.
That caution gives his large ambition a useful grain. “Robots are ready for today,” he has said. “They're just not ready for homes tomorrow.” The sentence makes room for both progress and reality. It also sounds like the researcher who once asked a machine to go beyond nouns and verbs. The useful intelligence was never in the name of the object. It was in the relationship between an action and what the world did next.
A career can look tidier in retrospect than it felt while being lived. Gupta's, though, has a remarkably visible through-line: scenes before labels, experience before annotation, recovery before choreography. At Carnegie Mellon, at FAIR and now at Skild, the machines have been given steadily more of the world to contend with. They began by looking at it. Then they tried to grasp it. Now they are being asked to show up for work.