The robot had wheels, two arms, a head crowded with gizmos, and a deceptively domestic assignment: tie a knot. John Schulman was a graduate student at Berkeley, newly diverted from neuroscience by Pieter Abbeel's lab and its mechanical menagerie. The knot was partly a rehearsal for surgical suturing. The machine was a PR2, a personal robot whose movements had to be arranged with laborious care. Years later Schulman would call that same machine an antique. The more important relic was the method: engineering one polished performance at a time.
Schulman could see the trap. A robot that mastered a single demonstration through layers of custom programming was impressive only until the room changed. It did not possess a general knack. It possessed choreography. Deep learning was beginning to make vision systems sturdier by letting them learn from large amounts of data. Schulman wondered whether robotics might acquire the same resilience. His answer was reinforcement learning, the old idea of learning through trial, error, and reward, newly paired with neural networks.
It is tempting to tell his career backward from ChatGPT, as if every robot arm had been pointing toward a chat box. The actual route was more charmingly untidy. As a boy he loved Isaac Asimov and became absorbed in BattleBots. He and his friends planned a combat robot, and Schulman read whatever might help them build it. The robot never materialized. The independent study did. In 2005 he made the United States Physics Olympiad team. At Caltech he studied physics because, as he later put it, he was curious about understanding the universe.
A useful change of mind
Physics did not hold him. Summer research projects felt flat, so he arrived at Berkeley in 2010 to study neuroscience. Then came a rotation in Abbeel's lab. The professor's work on autonomous helicopters and towel-folding robots was concrete, faintly comic, and irresistible. Schulman noticed that he was spending all his time on it. He switched into electrical engineering and computer sciences.
Berkeley had made a softer pitch, too. On the morning after he arrived, Schulman ran up the road toward the Berkeley Lab. It was around half past seven; almost nobody was out. He encountered a small herd of deer, with fawns among them. He remembered it as a great moment. For a researcher associated with dense equations and industrial laboratories, it is a revealing origin scene: a hunch about a place, confirmed by quiet attention.
His doctoral work gave the new direction mathematical footing. Trust Region Policy Optimization, or TRPO, sought to improve a policy without allowing a destabilizing leap. Generalized advantage estimation reduced the noise in judging which actions deserved credit. Then came Proximal Policy Optimization, or PPO, in 2017. PPO kept the central instinct while making it simpler to use: learn aggressively enough to improve, cautiously enough not to wreck the policy.
The jargon can obscure a plain concern. Learning systems change. An update that looks helpful can push them too far and spoil behavior that already worked. PPO placed a practical guardrail around that change. Researchers adopted it in games, robotics, and eventually language-model training. In 2018, MIT Technology Review included Schulman in its Innovators Under 35 list. The distinction recognized an algorithmic contribution, but the contribution's character mattered too. PPO was designed to work in laboratories, not merely look tidy on a blackboard.
The laboratory where AGI was acceptable conversation
In 2015, while Schulman was finishing his doctorate, Sam Altman drove to Berkeley and took a walk with him. Altman and Elon Musk were assembling an unusual research lab. Artificial general intelligence, meaning a system able to match or exceed people across broad domains, was still an impolite level of ambition in many technical circles. Schulman found the audacity attractive. OpenAI was prepared to discuss the subject seriously, and he wanted to do AI research in a place where that conversation was allowed.
The early group explored several directions. Schulman's reinforcement-learning experience fitted games and simulated agents, but language models gradually offered a more consequential arena. A 2017 OpenAI project led by Paul Christiano showed how human preferences could train agents in Atari games and simulated robotics. Later work applied related methods to language-model summaries. When GPT-3 arrived, Schulman joined that line of research.
The RLHF loop, without the fog
The resulting process became known as reinforcement learning from human feedback, or RLHF. People compare model outputs. A reward model learns from those comparisons. The language model is then trained to produce answers likely to receive higher approval. It is an imperfect proxy for judgment, but it changed the experience of talking to a model. Raw completion became response. The system learned to follow an instruction, decline some requests, admit uncertainty more often, and adopt a recognizable conversational posture.
WebGPT trained a model to browse and answer questions with references. InstructGPT showed that a smaller model fine-tuned with human feedback could be preferred to a much larger untuned one. Schulman led the reinforcement-learning team that produced ChatGPT, and from 2022 through 2024 he co-led OpenAI's post-training team. Post-training is where a general model acquires much of its useful behavior after the vast initial training run. It is also where research stops being abstract. Every preference implies a choice about how a machine should act around people.
“It seemed crazy to talk about AGI at the time, but I thought it was reasonable to start thinking about it.”John Schulman, on joining OpenAI
The surprise waiting outside the lab
By the time ChatGPT was ready, OpenAI had already trained GPT-4. The more capable model occupied the staff's attention, while the chat product used GPT-3.5. A trial group of perhaps 30 or 40 friends and relatives liked ChatGPT well enough, but nobody seemed overcome. Schulman expected modest interest. The public release in November 2022 rearranged his expectations.
His explanation for the reaction is refreshingly product-minded. ChatGPT was easier to use than previous models of similar quality. It may also have crossed a threshold where fewer fabrications and a little more self-awareness made conversation feel dependable. Then people showed one another what to do with it. Capability, interface, and imitation formed their own feedback loop. The research team had been comparing ChatGPT with a stronger system in the building. Everyone else compared it with the tools available the day before.
Thirty or forty familiar testers produced a polite response. A simple public chat box produced a cultural event. Sometimes distribution is an experiment the laboratory cannot simulate.
Schulman's public manner remained closer to the seminar room than the launch stage. He distinguished immediate misuse risks from more speculative scenarios, resisted treating every forecast as certainty, and returned often to the details of training. His interests listed on X include reinforcement learning, alignment, birds, and jazz music, a sequence that makes the technical concerns seem like members of a larger household rather than a personal brand.
Leaving, briefly landing, beginning again
In August 2024 Schulman left OpenAI after nearly nine years. He wrote that he wanted to deepen his focus on alignment and return to hands-on technical work. He joined Anthropic's alignment science team, making clear that the departure did not arise from a lack of support for alignment research at OpenAI. The move was a choice about where and how he wanted to spend his time.
The stay was brief. In February 2025 he joined Mira Murati, another former OpenAI leader, in founding Thinking Machines Lab. Schulman became chief scientist. The new lab framed its purpose around a widening gap: advanced systems were progressing faster than the public's ability to understand, customize, and use them. Its answer was not only another general model. It was a workshop for adaptation.
That idea appeared in Tinker, an API for fine-tuning language models, and in Schulman's 2025 technical essay “LoRA Without Regret.” LoRA is a resource-efficient way to adapt a model by training small additional matrices rather than revising every original parameter. The essay examined how to make those adaptations more faithful to full fine-tuning. Once again, Schulman was working on the boundary around an update: how to change a learning system economically without losing the behavior you wanted.
Thinking Machines extended the argument in 2026. It previewed interaction models designed for continuous collaboration, announced large-scale computing infrastructure with NVIDIA, and released Inkling as an open-weights model meant to be customized. The company also argued for staged access rather than treating “open” and “closed” as the only available positions. These are institutional choices as much as technical ones. Who may alter a model? What evidence should precede broader access? How can expertise living inside an organization reach the system it uses?
A career that began with a brittle robot demo has arrived at the same stubborn question on a much larger stage: how should intelligence learn from us without becoming trapped by our first instructions?
The durable loop
Schulman's trajectory has the satisfying recurrence of a jazz phrase. The BattleBots project never got built, but taught him how to teach himself. Physics yielded to neuroscience when the experiments did not excite him. Neuroscience yielded to robotics when the machines seized his attention. Hand-built robot routines yielded to reinforcement learning when they proved too brittle. Raw language models yielded to human feedback when fluent continuation was not yet useful conversation.
None of these turns erased the previous one. Physics supplied mathematical discipline. Neuroscience supplied the question of intelligence. Robotics made failure tangible. Reinforcement learning supplied a method for turning failure into information. At OpenAI, that method helped shape a product used around the world. At Thinking Machines, it has become an argument about agency: capable AI should not arrive as an inscrutable verdict from a remote laboratory. People should be able to work on it, adapt it, and understand what changed.
There is a modest ethic inside all this machinery. Begin with a goal, try something, observe the result, and update without moving so far that you lose what you have learned. It describes PPO. It describes good experimentation. It also describes Schulman's career. The universe did not yield to undergraduate physics research, the knot did not yield elegantly to hand engineering, and ChatGPT did not inspire awe in its small test group. Each disappointment carried information. He kept the signal and changed the policy.