Profile / 026
Lilian Weng ◆ Research, safety and the habit of learning in public ◆ September 29, 2026

People / Artificial intelligence

Lilian Weng and the Art of Learning in Public

She taught a robot hand to solve a Rubik’s Cube, helped build OpenAI’s safety systems, and turned her private learning notes into a public map of modern AI. After co-founding Thinking Machines Lab, she returned to OpenAI in 2026 to work on how research itself might accelerate.

A human hand can turn a Rubik’s Cube with scarcely a thought. A robot hand, confronted with the same plastic puzzle, has to contend with an entire catalogue of indignities: fingers miss, cubes slip, cameras misread colors, and the real world refuses to behave quite like a simulation. Lilian Weng spent her early years at OpenAI helping a team make that hand solve the cube. She remembers the project as exciting and difficult. She also remembers the people. The result required reinforcement learning, visual perception, hardware work and what she later called “crazy amounts of domain randomization.” A small puzzle demanded a large collection of skills.

That scene makes a useful opening for Weng’s career. She has repeatedly moved toward problems whose neat diagrams conceal messy systems. At OpenAI she went from robotics to applied AI research, then to model safety. At Thinking Machines Lab she helped start a new research company. Her current work, after returning to OpenAI in 2026, concerns research acceleration and recursive self-improvement: the possibility that AI systems could help improve the processes that produce their successors. Meanwhile, a blog she began as learning notes in 2017 has offered the public a remarkably clear view of the changing questions.

The child with the encyclopedia

Weng has told an audience that, as a child, she liked reading the family encyclopedia. Astronomy, biology and mathematics drew her in. She even mentioned astrology, a pleasing reminder that curiosity does not arrive with a peer-review stamp. She also competed in mathematics before university. There, she encountered students whose natural talent impressed her enough to reset her sense of herself. She described the lesson as humility: there would always be someone better, and hard work was the way to stand beside them.

She went on to study at Peking University and earn a PhD at Indiana University Bloomington. Her doctoral work examined how information moves through social networks, including the competition for attention among ideas online. It was a different research world from the one she now inhabits, but there is a thread: a concern with systems made from many interacting parts. A message spreads because people, networks and limited attention all meet. A robot hand succeeds because code, cameras, motors and training conditions agree for a moment. Weng’s later writing has the same taste for taking a complicated mechanism apart.

Her route into OpenAI passed through engineering work, including time at Affirm. By 2018 she was at the AI lab and working on robotics. The Rubik’s Cube project offered a particularly stubborn form of learning: a robot trained in simulation had to handle an object in the physical world. Weng said the team solved it without real-world training data for the task. She later recalled how closely the simulation, reinforcement-learning, vision and hardware groups worked together. A computer can practice more quickly than a person, but it still takes people to decide what practice should mean.

Lilian Weng speaking with Connie Chan at a HYSTA fireside chat in 2025
Weng, right, in conversation at HYSTA’s 2025 Intelligence Unbound event. Her account of math competitions and encyclopedias gave the research career a human beginning.

A cube, then a product

In 2021 Weng began leading OpenAI’s Applied AI Research team. The work changed scale. The question was no longer whether a hand could manipulate a cube. It was how powerful language models could be used by other people, with all the unpredictable behavior that invites. Weng described work on evaluations for unwanted outputs, a taxonomy and classifier for problematic content, and techniques to make models less likely to produce unsafe responses. The team also helped lay groundwork for fine-tuning and embedding APIs and moderation endpoints.

Those nouns can sound administrative next to a robotic hand. They are where research becomes a service. An API is a promise that somebody else can build on your system. A moderation endpoint is one of the ways a model’s behavior gets checked when it meets users. Evaluations turn a vague worry into something a team can test. Weng was listed as the applied research lead among contributors to GPT-4. Her account of the work suggests a practical habit: define a failure, build a way to see it, then improve the system in response.

“I believe in the power of learning and it is never too late to learn.”Lilian Weng, on her working philosophy

She had already begun doing some of that definition in public. In 2017, Weng started Lil’Log to document what she was learning. The name is modest; the entries are not slight. They walk through research on reinforcement learning, diffusion models, prompt engineering, language-model agents and safety, with diagrams, papers and the occasional equation intact. Weng has described the blog as a way to keep her own curiosity active. For readers, it became an unusually generous set of notes from someone doing the work inside a leading lab.

Her 2023 essay on LLM-powered autonomous agents arrived as the field was experimenting with models that could do more than answer a prompt. Weng organized the idea into parts that a reader could inspect: planning, memory and tool use. Planning breaks a goal into steps and revises them. Memory carries context beyond the immediate response. Tools allow the model to fetch information or act outside its own text. The essay did not make agents magically reliable. It gave people a vocabulary for asking why one fails.

The work behind the model

After GPT-4, Weng moved further into safety systems. OpenAI named her head of Safety Systems in 2023 and placed her among technical experts on its Safety and Security Committee in 2024. By the time she left OpenAI that November, she said the safety team had grown to more than 80 scientists, engineers, project managers and policy experts. She had briefly served as vice president of research, safety. A group of that size has to do more than publish papers. It needs standards, experiments and working machinery that can be used across products.

80+People Weng said were on OpenAI’s Safety Systems team when she left in 2024.

The public writing from those years reflects the breadth of that assignment. One Lil’Log post examined adversarial attacks on language models: ways a user might try to make a model ignore its intended behavior. Another examined hallucinations, narrowing the term to claims that could be checked against context or outside knowledge. Still another considered the quality of human data used to train models. These are distinctly unglamorous questions compared with the spectacle of a cube-solving robot. They matter precisely because products have to survive ordinary use, odd use and hostile use.

Weng’s description of teamwork is similarly unsentimental. In an OpenAI interview, she said she was willing to take on “dirty” work when it was the largest blocker or created the most value for a project. She encouraged colleagues to do the same. That is a particular idea of leadership: remove the obstruction in front of the team, even if the task will make a poor conference slide. It also explains a connection between her blog and her job. The act of laying out a field carefully can be another way to clear a path.

A new lab, and a return

At the end of 2024, Weng left OpenAI. In 2025 she joined Mira Murati and other researchers in founding Thinking Machines Lab. The new company set out to build AI systems that people could use and adapt, drawing attention because its founders had helped make some of the field’s best-known models. Weng’s own route into the venture had already crossed research, deployment and safety. At a 2025 public conversation, she spoke about the learning curve at OpenAI and the pull of new work. The scientist who had learned from math contests was still measuring a career by what remained to be understood.

She left Thinking Machines Lab in July 2026 and returned to OpenAI. The company said she would lead a top-level team supporting internal research acceleration, including cross-team work on recursive self-improvement. The phrase can summon science-fiction images of a machine rewriting itself in an instant. Weng’s July 2026 Lil’Log essay takes a more concrete route. It asks what happens when an AI model can help improve the workflow around it: its tools, memory, context, tests and the code that coordinates them.

She calls that surrounding system a harness. A harness determines how a model plans, calls tools, stores artifacts and evaluates its results. It can make the same underlying model more useful across a long task, much as a good laboratory procedure makes an experiment easier to repeat and inspect. Weng’s essay traces research on automating workflows and improving the harness itself. She is careful about the limits: a clever loop does not create intelligence from nothing. The underlying model must be capable enough to make useful improvements. That caution is part of what makes the argument worth reading.

Her 2026 writing also revisited scaling laws, the relationships between model size, data, compute and performance. The subject sounds distant from a child’s encyclopedia, yet the motive is familiar. She is trying to understand where improvement comes from, what a graph actually says, and which detail changes the answer. The blog’s chronology now doubles as a history of AI’s shifting center of gravity: from learning to control an object, to deploying models, to designing systems that can sustain longer work.

The notebook stays open

There is an appealing irony in Weng’s public identity. She has held senior titles and helped build large teams, but many people know her through a website that opens, simply, by saying she has been documenting learning notes since 2017. A note is provisional. It invites correction. It can also be shared before every question has a tidy answer. For a field that moves as quickly as AI, this may be a more useful posture than pretending the map is complete.

Weng has said she wants AI to reduce repetitive work, speed scientific discovery and interact safely with the physical world. Her career has visited each part of that sentence: the robot hand in the physical world, the applied research that reached users, safety systems built for deployment, and now research about improving the research process. It is too early to know what the current chapter will produce. The pattern, though, is visible. When the mechanism gets complicated, Weng opens the notebook, names the parts and starts learning again.