On the first day of the 2020 International Olympiad in Informatics, Walden Yan found himself tied at the edge of a gold medal. The contest had moved online that year. Students worked through virtual machines hosted in Singapore, watched by cameras and local proctors, with five hours to solve each day’s set of problems. Yan, then a first-year Harvard student, had enough points to be in the gold group, just barely. Day two would decide whether more than a thousand hours of practice ended with a medal or a near miss.
He finished 19th among nearly 400 competitors from more than 100 nations. The result gave him gold, and a public story with an irresistible shape: a self-taught programmer, a remote contest, a last-day climb. Yan had learned to code with YouTube videos. He had trained on old problems until the opening lines of a new one could suggest an algorithm before he finished reading it. He compared the repetition to sport. Stop training for months, he said, and the skill goes rusty.
The interesting sequel came outside the contest. Competitive programming rewards a clean answer to a defined problem. Real software work arrives with old code, partial instructions, forgotten decisions and a colleague who knows a detail nobody wrote down. Yan eventually became a co-founder and chief product officer of Cognition, the company behind Devin, an AI software engineer built to work through those messier tasks. The contest clock had been replaced by something harder to see: whether an agent could keep the right facts in mind long enough to finish.
The education of a pattern hunter
Yan’s early work already joined two interests that would later meet at Cognition: programs and the possibility of having machines produce them. While at Avon High School in Connecticut, he joined MIT PRIMES, a research program for high school students. With William S. Moses, he co-authored a 2020 paper on neural program synthesis. The research tried to improve models that turn a description into code by accounting for the way programming languages pair tokens and nest structure. It was an early encounter with a stubborn fact about code: syntax is visible, but the relationships that make a program work are deeper.
The Olympiad taught a different version of that lesson. One 2020 problem required constructing a graph with a specified number of paths between pairs of nodes. Another had such an odd circular description that Yan struggled to get hold of it. Practice helped him recognize familiar shapes quickly; unusual shapes still demanded a fresh idea. His reward for those hours was an impressive rank. His longer-term reward may have been a habit of treating apparent magic as the product of a system that can be studied.
After the contest, Yan said he wanted to explore machine learning and website design, then put his programming ability to good use. Those ambitions sound modest beside what followed. They also sound like the young engineer who would later write about the unromantic details of AI agents: what they see, what they forget and how one apparently small decision changes the work that comes next.

A green line on a terminal
Yan has described one early Devin moment with the detail of someone who still remembers the suspense. The team was trying to start a MongoDB server. The documentation was difficult, and the humans were stuck. He turned to a rough prototype and asked it to get the server working. The agent ran commands. It deleted files. The team watched with understandable concern. Then the command line came back green.
This was a small job, at least when compared with the ambition of building an autonomous software engineer. It mattered because the prototype crossed from suggestion into action. It had encountered a real environment, made changes and produced a result the team could inspect. Yan recalled the room asking how far they could push it. Cognition’s founding team, including Scott Wu and Steven Hao, built the answer into Devin, which launched in 2024.
The temptation in a story like this is to let the green line stand for a solved problem. Yan’s own writing has taken the opposite direction. As Devin moved from a striking demo into product work, he concentrated on why capable agents can go astray. A coding task may last long enough for early assumptions to disappear from view. Its instructions may leave out the very convention that determines whether a change belongs in the codebase. The agent can edit a file correctly and still miss the larger intention.
The argument against the crowded room
In June 2025, Yan published an essay with a title designed to interrupt a fashionable conversation: “Don’t Build Multi-Agents.” His target was a common architecture in which one agent divides a task among several others and later tries to combine their work. The diagram looks efficient. The difficulty is that each branch makes choices the others cannot see. Two agents asked to build different parts of the same game might each produce perfectly plausible pieces that belong to different games.
Yan identified two principles. Share the full context of a task, including the trail of earlier actions, rather than a tidy summary of the latest instruction. And recognize that an action carries decisions even when nobody has named them. A variable name, a visual style, an edge case or an interpretation of the brief can become a commitment. Give separate agents separate pieces of the story and they may make incompatible commitments before anyone notices.
He called the larger discipline “context engineering.” The phrase points to a practical job: deciding what information an agent needs at each step and how to preserve it over a long run. It lacks the theatrical quality of a swarm of digital workers. That is part of its appeal. Anyone who has joined a project halfway through knows how much time is spent reconstructing why the code looks the way it does. An agent faces a similar problem, except that its memory is designed, constrained and continually filled by tool output.
The 2025 essay also suggested a straightforward starting point: a single agent working in sequence, with its context kept continuous. When that history grows too large, the system needs a careful way to compress what has happened. Cognition, Yan wrote, had experimented with smaller models for that job. The picture was less exciting than sending ten agents racing in parallel. It was also easier to reconcile with what a coherent software change requires.
A public revision
Ten months later, Yan published “Multi-Agents: What’s Actually Working.” The title answered the earlier warning with a useful qualification. Models had improved, users were running more agents, and Cognition had begun to deploy arrangements that helped in practice. Yan did not erase his concern about parallel writers. He described a narrower class of systems in which extra agents contribute intelligence while code changes remain coordinated.
One example pairs Devin with a dedicated reviewer. The reviewer starts with a clean view of the diff, free from the coding agent’s long trail of attempted fixes. It can question an implementation that the original agent has grown accustomed to. Devin then interprets the review in light of the user’s request and its own broader context. Another pattern gives a working model a “smart friend” it can consult on difficult decisions. The problem becomes how to ask for help, how much context to pass and how to use the answer.
There is a pleasing engineering modesty in this shift. The public claim changed because the team’s experiments changed. The old warning still applied to agents writing in parallel without a shared understanding; the new evidence supported review loops and advisory roles. Yan’s follow-up also reported failures. A smaller primary model in one experiment did not reliably know when to seek the stronger model’s help. The system was faster and cheaper, but the quality ceiling remained. It is a rare product essay that makes room for a disappointing result.
That willingness to refine a position fits the career that preceded the essays. Contest practice is a cycle of hypotheses and correction. Research papers are arguments exposed to evidence. Building a product asks for the same habit under less orderly conditions. An agent may produce code that runs today and creates maintenance trouble two weeks later. A clean demo is a beginning, not a verdict.
What the agent should do next
As chief product officer, Yan now speaks about Devin in terms that extend well beyond code generation. In an interview with Anthropic, he described an ambition for an agent that could take higher-level goals, propose its own tasks and coordinate an engineering team to execute them. Cognition’s public work has also widened toward review, debugging and incident response. In July 2026, Yan announced that the TierZero team was joining Cognition, pointing to the work of keeping software running after it ships.
The aspiration is expansive, but the path remains full of small, stubborn questions. How does an agent know which repository detail matters? When should it ask a human to clarify? Can a reviewer catch an error without losing the original intent? How does a manager agent pass a discovery to another worker before that worker makes a conflicting choice? Yan’s essays return repeatedly to these questions because they sit between an impressive answer and dependable work.
There is a neat symmetry in the journey from the 2020 contest to Cognition’s current work. At the Olympiad, Yan trained himself to see the hidden structure inside a strange problem statement. At Cognition, he is trying to give software agents enough context to do something similar inside a living codebase. The first challenge ended after two days and had a scoreboard. The second has no final whistle. A green command line is still satisfying. The harder achievement is making it green for the right reasons, again and again.