At the small London office of DeepMind in 2012, Ioannis Antonoglou expected a first interview. A recruiter had sent him to meet Shane Legg, one of the lab's founders. Legg explained the proposition: the team wanted to build intelligence that could learn across tasks. Antonoglou, fresh from postgraduate study, found the ambition absorbing enough to offer his help even if the interview failed. Legg asked whether he had time to stay. Antonoglou did. He moved through interviews from noon until seven in the evening, and the next day the company called to hire him.
He was 23, by his own recollection, and DeepMind had roughly 25 people. He thought the work might make for two fascinating years before he found a “real job.” He stayed eleven. The joke is gentle, but the distance between those two estimates contains much of his story: a young engineer follows a question into a small company, and the question grows large enough to occupy a career.
The workshop rule
Antonoglou was born and raised in Thessaloniki. His father ran a small workshop with his brother; his mother raised the children. He remembers the 1990s as years of neighborhood games, soccer and time outside. He also remembers a family attitude toward effort. If a child said he wanted to become an astronaut, his father would not wave the idea away. He would ask what steps the child would need to take. The dream was allowed to remain extravagant; the plan had to become concrete.
That habit stayed with Antonoglou. He liked mathematics and technology, and chose electrical engineering at Aristotle University of Thessaloniki because it left several doors open. He has said he applied only to the university in his hometown. There was no romantic exile in the first chapter of this AI career. The local option was the one he wanted. Later, postgraduate study took him to Edinburgh, then a doctoral program at University College London.
At university, work with CUDA programming drew him toward artificial intelligence. The subject offered an unusually direct version of the puzzle that had interested him: how can a machine make decisions, and how can it get better at them? The answers available in 2012 were modest beside today's claims. DeepMind was still an obscure startup. Its ambition, though, had a precision he found hard to ignore. Build systems that learn. Measure them. Improve them. Repeat.
“When you dream of something, you shouldn’t ask if it can work - but how you can make it work.”Ioannis Antonoglou
A board with too many possibilities
Games were useful laboratories for that ambition. They offer rules, scores and endless opportunities to try again. At DeepMind, Antonoglou first worked on deep Q-networks, a way to learn decisions from experience in video games. He later worked with David Silver, whom he describes as both a professor and a long-term mentor. Their next famous test had a wooden board and black and white stones.
Go had long resisted the kind of brute-force search that helped computers master chess. There are too many plausible moves, and the quality of a position can be difficult to explain even when an expert senses it. Antonoglou described the challenge as giving a machine something like an instinct for whether a move looks promising. AlphaGo combined learning and search to make that judgment useful. In March 2016 it beat Lee Sedol by four games to one in Seoul. Antonoglou was part of the team there.
The match had the strange atmosphere of sport and experiment occurring at once. It was front-page news in Korea, Antonoglou recalled, and screens in public spaces carried the games. Move 37 in the second game became a symbol of the machine's unexpected choices; Lee's own Move 78 in the fourth showed that a human could still find a way through. The result was decisive, but the memorable drama came from the back and forth: neither side behaved like a tidy illustration in a textbook.
The follow-up research kept removing scaffolding. AlphaGo Zero learned through self-play without relying on a library of human games. AlphaZero extended the approach to chess and shogi as well as Go. MuZero pushed further: it learned to plan in environments without being handed their rules. Antonoglou appears among the authors of the papers behind those systems. His UCL thesis, completed in 2023, bears a fitting title: Learning to Search in Reinforcement Learning.
This sequence can sound like a march toward inevitability when compressed into a line of names. Research is rarely so polite. Each system solved a particular problem under particular conditions. A Go board supplies a clear definition of winning; an office computer does not. Antonoglou has been frank about that distance. Games make good tests because they are contained. The world outside them is messy, ambiguous and capable of changing the task halfway through.
The problem with “sometimes”
Near the end of his DeepMind tenure, the work moved from games to language models. Antonoglou says he led reinforcement learning from human feedback work for Gemini. That branch of AI uses human preferences and other feedback to shape a model's behavior. It also brought him into closer collaboration with Misha Laskin, the colleague with whom he would later start Reflection AI.
The attraction of agents is easy to understand. A person can ask a system to arrange a schedule, work through a codebase or complete a task across several pieces of software. A useful agent would do the work and return a result, rather than stopping after a plausible paragraph. Antonoglou describes a continuum from tasks that take a person minutes to ones that take hours or longer. The farther along that continuum an agent travels, the more chances it has to misunderstand, lose context or make a wrong turn.
In a Sequoia conversation, Antonoglou called the central obstacle the “sometimes” problem. A model that does something impressive once is still a poor colleague if it cannot do it consistently. The language has a dry engineer's edge. It asks whether a demo survives the ordinary Tuesday on which someone needs an invoice checked, a repository changed or a series of calendar conflicts resolved. A victory on a fixed board can be counted. Trust in an agent has to be earned across untidy repetitions.

A company between three clocks
Antonoglou and Laskin founded Reflection AI in March 2024. Laskin became CEO; Antonoglou became president and CTO. The company began with a bet that reinforcement learning could make AI agents more capable. It later expanded its ambition to training frontier base models and releasing open weights, so others could run, inspect and customize them. In October 2025 the company announced that it had raised $2 billion to pursue that strategy. Money can buy computing capacity; it cannot settle the “sometimes” problem by itself.
Even the choice of city had a practical explanation. Antonoglou had spent years in London, knew people he wanted to work with there, and needed access to the talent and investors of California. London and San Francisco sit eight hours apart, an awkward gap for a shared workday. New York offered a bridge. Reflection established offices in New York, San Francisco and London. It is a small anecdote with a useful lesson: before a research team can coordinate its ideas, its members need hours in which they can speak to one another.
For Antonoglou, openness is part of the technical plan, not simply a distribution choice. He argues that releasing model weights gives other researchers the chance to test systems, find blind spots and improve them. Organizations can also adapt and operate models under their own control. He has discussed this in terms of sovereign AI: the relevant question for a country or company is whether it can deploy and shape the system it relies on. His public case is ambitious, and Reflection's results will have to support it as models are released.
The argument is connected to his experience at DeepMind. Earlier generations of AI research were published openly; the papers on DQN and the game-playing systems gave outsiders methods to study and challenge. After the competition around large language models intensified, he says, frontier labs published less. Reflection was created in part to recover a culture in which capable models and their methods can be examined by a wider community. There are commercial questions inside that choice, and hard questions about how capable systems are tested before release. He discusses both as engineering work that cannot be wished away.
What counts as a finished move?
There is an appealing symmetry in Antonoglou's career. His father met an impossible-sounding childhood ambition by asking for the steps. DeepMind turned a broad question about intelligence into a sequence of games that could be measured. Reflection is trying to bring some of that discipline to the open-ended work people do on computers. The symmetry should not make the new task look easy. A game ends when the rules say it ends. A work task is finished when a person can use the result and trust what happened along the way.
Antonoglou's practical definition of an advanced agent is a system that can interact with software and do work across ordinary computer workflows. That definition is less theatrical than talk of machine minds, and more demanding. It asks the model to operate tools, recover from errors and keep a goal intact across steps. It also gives a way to judge his new company: watch what its agents can finish, how reliably they finish it, and how much control users have over the models underneath.
The young engineer who spent an unexpected afternoon interviewing at DeepMind has become a founder with a much wider stage. He still speaks in terms of learning, testing and iteration. A stone placed on a Go board has an immediate consequence. A model placed in other people's hands has many. Antonoglou is betting that the same patience that helped a machine learn to play can help make one useful beyond the board, and that more people should be allowed to see how it works.