Field notes
2026 Sakana AI opens a Frontier Intelligence Group2025 Jones asks what follows the transformer2017 Eight researchers publish Attention Is All You Need

People / Artificial intelligence

Llion Jones Helped Invent the Transformer. Now He Wants Room to Think Beyond It.

A paper he co-wrote changed artificial intelligence. In Tokyo, the Sakana AI co-founder is trying to protect the freedom that made such a paper possible - and see what comes after it.

A person can spend a career trying to make one idea work. Llion Jones has the rarer problem of having helped make an idea work so well that it became the default. In 2017, he and seven colleagues at Google published Attention Is All You Need, a paper about machine translation. Its transformer architecture later became a central ingredient in the large language models now familiar to millions. The paper's title sounds like a dare. Nine years on, Jones is interested in what attention may have left out.

At Sakana AI, the Tokyo company he co-founded in 2023, Jones has a title that comes with meetings, hiring and public talks: chief technology officer. He has described his days as different from one another, shifting between managing people, managing projects and speaking at conferences. Yet one question links the research and the management. What conditions let a group of people stumble onto an idea no one knew to request?

He has an unusually direct answer. In a 2025 talk, Jones recalled the freedom the original transformer team had to follow its curiosity. Today, he worries that a crowded, expensive AI industry gives researchers strong reasons to work on the ideas most likely to pay off soon. Sakana AI is, in part, an attempt to make a little space around the unknown again.

Before the paper had a title

Jones studied artificial intelligence and computer science at the University of Birmingham, then completed a master's degree in advanced computer science. He remembers the lecturers as supportive and the instruction as strong. In a video for the university, he even suggested that the school's reputation helped him get into Google without a referral. It is a modest, concrete account of how a research career begins: with teachers, a CV and somebody willing to read it. Birmingham awarded him an honorary degree at its July 2026 graduation ceremonies, a return to the institution he credits for setting him up.

Before Google Research, Jones worked as a software engineer at Delcam and then at YouTube. He joined Google Research in 2015, where he joined a group working on a stubborn language problem. Earlier systems processed sequences in ways that made training slow and difficult to parallelize. The 2017 paper proposed a model built around attention mechanisms. For a reader outside the field, the useful picture is that the system could weigh relationships among words across a sentence without marching through them one by one in the older fashion. The immediate target was translation. The wider consequences took years to unfold.

Eight names appeared on the paper: Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan Gomez, Łukasz Kaiser and Illia Polosukhin. Listing all eight matters. This was a collaboration, and the remarkable result belongs to a team. Jones has spoken about the informal, low-pressure setting in which the work developed, including ideas sketched on whiteboards over lunch. That detail is easy to romanticize, but it points to something practical: no manager can schedule a surprising result for Thursday afternoon.

8Authors on the 2017 paper
2023Sakana AI founded in Tokyo
2025Continuous Thought Machines paper

The transformer traveled far beyond the translation task. It became the architecture behind many modern generative systems, and the letter T in ChatGPT. Jones can tell a new audience that he helped build that T, a line short enough to fit into a lecture and large enough to explain why people listen. The risk, as he sees it, is that success can turn a good tool into the only tool people care to investigate.

A company named for fish

Jones left Google Research to start Sakana AI with David Ha and Ren Ito. Its headquarters are in Tokyo, where Jones is based. Sakana is Japanese for fish. Jones has explained the name with a simple contrast: a single fish may not do very much, while a school can display coordinated behavior. In the company's research, that image becomes an invitation to ask what multiple smaller systems might achieve together.

The company's founding pitch in 2023 was deliberately exploratory. Instead of simply building larger versions of the prevailing model, its founders talked about new architectures and the possibility of many specialized models working together. These were research directions, not promises of a finished successor. In an industry accustomed to enormous models and even larger funding rounds, the fish were a useful reminder that intelligence might come from arrangement as well as scale.

Jones also wanted to recreate a working atmosphere. In his Birmingham alumni interview, he described telling new hires, “I want you to work on what you think is interesting, is important.” That is a management instruction with a hole in the middle where a conventional assignment would sit. The researcher must decide what is worth pursuing. The company has since described research freedom, distinctive ideas and a willingness to fail as explicit principles of its Frontier Intelligence Group, announced in September 2026.

“I want you to work on what you think is interesting, is important.”Llion Jones, describing his message to new colleagues

There is no shortage of work for those colleagues. Sakana AI has explored model merging, collective intelligence and automated research, alongside products for Japanese users and organizations. Jones's personal thread runs through the deeper question: how much of today's AI progress comes from improving an established recipe, and how much might come from changing the recipe itself? He is careful about the distinction. In a 2026 discussion, he said he had not suddenly become opposed to transformers and acknowledged the progress still being made with them.

Giving the neurons a clock

One answer arrived in a 2025 research paper called Continuous Thought Machines. Jones was one of five authors, alongside Luke Darlow, Ciaran Regan, Sebastian Risi and Jeffrey Seely. The work asks what happens when an artificial network pays closer attention to timing and coordination among its units. Ordinary artificial neurons are simplified mathematical devices. The new model gives them richer behavior over time and uses their synchronization as part of its internal representation.

That description is technical, but its origin is plain enough. In brains, when activity happens can matter as much as which cell is active. The team borrowed that observation without claiming to have built a literal brain. Its demonstrations included solving mazes and handling images, and the researchers released code and interactive material for others to inspect. The paper presented an approach to investigate, with encouraging behavior, rather than a final verdict on the architecture that should run tomorrow's AI.

Jones described the method at a March 2026 event for MUFG Innovation Partners. The researchers knew synchronization mattered in biological systems, he said, so they added a version of it to artificial networks and watched for what emerged. He pointed to calibration, how closely a model's expressed confidence matches its actual accuracy, as one of the interesting results. The larger habit is recognizable: start with a serious question, make something testable, and allow the result to surprise you.

Llion Jones speaking on stage during MUIP Innovation Day 2026 in Japan
Jones talks beyond the transformer at MUIP Innovation Day, March 2026. Photo: MUFG Innovation Partners.

The missing million examples

At that same event, Jones singled out two gaps in current AI. One is data efficiency. A person can learn a new task from relatively few examples, while many AI systems need immense training sets. The other is generalization: carrying a lesson from one setting into a genuinely different one. Jones used driving as an example. A human who learned in one city can adapt to another; a machine trained on one city's roads may have a much harder time.

These questions have consequences beyond the laboratory. A company may have only a small amount of data about a particular client or process. A useful system must make sense of sparse evidence and unfamiliar situations, not merely perform well on examples that resemble its training set. Jones argued that progress on data efficiency and generalization could open new uses for AI in finance. He also warned against treating scale alone as a complete answer to those problems.

His way of talking about the commercial moment is less apocalyptic than some headlines. Asked whether AI was a bubble, he offered a balloon as the better image: it might need adjustment, but the underlying technology is producing real value. He also told the audience that capability moves so quickly one has to try the tools again each month to see what has become possible. It is a useful correction to the idea that someone exploring beyond transformers must think the current systems are useless. Jones clearly does not.

From one architecture to an open question
2006BSc in artificial intelligence and computer science at Birmingham.
2009MSc in advanced computer science.
2012Begins work at YouTube after software engineering at Delcam.
2017Attention Is All You Need introduces the transformer.
2023Jones co-founds Sakana AI in Tokyo.
2025Continuous Thought Machines and a TEDAI talk on research freedom.
2026Sakana AI describes its Frontier Intelligence Group.

The permission to wander

When Jones spoke at TEDAI San Francisco in October 2025, his subject was competition. The talk's argument was unsettling because it came from someone who had benefited from a breakthrough. Money and attention had poured into AI after the transformer; at the same time, he believed the pressure to publish and beat rivals was narrowing the kinds of questions researchers could ask. The architecture had succeeded. The conditions that produced it were harder to reproduce.

Sakana AI's Frontier Intelligence Group makes that concern tangible. Its September 2026 announcement described regular meetings to discuss ideas outside the mainstream, invitations to outside researchers, and time for long-term experiments. Among its stated values are freedom to fail and work a researcher believes is unique. The company also points to curiosity about nature, from groups of fish to the dynamics of individual neurons. Those examples make the group's ambition more specific than a slogan about innovation.

None of this provides a reliable map to the next breakthrough. That is precisely the awkwardness. A lab can hire people, reserve time, fund experiments and make failure survivable. It cannot order originality as if ordering office chairs. Jones knows this from the paper that gave him his public reputation. What the group can do is keep multiple questions alive long enough for evidence to distinguish an odd idea from a dead end.

There is a pleasing loop in his account of his own career. Birmingham's supportive teachers helped set him on a path to Google. An unusually free research environment helped his team produce a consequential paper. He now spends part of his working life trying to give colleagues that same latitude. The scientist's job and the executive's job meet at a deceptively plain sentence: work on what matters to you.

The transformer remains useful, and Jones says so. Meanwhile, the fish keep swimming, the neurons in Sakana's experiments keep time, and researchers keep asking whether intelligence has other shapes. Jones's wager is that the answer is worth looking for, even when nobody knows what it will look like.