Latest
Jimmy Ba announced his departure from xAI in February 2026 ◆ Adam received ICLR's 2025 Test of Time Award

People / AI research

Jimmy Ba and the Algorithms That Learned to Learn

Before he helped build xAI, Jimmy Ba helped solve two quieter problems inside modern AI: how a network takes its next step, and how it stays steady along the way.

Ten years after Jimmy Ba and Diederik Kingma introduced a way to train neural networks, the paper received a prize with an unusually honest name: the Test of Time Award. The 2025 honor from ICLR recognized Adam, an optimizer whose short name had long since escaped the title page and entered the working vocabulary of machine learning. It was a curious kind of fame. You could use the method every day and still know little about the person who helped make it.

Ba’s story is easiest to see from that distance. He studied electrical and computer engineering at the University of Toronto, earned a doctorate under Geoffrey Hinton, joined the computer science faculty, and later helped found xAI. By the time he appeared onstage for the Grok 4 launch in July 2025, he had already spent more than a decade working on the less theatrical problem beneath AI products: how a machine learns without wasting its next move.

There is an elegant mismatch here. The public conversation about AI often begins with what a chatbot says. Ba’s research begins several steps earlier, with the calculations that determine whether the system can learn to say anything useful at all. In that work, the drama is measured in gradients, training steps, and stability. It seldom gets a stage.

2014Adam paper first posted
2016Layer Normalization paper
2025ICLR Test of Time Award

The patient machinery of learning

A neural network trains by repeatedly adjusting its internal numbers. Each adjustment needs a direction and a size. Take steps that are too large and the process can become erratic; take steps that are too small and learning crawls. Adam offered a practical way to adapt those steps using running estimates of the gradients. Its authors described a method that was straightforward to implement, efficient with memory, and suited to noisy or sparse information. Those qualities gave it a long working life.

Kingma and Ba posted the paper in December 2014 and presented it at ICLR in 2015. The award ten years later was recognition of use, not merely promise. ICLR described Adam as a default optimizer for countless models across computer vision, language processing, and reinforcement learning. The research had settled into the infrastructure of other people’s research. That is a peculiar achievement: an idea can become so familiar that its author disappears behind a setting in a training script.

Two years after Adam’s first posting, Ba was first author on another paper with Jamie Ryan Kiros and Hinton. Layer Normalization proposed a way to normalize the activity within a layer for each individual training example. The distinction mattered because an earlier approach, batch normalization, depended on a group of examples processed together. Ba and his co-authors showed how their method could work consistently during training and evaluation and could be applied to recurrent networks.

The pair of papers makes a useful portrait of his interests. Adam asks how to choose the next adjustment. Layer Normalization asks how to keep the calculations orderly enough for the next adjustment to matter. Neither paper is a grand declaration about what a machine should become. Both are attempts to make the act of learning less fragile.

“How can we build general problem-solving machines with human-like efficiency and adaptability?”

Jimmy Ba, on his research goal

Toronto, three times over

Ba’s academic record has an uncommon consistency. He earned his bachelor’s degree in 2011, master’s in 2014, and PhD in 2018, all in electrical and computer engineering at the University of Toronto. His personal biography names Brendan Frey and Ruslan Salakhutdinov in connection with his earlier degrees and Hinton as his doctoral supervisor. The names map a research community in which machine learning was both a technical discipline and a shared conversation.

His early publications show a willingness to interrogate the tools as much as the tasks. With Frey, he wrote about adaptive dropout, a method for training deep networks. With Rich Caruana, he asked in a paper title whether deep nets really needed to be deep. In a later collaboration, “Show, Attend and Tell,” he helped explore how attention could guide image captioning. These were different problems, yet they return to a recognizable habit: inspect the mechanics of learning before assuming the architecture alone will do the work.

In 2018, Ba joined Toronto’s computer science department as an assistant professor and became part of the Vector Institute community. He was named a Canada CIFAR AI Chair. Teaching followed the research: his website lists courses on neural networks and deep reinforcement learning. The pages of publications continued to grow, including work on the Lookahead optimizer, whose memorable subtitle was “k steps forward, 1 step back.” Even its name sounds like advice an impatient engineer might need.

Recognition arrived in increments. Toronto announced his Sloan Research Fellowship in 2023, citing contributions that made deep learning more efficient and reliable. In July 2024, the university promoted him to associate professor with tenure. The Adam Test of Time Award came the next spring. It is difficult to design a better sequence for a researcher interested in durable methods: first a faculty appointment, then a fellowship, then a prize awarded only after others have spent years putting the idea to work.

  • 2011Completes a Toronto engineering bachelor’s degree.
  • 2014Completes a master’s degree; Adam appears later that year.
  • 2018Finishes the PhD and joins Toronto’s computer science faculty.
  • 2023Joins xAI’s founding team.
  • 2026Announces his departure from xAI.

The equation reaches the stage

When xAI announced its formation in July 2023, Ba was among its named founding researchers. The move put him in a company whose work would be judged in public, quickly and sometimes noisily. Toronto remained part of his professional identity even as xAI put him closer to large model launches. The combination was unusual but coherent: a scientist of learning systems inside a firm trying to build them at scale.

The July 2025 Grok 4 presentation offered an unusually literal picture of that transition. Ba sat with Elon Musk and fellow co-founder Tony Wu before a dark stage backdrop and a chart about reinforcement learning compute. The research vocabulary had become product-launch vocabulary. Years earlier, Ba had been asking how to make neural networks train efficiently. Now he was helping present a model built by a team attempting exactly that, at an industrial scale.

Jimmy Ba at right with Tony Wu and Elon Musk during the Grok 4 presentation
01 / ON STAGEJimmy Ba, at right, with Tony Wu and Elon Musk during the July 2025 Grok 4 launch. A training chart is visible behind them.

A few months later, at a San Francisco AI summit, Ba told the audience he had nearly missed his own appearance after his self-driving ride turned onto the Bay Bridge by mistake. The detour gave a researcher of adaptive systems an unscheduled demonstration of why the next move matters. The image is useful because it resists a too tidy account of a career. A stage appearance does not replace a research paper; a research paper does not tell you what it is like to turn an idea into a product. Ba had moved between those worlds. He was part of a company founded in 2023, remained on Toronto’s faculty, and was recognized in 2025 for a method published a decade earlier. Several clocks were running at once.

That same year, he appeared on the Ray Summit speaker lineup, another sign of how his work crossed from academic circles into the operating concerns of AI builders. The public profile of a researcher can be misleadingly episodic: a paper, an award, a launch, a talk. Between them is the continuous work of debugging an idea, testing it against alternatives, and deciding which part of a complex system actually needs attention.

A new direction, in his own syntax

In February 2026, Ba announced that he was leaving xAI. His message thanked the people there and said he would remain close to the team as a friend. He also borrowed a phrase from the technical language of his career: it was time to “recalibrate my gradient on the big picture.” The sentence was characteristic of a researcher who had spent years studying how to choose a next step. It stated a change of direction without laying out a destination.

Ba wrote about the prospect of far greater productivity from better tools and the possibility of recursive improvement in AI. Those were his expectations, not announced products or a disclosed new role. The public record leaves the next professional chapter open. What it does show is a consistent fascination with efficiency and adaptability, from his own research statement through the papers and into his farewell note.

His long-term question on his website asks how to build general problem-solving machines with human-like efficiency and adaptability. It is a large question, almost comically large beside the modest names of individual techniques: Adam, Layer Normalization, Lookahead. Yet the career makes sense when read in that order. An ambitious question gets answered, if at all, through smaller choices about how learning works.

The Test of Time Award provides one answer to why those choices matter. A method written by two researchers in 2014 was still useful enough in 2025 to be honored for its persistence. Ba’s xAI chapter gave his name a place in the public story of AI companies. His earlier work has a different kind of public life: repeated quietly wherever someone starts a training run and asks a machine to take the next step.