YESPRESS / PROFILE◆Guodong Zhang / Machine learning researcher◆Toronto → DeepMind → xAI◆March 2026 / A new chapter◆YESPRESS / PROFILE◆Guodong Zhang / Machine learning researcher◆Toronto → DeepMind → xAI◆March 2026 / A new chapter◆
People / Artificial Intelligence

Guodong Zhang and the Question Hidden Inside Every AI Answer

He studied how neural networks learn, helped build xAI, and then walked away. The thread through his work is a deceptively simple question: when should a machine admit what it does not know?

In June 2023, as he prepared to finish his computer science doctorate at the University of Toronto, Guodong Zhang described a problem that sounds almost impolite in an industry built on answers. Could a neural network recognize the gaps in its own knowledge? It was one of three questions he gave the university when asked to explain his research. The others concerned how to train a model quickly and why it can perform well on examples it has never seen. Together they amount to a practical examination of the machine behind the polished sentence on a screen.

The timing gave those questions unusual weight. Large language models had moved from research labs into ordinary conversation. Zhang was about to leave the university for industry, and within weeks he would be named among the founding team of xAI. The company’s ambitions were expansive. His own published work had a narrower, patient rhythm: isolate a training problem, build a model simple enough to reason about, and see which conclusions survive contact with a real neural network.

That rhythm carried him from Zhejiang University to Toronto, from research collaborations with Google Brain and DeepMind to an operational role in xAI’s coding and image generation efforts. In March 2026, he announced that his time at xAI was over. The distance between the thesis and the departure is about three years. The questions that connect them are older and more durable.

2019Noisy quadratic model paper
2023Toronto PhD and xAI founding team
2026Departure from xAI

A useful appetite for simple models

Zhang’s personal research page has the plain look of an academic homepage. It lists papers on multi-agent optimization, deep learning dynamics and Bayesian deep learning. It also begins with a line attributed to the statistician Herman Chernoff: “Many things are only trivial once you know them.” The line suits a researcher trying to make complicated systems legible. A result can look obvious on a slide after somebody has done the work of finding it.

In 2019, Zhang and seven collaborators asked a costly question: how much faster does training get when a neural network processes more examples at each step? Larger batches can reduce the number of steps needed, but their benefit eventually tapers. The team compared different optimization methods and used a deliberately simple noisy quadratic model to predict where those gains change. In their experiments, methods including Adam and K-FAC could make use of larger batches than momentum-based stochastic gradient descent alone.

There is a practical reason to care about a graph full of curves. Full scale training experiments consume time and computing power. A smaller model that captures the right behavior lets a research team ask better questions before paying for the larger run. Zhang later told Toronto that his doctoral work on this model had been used by industrial labs including Anthropic, DeepMind and OpenAI. That account is his, given in the university’s 2023 interview; the published paper provides the technical argument beneath it.

Graph from Zhang and collaborators' noisy quadratic model paper showing risk over training steps for different curvatures
In a 2019 paper, Zhang and collaborators used a simplified model to reason about the trade-offs in neural network training. The original figure tracks risk as training steps accumulate.

That work sat among a larger set of collaborations. A 2017 paper on deformable convolutional networks, with Zhang among its co-authors, was selected for an oral presentation at the International Conference on Computer Vision. At Toronto, he worked with Roger Grosse and others on optimization, uncertainty and game-like learning problems. The titles can sound abstract. The questions are less remote: how should a model change when it learns, how can it say how sure it is, and what happens when multiple learning systems interact?

Toronto, with a classroom attached

Zhang said Geoffrey Hinton’s work helped draw him to Toronto. In the university interview he recalled the impact of AlexNet’s 2012 image recognition result and the prospect of working around the Vector Institute. This was a young researcher choosing a place where a field had changed direction. It was also a place where he would teach. His site records that he instructed an introductory machine learning course in fall 2021, gave a guest lecture on neural network training dynamics, and served as a teaching assistant across courses from artificial intelligence to neural networks.

The teaching record is an instructive counterweight to the familiar mythology of the secretive AI lab. An instructor has to take concepts that appear in compressed notation and explain them to people who have not yet learned the shorthand. Zhang’s papers did the reverse as well: reduce sprawling training behavior to models that can be tested and argued about. Both activities demand clarity. His site also lists an Apple PhD Fellowship, a Borealis AI Fellowship and earlier awards at Zhejiang University. They mark a long run of academic recognition before a company biography entered the picture.

“I recommend everyone maintains a curious mindset and pursues their passions.”Guodong Zhang, University of Toronto interview, 2023

In that same interview, he spoke about AI’s possible benefits in everyday work, from translation to writing and answering questions. He also argued that powerful systems should remain under human control and that the public should have a part in regulation. Those positions did not require him to give up on the technology. They demanded that the people building it pay attention to its limits. The idea that a model should know when it lacks an answer fits the same habit of mind.

A research career meets a product clock

After Toronto, Zhang worked at DeepMind. Then came xAI. The founding roster published in 2023 included him alongside researchers and engineers from several labs. The move put a scholar of training dynamics into a company trying to build and ship models on a compressed schedule. The work would eventually reach beyond text. By early 2026, accounts of xAI’s structure associated Zhang with its coding effort and Imagine, the image and video product.

A public xAI all-hands presentation in February 2026 made those responsibilities visible. The company laid out work streams for its models, coding, Imagine and other projects. Zhang appeared in the discussion of the generation effort. One colleague described how quickly the team had moved from having no internal diffusion code to releasing Imagine across product surfaces. Public company presentations naturally emphasize momentum; the event still showed how far Zhang’s work had travelled from a paper on optimization curves. A model’s behavior was now bound up with product decisions, updates and a large audience.

The transition is easy to describe and harder to live. A research paper can state assumptions in a paragraph and hold a conclusion within them. A product encounters every input its users think to try. The distance between those settings is where uncertainty becomes operational. Training speed matters, but so do evaluation, the shape of a generated image, and how much confidence a user should put in a machine’s response. Zhang’s public research does not tell us how he settled every trade-off inside xAI. It does explain why the trade-offs belonged to his territory.

The public record / selected stops
  1. Deformable Convolutional Networks appears at ICCV; Zhang begins doctoral study at Toronto.
  2. His team publishes a NeurIPS study of batch sizes and a noisy quadratic model.
  3. He completes his PhD and joins xAI’s founding team.
  4. He works across coding and Imagine, then confirms his departure in March.

The short farewell

On X in March 2026, Zhang wrote that it was his last day at xAI. He called the preceding three years a “wild journey” and said he was excited about the next chapter. He thanked people for their support and wrote that he would miss the friends he had made. It is a modest note for the end of a conspicuous tenure. There were also reports of a wider reshaping at xAI and other founder departures. Zhang’s own words confirm the fact of his exit; they do not lay out a detailed account of his reasons.

The departure changed the present tense of his biography. “xAI co-founder” remains true. “xAI leader” belongs to a particular period. His current public profile offers a less tidy answer to what comes next. On X, he has written about machine learning and the engineering behind large models. His personal site still points readers toward publications, a curriculum vitae and code. Those traces belong to a researcher as much as to a former founder.

The speed of the xAI years can obscure the slower formation that preceded them. Zhang had been thinking about optimization and uncertainty well before AI products became an everyday subject. He had co-authored work in computer vision, examined learning dynamics, taught undergraduates, and discussed what human control should mean for increasingly capable systems. The company gave those interests a new scale. It did not invent them.

A question that travels

There is a faintly comic tension in asking a machine to admit ignorance. Its most persuasive talent is often fluency. It can produce a sentence before the person reading it has decided what to ask next. Zhang’s 2023 question about awareness of knowledge gaps cuts through that performance. To be useful, a system should not merely speak; it should give people a sound basis for deciding whether to trust what it says.

That concern joins the small and large scales of his career. On one end is a carefully bounded training model, drawn with curves and assumptions. On the other is a public product built by a company with worldwide reach. Both ask for a disciplined account of what a model can do and where its confidence runs out. They ask the same of the people building it. There is no final theorem for a career, and no neat moral attached to a resignation post. There is, though, a clear line in the work Zhang chose to publish and the questions he chose to raise.

When he told Toronto students to keep a curious mind, he added that knowledge in this field can become outdated in only a few years. He had reason to know. The tools, companies and product names can change almost overnight. The harder task is to keep asking good questions after the answers that once felt settled begin to look easy.