A large neural network can make a small question very expensive. Which learning rate should it use? How should its weights be initialized? Answering by trial and error means running a costly training job again, perhaps several times. Greg Yang’s answer began with a mathematical proposition: under the right conditions, a smaller network can tell you something useful about a much bigger one. That idea has shaped a research career whose vocabulary sounds remote from ordinary life - infinite width, tensor programs, maximal updates - but whose practical promise is plain. Make the small experiment count.
Yang came to this problem by a winding route. He was born in China’s Hunan province, moved as a child to Guangzhou and Beijing, then to Houston before studying at Harvard. For a year and a half, he left university to produce electronic dance music and work as a DJ under the name Zeta. A mathematician’s prize biography records that this period exposed him to the ideas of artificial intelligence. There is something fitting about that origin. A DJ works with patterns and timing; Yang would eventually spend years asking how patterns change when a network gets larger.
A detour with a beat
At Harvard, his interests were already unusually wide. His senior thesis developed a homological theory of functions, connecting computational complexity, learning theory and algebraic topology. His master’s thesis described a mathematical theory of neural memory and algorithmic learning based on Lie groups. The senior thesis won a Thomas Temple Hoopes Prize, and his undergraduate research earned an honorable mention for the 2018 Frank and Brennie Morgan Prize. These were not prizes for building a product. They recognized a habit that would become useful in AI: choosing a difficult object, inventing a language for it, and then seeing what the language can prove.
After Harvard, Yang worked at Microsoft Research. The field around him was making neural networks larger and more capable, while many of the rules for training them remained empirical. An engineer could tune one model and still have to ask whether the same settings would work when the model grew. The common response was another experiment. Yang wanted a theory that would tell researchers when the answer should carry across sizes.

He called the framework Tensor Programs. He gives an approachable version: a tensor program is a composition of matrix multiplication and coordinatewise nonlinear operations. Those are the basic moves behind many neural-network computations. A formal language for those moves lets him ask what happens as a network’s width tends toward infinity. Infinity here is a mathematical tool, an idealized limit from which finite networks may become easier to understand. It is not a claim that anyone can train a model with infinitely many parts.
The series began with the mechanics of wide networks and moved through Gaussian processes, neural tangent kernels, matrix laws, feature learning and optimization. Its papers are difficult, and Yang says so with unusual frankness. He calls the original paper dense and its notation as cumbersome, explaining that several successors were written partly to present the ideas more accessibly. It is a charming admission from the author of a formidable mathematical program: even a master theorem needs good directions.
The small rehearsal
The practical payoff arrived through a particular way of setting up a neural network called maximal update parameterization, usually shortened to μP. In ordinary language, it is a recipe for how certain aspects of a model should scale when its width changes. In this setup, Yang and his collaborators found that useful training settings can stay stable across model sizes. They named the procedure μTransfer: tune the smaller model, then transfer the settings to the larger one without tuning that larger model directly.
The 2022 paper did not leave the idea as a neat line on a chalkboard. Its authors tested it on Transformer and ResNet models. In one experiment, settings tuned on a 40-million-parameter proxy were transferred to a 6.7-billion-parameter language model. The team reported results better than published numbers for a comparable GPT-3 model while spending the equivalent of 7 percent of the large model’s pretraining compute on tuning. In another experiment, a 13-million-parameter proxy guided a 350-million-parameter BERT-scale model. These were specific experiments, with specific model setups. Their force came from showing a mathematical expectation survive a demanding practical test.
The attraction is easy to see. A lab can afford many small experiments and only a few enormous ones. If the small experiments reliably identify useful settings for the large run, research becomes less dependent on expensive guessing. The result also gives theory a practical burden: it must predict something an engineer can use. Yang’s work occupies precisely that meeting place, between a limiting equation and a training schedule on real hardware.
An appetite for foundations
Yang’s ambition stretches beyond cheaper tuning. He wants a theory that tells researchers how to scale large networks and gives them a stronger understanding of the models, including questions of safety and alignment. He has kept extending Tensor Programs toward that aim. A later paper studies adaptive optimizers such as Adam. Another examines networks that grow in depth as well as width, introducing depthwise hyperparameter transfer for some residual-network designs while identifying limits in others. The pattern is revealing: new claims sit beside conditions under which they do, or do not, apply.
His path also connected him to a wide group of researchers. The μTransfer paper lists collaborators from Microsoft and OpenAI, including Edward J. Hu, Igor Babuschkin, Jakub Pachocki and others. Its open-source implementation lives in Microsoft’s μP repository. His GitHub page points readers to the code, and he offers a reading guide for the papers. He suggests starting with whichever result draws your interest and moving among papers when the mathematics gets difficult. It is advice from someone who knows what it feels like to have a theory outrun its explanation.
In 2023, Yang joined the founding team of xAI. The company’s early public materials placed him among researchers from Microsoft, OpenAI, DeepMind and Google. xAI soon released Grok, its conversational AI product. Yang’s public record is strongest on his mathematical research; it does not assign him a neat, individually measured share of a product built by a team. His move to xAI is still significant. It brought a researcher intent on a general theory of deep learning into a lab operating at the scale where those questions become costly and consequential.
The fit is not hard to understand. If an organization plans to build larger models, it has reason to care about the cost of getting the setup wrong. If it wants the systems to be more understandable, it has reason to care about theory that connects one architecture to another. Yang’s framework aims at both concerns. It does not make the uncertainty disappear; mathematics rarely offers so convenient a service. It provides conditions, predictions and a more disciplined way to ask the next question.
The work also invites a useful distinction between a successful transfer and a universal shortcut. The paper’s examples depend on how the model is parameterized and which settings are being transferred. Yang’s own explanation stresses that the favorable behavior belongs to μP, not automatically to every default setup. The qualification matters. A claim that small models always predict large ones would be easy to remember and hard to trust. A claim with explicit assumptions can be tested, improved, or rejected. That is the mathematician’s discipline inside an industry often tempted to turn each promising graph into a slogan.
Even the title Tensor Programs has a modest utility beneath its imposing sound. It names a shared form for many computations, so each new architecture does not have to be treated as an entirely separate mystery. The language is abstract because it must cover a lot of ground. Its value is measured in concrete consequences: a prediction about feature learning, a rule for initialization, a training setting that still works when a network grows. That repeated passage from general rule to particular test is the central rhythm of Yang’s research.
The question he kept
In January 2026, Yang said he would step back from an operating role at xAI and continue as an informal advisor. The title changed, but the arc of his published work remains clear. The student who left Harvard for music, returned to finish two degrees, then pursued a mathematical account of vast neural networks has kept returning to the same puzzle: can the behavior of a giant system be understood before the giant experiment is run?
There is a quiet wit in the scale of the answer. The network may have billions of parameters; the useful clue may come from a model small enough to test repeatedly. Tensor Programs supplies a language for the leap between them. μTransfer tries to make that leap operational. For Yang, the ideal outcome is larger still: a theory that helps people choose how to build these systems and understand what they have built. He has described that ambition in grand terms. The work, paper by paper, proceeds through smaller and more exact ones.