There is a peculiar indignity in reaching the end of a page and forgetting what was on the one before it. For a reader, it would make a novel nearly impossible. For an early transformer language model, it was a design constraint. Text arrived in fixed-length pieces; when a piece ended, useful context could fall away. Zihang Dai's research career can be read, in part, as a series of attempts to make that loss less inevitable.
The name may be familiar from the founding roster of xAI. Dai joined the company at its 2023 announcement, after work at Google Brain and a doctorate at Carnegie Mellon University. But the earlier papers are where the story gets its shape. They concern memory, the order in which a model learns words, and the cost of making a model read more carefully. Those problems sound technical because they are. They are also recognizably human: remembering a passage, considering what comes before and after a sentence, and deciding which details deserve attention.
A page break inside the machine
Dai's 2020 Carnegie Mellon dissertation collects several ideas under the title Improving Deep Generative Modeling with Practical Applications. One of them became Transformer-XL, developed with Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V. Le, and Ruslan Salakhutdinov. Dai and Zhilin Yang were marked as equal contributors. Their question was pointed: how could a transformer learn from a longer stretch of text without losing the order of events when one segment gave way to another?
Transformer-XL's answer was a form of recurrence. A model processed one segment, kept information from it, then reused that information while reading the next. The paper paired that idea with a way of representing relative positions, so words could retain a sense of where they stood in relation to one another. In plain language, the model could turn a page without treating the next page as the first one it had ever seen.
The published results gave the idea weight. On the paper's tests, Transformer-XL learned dependencies longer than those captured by the recurrent and vanilla transformer baselines the authors compared it with. The researchers also reported faster evaluation than a vanilla transformer in the long-context setting. Those are results within particular benchmarks, not a promise that every future model would behave the same way. Still, the direction was clear: the boundary between chunks of text did not have to be a hard wall.
Learning a sentence from more than one direction
The next major paper changed the training question. XLNet, released in 2019, was written by Zhilin Yang, Dai, Yiming Yang, Carbonell, Salakhutdinov, and Le, with Yang and Dai again listed as equal contributors. The model explored permutation language modeling. Rather than always learning to predict words in one fixed left-to-right order, it learned across different orders of factorization. The purpose was to bring information from both sides of a word into the training process while retaining an autoregressive objective.
XLNet also drew on Transformer-XL as its backbone. The work thus moved from one question about memory to a second question about how to learn from the words memory could hold. The team's public code repository said XLNet outperformed BERT on 20 downstream tasks and set the leading result on 18 at its June 2019 release. Such rankings are snapshots of specified tests at specified times. Their lasting interest is in the method: the team tried to change the learning objective as well as the architecture that supported it.
Dai's GitHub account hosted the XLNet code and model-release instructions. That detail matters. A paper can make a claim; released code gives other researchers a route to inspect, reproduce, and adapt the work. The repository included training and evaluation scripts alongside the published model details. It was a research result delivered as working material, not merely a chart of scores.

His dissertation gives those projects a common address. Carnegie Mellon's Language Technologies Institute lists Dai as a 2020 Ph.D. graduate in Language and Information Technologies, advised by Yiming Yang. It records his employer upon graduation as Google. His own older site identifies him as a research scientist at Google Brain, a useful timestamp on an earlier stage of the career even though that page has not been updated to reflect the later xAI chapter.
Efficiency becomes part of the plot
After the question of how much a model could remember came another: how much computation should it spend carrying the representation of a sequence forward? Funnel-Transformer, a 2020 paper by Dai, Guokun Lai, Yiming Yang, and Quoc V. Le, proposed compressing the hidden-state sequence. The design traded some sequence resolution for room to make a model deeper or wider under a similar computation budget. The image in the name was apt: the representation narrowed as it passed through the model.
There was no single grand trick across these papers. Transformer-XL concerned context over time. XLNet reconsidered how a language model learns from a sentence. Funnel-Transformer looked for redundancy in the work a model performs. Taken together, they suggest a research program with a clear recurring pressure: a model should extract more from text without treating memory or compute as infinite. The papers were collaborative, and the author lists make that visible. Dai's name appears alongside colleagues who kept working on related problems across institutions and methods.
That small public remark offers a rare glimpse of his own phrasing outside a paper. He wrote that he had decided to learn more about product and suggested that real, painful, and specific were useful words for the problem he wanted to understand. It is a practical list. For someone whose published work spent years identifying exact limits in language models, the vocabulary has a recognizable continuity. It names the kind of difficulty that can be described clearly enough to work on.
A new company, then a departure
In July 2023, xAI announced its founding team. Dai was listed with researchers from several major AI labs. His path to that roster had run through Carnegie Mellon and Google Brain, and his published work already covered several strands of model design. The new company described an ambition to understand the nature of the universe. For a scientist whose papers had focused on precise problems inside language models, it was a striking change in scale of stated mission.
The public record offers little basis for assigning Dai sole ownership of any particular Grok release, and company products emerge from large teams. It does show his presence on xAI's original roster. A later public post found him looking forward to seeing friends at a NeurIPS gathering after time away from conferences. It was a brief note, but a revealing one: even in an industry increasingly organized around companies and product launches, research still has a social geography of conferences, collaborators, and old friends.
Dai left xAI in March 2026, during a period when several members of its founding team departed. His exit closed that chapter without closing the earlier paper trail. The sequence of research, company work, and departure resists a tidy moral. What it offers instead is a way to see how an AI career can move between the laboratory and the company while carrying the same technical questions along.
- Transformer-XL and XLNet put long context and language pretraining at the center of Dai's published work.
- He completed his Carnegie Mellon doctorate and moved into research at Google.
- He appeared on xAI's original founding-team roster.
- He departed xAI in March.
What remains on the page
The papers are especially useful now because they preserve questions in full. In a product announcement, the machinery may be reduced to a model name and a performance claim. In Dai's research, the machinery is the story. Where does context go when a segment ends? Can training order change what a model learns? Which parts of a sequence need the most computation? Each question is narrow enough to test, yet broad enough to influence how later systems are built.
A short profile cannot turn collaborative research into a solo legend, and there is no need to. The names beside Dai's recur for a reason. His work was part of a community that moved between universities and industry labs, published code, compared results, and revisited old assumptions. The fascination is in that shared process: a researcher spots the point where a system forgets, wastes effort, or reads too rigidly, then helps devise a more capable way through it.
The title on a company roster may change. A carefully stated problem tends to last longer. Dai's most public contribution began with a page break that an algorithm could not cross gracefully. The paper that followed offered it a memory. From that small obstruction came a line of work that reached well beyond the page.