The chart was drawn in crayon. That is the part Noam Shazeer remembered a quarter century later: a rising line on a Google office wall, marking the daily number of searches. He had seen the company at a job fair in 1999 and assumed it was already too large to be interesting. Everyone he knew used it. By the next year he had sent a résumé anyway. Inside, there were people he liked, problems worth having, and this rather persuasive little picture of growth.
He took the job. Google assigned him a mentor named Jeff Dean, who seemed to know the answer to every question. Shazeer later joked that he had mistaken Dean's command of the place for a universal quality of Google employees. The joke is funny because it comes from someone whose own work would eventually fill a substantial corner of the company's technical history. His first contributions were practical and invisible to most users: spelling suggestions in Search, then the PHIL algorithm that became central to early AdSense. A person could encounter his engineering several times in a day without ever hearing his name.
The name became much harder to avoid after 2017. Shazeer was one of eight authors of Attention Is All You Need, the paper that introduced the transformer. That architecture became a central building block for contemporary language models. Yet the tidy origin story obscures a longer, odder route: a perfect-score math contestant, a search engineer, a researcher who kept making huge models more practical, a chatbot founder, a Gemini leader, and, in 2026, an OpenAI architecture researcher. The line on that crayon chart kept climbing. His career took several turns beside it.
Six problems, six full marks
In 1994, Shazeer represented the United States at the International Mathematical Olympiad. Every member of that six-person team earned the maximum 42 points; together they posted a perfect team score of 252. Shazeer got seven points on each of the six problems. This is a lovely fact, partly because it resists embellishment. There is no need to call the test impossible. The score is enough.
1994 / Math olympiadShazeer earned full marks on all six problems. The US team did the same, finishing with a perfect 252.
At Duke University, he continued to compete. A 1998 mathematics department account lists him 20th among 2,510 contestants in the previous year's Putnam competition, a six-hour exam notorious for giving even strong students a rough afternoon. Mathematics was a public early talent; his later account of Google suggests that artificial intelligence was already a private ambition. In a 2025 interview, he recalled imagining that he might earn enough from a technology job to spend a long time working on AI. His punch line, delivered after a career in Google's AI labs, was that the plan had worked out exactly as intended.
The early Google work mattered for another reason. Search spelling correction asks a machine to make a useful guess from messy language, at enormous scale, quickly enough that nobody notices the wait. Advertising systems demand the same indifference to slow or fragile code. Neither project looks like a modern chatbot, but both taught the same stubborn lesson: an idea becomes consequential when it runs reliably in the hands of strangers.
The useful trick inside the machine
Around 2015, Shazeer became absorbed by language modeling. His 2025 Hot Chips presentation described text as distilled abstract thought and next-token prediction as an unusually simple way to frame a very large problem. The attraction was partly scientific, partly logistical. There was plenty of text to learn from. The computers were improving. Models could be tested against a clear task: given what came before, what comes next?
One obstacle was cost. Bigger neural networks could hold more patterns, but asking every part of a gigantic network to work on every input was expensive. In the 2017 paper on sparsely gated mixture of experts, Shazeer and colleagues proposed a way to call on selected pieces of a larger model when needed. Imagine a busy newspaper with many specialist desks but only the relevant desks summoned for each story. The analogy is imperfect; the computing idea is precise. More model capacity could be available without using every parameter for every step.
The transformer paper offered a different advance. Earlier sequence models commonly processed words in order or relied on convolutions. The team's model used attention to connect parts of a sequence more directly and allowed much of the work to run in parallel. On the translation tasks in the paper, it achieved strong results with less training time. The architecture was a collaboration, and credit belongs to all eight authors: Ashish Vaswani, Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan Gomez, Łukasz Kaiser, and Illia Polosukhin.
Shazeer's subsequent work kept addressing bottlenecks that show up after a fine idea meets a big computer. Mesh TensorFlow gave researchers a way to divide huge calculations among processors. T5 explored a unified text-to-text approach to language tasks. Multi-query attention reduced the memory bandwidth needed while a model generates text. Each sounds like plumbing until a system begins to stall. Then the plumbing is the story.
“Mostly these things fail, but if you try 100 things or 1,000 things or a million things, then you might hit on something amazing.”Noam Shazeer, 2025 interview
He has spoken with cheerful severity about experimental failure. The point of his two-percent estimate was to argue for trying many ideas, not to crown any single one as inevitable. The familiar image of a lone inventor having a perfect thought does little justice to this kind of work. Models advance through experiments, failed runs, revised architectures, and colleagues who notice a better path.
Someone to talk to
By 2021, the research had become a product question. Shazeer and Daniel de Freitas left Google and founded Character.AI. They had worked around conversational systems at Google; de Freitas led work on Meena and LaMDA, and Shazeer contributed to LaMDA. Their new company would let users create and speak with AI characters. A chatbot could be a tutor, a fictional companion, or an invented personality. The point was to let people shape the other side of the conversation.
Character.AI's public beta launched in September 2022. In a December introduction signed by Shazeer, de Freitas, and the team, the company said the beta generated a billion words a day two months after launch and that users had made more than 350,000 characters. Those numbers capture a change in scale, but the product's more revealing feature was ordinary: people kept coming back to talk. The interface gave a technical research program a social life.

The range of characters made the product difficult to describe with one clean noun. Some conversations were practical; others were playful performances. In a TIME profile, a reporter interviewed an AI version of Shazeer before meeting the real one. The imitation gave an orderly explanation of why its subject had founded a startup. The actual Shazeer offered a similar reason: he wanted to get the technology to billions of people, and a startup could move faster. It is an amusing scene, though it also says something serious about the company. Character.AI made the gap between imitation and person part of the daily user experience.
Shazeer did not present the business as a mere demonstration of clever mimicry. In the founders' introduction, he and de Freitas emphasized education, assistance, mentorship, creativity, and social connection. Ambitions that broad can sound like a pitch deck, but they also explain why the team chose open-ended conversation as the interface. A user can ask, disagree, change direction, or invent a character the founders never imagined. The product grows out of that freedom.
The return ticket, and the next one
In August 2024, Google entered a nonexclusive agreement to license Character.AI technology. Shazeer and de Freitas returned to Google along with other researchers, while Character.AI continued as its own company. Shazeer became a co-technical lead on Gemini, Google's flagship model effort. It was a striking loop: the engineer who had left to put conversational AI in front of the public came back to help build the models that would power Google's own offerings.
There is a tempting boardroom version of that episode, with large numbers and a tidy winner. The more interesting version lives with the people and the work. Shazeer's conversations with Jeff Dean in 2025 still sounded like two engineers thinking aloud about compute, data movement, experimentation, and the surprisingly hard task of making models useful. Their first relationship at Google had been mentor and new hire. A quarter century later, they were discussing Gemini together in public.
The same year, Shazeer stood before a hardware audience at Hot Chips and opened his keynote with a simple thank-you to the people building supercomputers. It was half joke, half accounting. Large models ask an enormous amount of the machines beneath them. His slides moved from the 2016 mixture-of-experts work through the transformer and Mesh TensorFlow, then returned to memory, bandwidth, and processor networks. The technical interests had survived the startup detour intact.
In 2026, the National Academy of Engineering elected Shazeer for contributions including transformer and mixture-of-experts architectures. That June, he announced he was leaving Google for OpenAI. OpenAI research leadership said he would lead architecture research there. The move placed him inside another organization trying to answer familiar questions about model design, scale, and usefulness. It also gave the career one more turn that would have looked improbable from that 2000 crayon chart.
A conversation still in progress
Shazeer's public manner tends toward engineering candor, dry humor, and large horizons. He has talked about a million automated researchers and about the uncomfortable odds of any proposed machine-learning idea succeeding. These remarks sit together more comfortably than they first appear. If experiments fail often, the way to learn more is to make better systems for trying them. If conversational machines are going to serve many people, the architecture behind them must survive contact with scale.
His story is often reduced to a few famous nouns: transformer, Character.AI, Gemini. The quieter verbs are more useful. He noticed the chart. He built spelling correction. He collaborated on architectures, made them run across machines, founded a place for people to use them, went back to build more, and moved again. From a perfect score in 1994 to architecture research in 2026, his route has followed the same curiosity through very different rooms. The question on the other side of the screen remains open: what, exactly, should a machine be able to say back?