In a USC office, long before a paper with a mischievously confident title changed the vocabulary of artificial intelligence, Ashish Vaswani built himself a GPU workstation. A professor remembered the machine years later because it marked a small act of conviction. Deep learning was still an uncertain bet in natural language processing. Vaswani wanted to see what its machinery could do. A workstation assembled by hand is a rather physical way of asking a research question.
The question eventually grew much larger. In 2017, Vaswani and seven colleagues at Google published “Attention Is All You Need.” Their transformer architecture offered a way to process relationships among words without moving through a sentence strictly one word at a time. Its first public proof was machine translation. Its descendants would turn up across language models and other forms of AI. The history of the transformer is often told through the technology it enabled. Vaswani’s story is more interesting when it begins with the people who made it, and with the habits that carried him from a university lab to Google Brain, two startups and, in 2026, NVIDIA.
A translation problem with a long shadow
Vaswani was born in India and raised there and in the Middle East. He studied computer science engineering at Birla Institute of Technology, Mesra, then went to the University of Southern California for graduate work. At USC, machine translation gave his research a useful constraint: a system has to preserve meaning while turning one sequence of words into another. The task is unforgiving about context. A word at the start of a sentence may depend on one near its end, and the same word can mean something different when its neighbors change.
He worked in USC’s Information Sciences Institute on neural language models, including research that improved translation and methods for training large vocabularies more efficiently. In 2012 he went to Papua New Guinea on a project using language technology to help document endangered languages. That trip places his work in a wider frame. Language was a research problem with people and places attached, not merely a spreadsheet of benchmark scores.

USC also supplied a collaborator who would recur throughout his career. Vaswani gave a guest lecture on neural networks and met Niki Parmar, then a graduate student. They became friends and research partners. Both later worked at Google, appeared among the eight names on the transformer paper, and went on to found Adept and Essential AI together. It is tempting to make a grand point about fate. The smaller observation is enough: a lecture can be a career event for more than the person at the podium.
A former adviser, David Chiang, recalled Vaswani’s early interest in deep learning. Another, Liang Huang, remembered the homemade GPU workstation. Their memories point to a researcher drawn toward an approach before it became the obvious approach. Vaswani later said USC gave him a culture that prized practical results and clear communication, as well as room to pursue bold ideas. He earned his doctorate in 2014 and joined Google Brain in 2016.
Eight authors, one urgent deadline
The transformer did not arrive as a solitary flash. At Google, Jakob Uszkoreit and Illia Polosukhin had been exploring self-attention. Vaswani joined their discussion; other researchers joined as the idea gathered force. The eventual paper listed Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan Gomez, Łukasz Kaiser and Illia Polosukhin. Each brought work that made the design function. The eight-author list is a useful corrective to the tidy myth of the lone inventor.
Their immediate goal was to translate text. Older leading systems relied on recurrence or convolution. The team proposed a network based on attention alone. In plain terms, attention allows a model to weigh how parts of an input relate to one another. A sentence can be considered through many relationships at once, which also lets training make good use of parallel computing. The idea sounds neat when written down. Getting it to beat strong alternatives demanded months of experiments, model revisions and the sort of deadline that makes an office couch look inviting.
Simplified from the paper’s architecture and experiments.
The sprint had its own little comedy. The researchers gathered in the building with the better espresso machine. Vaswani recalled sleeping on a couch while the paper was being written. Looking at curtains patterned like neurons, he told Gomez that the work might reach beyond translation to other kinds of information. Gomez remembered an all-nighter with Vaswani before submission. Research history can sound immaculate in a conference proceeding; the coffee and the couch restore its human dimensions.
The results were concrete. The larger transformer model trained for three and a half days on eight GPUs and achieved a then-leading English to French translation score. Its English to German result improved on previously reported systems. The model was easier to parallelize than recurrent designs. Those facts mattered to researchers deciding whether the architecture was worth trying in their own work. The paper’s title may have invited a smile, but the numbers gave colleagues a reason to open the code.
Even the name had a draft. Vaswani recalled that “Attention Net” lacked much appeal. “Transformer” won. There is a pleasant mismatch between the name’s cinematic energy and the team’s first application, a translation benchmark. By the following years, the word had become part of the naming convention for systems far outside the original experiment. The researchers had designed a useful way to handle relationships; the field supplied more relationships for it to handle.

When an architecture becomes a company
After Google, Vaswani co-founded Adept AI Labs with Parmar and other colleagues. The venture pursued AI systems able to act across software tools, a move from building a general capability toward making it useful at a desk. Vaswani served as chief scientist. He and Parmar later left and founded Essential AI in 2023. Its early public description focused on a deeper partnership between people and computers, with models that could help automate repetitive enterprise work.
Essential AI announced a $56.5 million Series A that December, with backing that included AMD, Google and NVIDIA. Funding is a tool, not a biography, but the investor list shows how many large players wanted to stay close to the work of two transformer co-authors. The company was based in San Francisco. Vaswani became its chief executive, a role that asks a researcher to keep a laboratory’s questions alive while dealing with the payroll, customers and hardware bills that make those questions possible.
By 2025, his public argument had become sharper. Vaswani warned that an AI race can reward the scaling of familiar ideas and leave less room for riskier research. He argued for open work so that more people could test, improve and control frontier models. The view was rooted in a memory of the period that produced the transformer: several labs could share consequential ideas, and advances accumulated across groups. He did not describe progress as a single lightning strike. It was a chain of contributions, each making the next one possible.
Essential AI gave that position a tangible form in December 2025 with Rnj-1, an eight-billion-parameter model released with base and instruction-tuned versions. The name honored Srinivasa Ramanujan. Vaswani’s accompanying essay explained the model, its focus on coding and technical tasks, and the team behind it. The release made his case for openness something other researchers could inspect and use. It also showed a familiar pattern in his career: choose a demanding technical problem, then put a working artifact in front of people.
The next laboratory
In June 2026, reporting said NVIDIA had hired Vaswani and other Essential AI researchers for work connected to its open Nemotron models. His professional profile now lists NVIDIA in San Francisco. The move has an obvious historical symmetry. The original transformer paper measured training on NVIDIA GPUs; its first author now works at the company that makes them. Symmetry can be too convenient a storyteller’s device, but here it is also a practical fact about where model research and computing infrastructure meet.
The transition leaves open questions about Essential AI’s future and how Vaswani will carry his argument for shared research into a larger organization. Public evidence establishes the move and his earlier position; it does not supply a script for what comes next. What can be seen is the continuity of subject matter. Translation, attention, agents and open models all ask how a machine can use context to do something useful. The institutions changed as the tools and stakes changed.
That continuity gives the title of the 2017 paper an amusing afterlife. “Attention Is All You Need” was a deliberately bold claim about a network architecture. For Vaswani, attention has also meant attending to collaborators, to the conditions that let ideas circulate, and to the neglected paths a race might leave behind. He was once the student building a workstation to test an unfashionable approach. He is now a researcher inside one of the companies supplying the machinery for the next generation of models. The distance between those scenes is considerable. The question that connects them is simple: what becomes possible when someone is allowed to try the idea that does not yet look inevitable?
Explore his work
- ResearchRead “Attention Is All You Need”
- Research blogVaswani on Rnj-1 and open models
- InterviewHis 2025 conversation about AI research
- VideoIndia Science Festival fireside chat, 2026
- ProfilesLinkedIn · X