Jay Alammar’s first audience for his machine learning writing was Jay Alammar. He had a computer science background and had worked as a software engineer, yet the new deep learning literature offered plenty of ways to mistake recognition for understanding. A tutorial could feel clear while it was open in a browser. Three months later, he wanted something firmer than the sensation of having followed along. So he began writing explanations and posting them online. The prospect of being wrong in public sent him back to the papers, then back to the code. A reader came later. The first job was to find out what he himself knew.
That origin explains the peculiar usefulness of his work. Alammar does not simply define a term and move on. He opens a mechanism, takes out a piece, and draws how it relates to the next piece. The method gave countless readers a path into language models at the moment those models were becoming central to artificial intelligence. It also made his name familiar in a field where the research arrives faster than anyone can sensibly digest it. Good technical writing is often described as translation. His is closer to disassembly, with the screws arranged neatly on the table.
The learner at the drawing board
In an interview about his teaching, Alammar traced his entry into machine learning to the period when TensorFlow made the field feel newly accessible to him. He read tutorials, followed blog posts, and wrote his own accounts to test what had stuck. Early pieces went to Reddit. He found that publishing one or two made him more comfortable with a topic, partly because a public explanation demanded more care than private notes. It forced a narrative: where does the reader start, what changes, and why should the next step follow?
The visual element developed alongside that habit. Some early work used interactive demonstrations; later, graphical explanations proved especially effective. Each drawing gave him a chance to catch a flaw before a reader did. He could see an assumption sitting awkwardly between two boxes, or an arrow making a claim the prose had not earned. He has described revising figures for The Illustrated Transformer roughly seven times apiece. This is the secret inconvenience of a simple picture: the simplicity has to be made.
“Writing as the best form of learning yourself and sort of deepening your understanding.”Jay Alammar, on how the blog began
The impulse to build things began earlier than the AI blog. As a child in Saudi Arabia, he was drawn to computers and science fiction. A profile from Mawhiba recounts a family exercise in enterprise: his father gave him and his brother sweets to sell outside a mosque. It was brief, but Alammar later described the discomfort of selling before an audience of children as a lesson in persistence. He eventually studied computer science at the University of Kansas. During college he taught himself web development beyond the syllabus, and a book about the rise of search engines helped him imagine a company that could begin with one person and a small idea.
In 2006, he launched Qaym, a restaurant review site in Saudi Arabia. It grew into a startup with a web application and mobile apps. He later worked in venture investing with STV, helping technology companies grow. Alongside those roles, he continued making small open source contributions when time allowed. One visual post about pandas functions found its way into that library’s official documentation. The pattern is recognizable: learn a tool, notice where the explanation could be kinder, and leave a map behind.
A transformer, one layer at a time
The decisive map appeared in June 2018. The transformer architecture had been introduced in a 2017 research paper with the thrillingly declarative title Attention Is All You Need. It promised an important way to process sequences, but the machinery was difficult to picture if you had not spent months among the equations. Alammar’s The Illustrated Transformer began with a large view and then opened it: encoder, decoder, embeddings, attention, and the path from input to output. The pictures gave readers a place to stand before the mathematics got crowded.
Its power lies in sequence. He does not ask a new reader to hold every component in mind at once. First comes the black box. Then the box has two halves. Each half acquires layers. A word becomes a vector; that vector meets other vectors through attention. The technical terms remain, because they are the terms a learner will meet in the original work. But by the time they appear, each has a job. Even a reader who cannot yet reproduce the calculations can say what question the calculation answers.
The post traveled into classrooms. Alammar’s site notes references in AI and machine learning courses at MIT and Cornell; his related BERT essay reached courses at Stanford and Carnegie Mellon. It was also revised as the subject moved on. A 2025 update points readers toward a free animated short course on how transformer language models work now. A diagram from 2018 can remain useful without pretending the field stood still in 2018.
He sought feedback from researchers connected to the transformer paper, thanking several of them in the post. That detail matters: a lucid explanation is only valuable if its lucidity survives contact with the thing it explains. Technical accuracy and approachability are sometimes treated as rivals. His best work makes them negotiate, then sends both to the printer.
Two hundred slides for a roomful of people
The same effort is visible outside the blog. For his first technical conference talk, at QCon London, Alammar prepared a presentation on word embeddings and their use in recommendation systems. He later wrote that the six weeks before the event consumed about 100 hours and produced 200 slides. That is an improbable amount of material for one talk and a perfectly plausible amount of thinking for a good one. He took a concept often introduced as a compact formula and carried it toward examples from companies such as Airbnb and Alibaba.

At Udacity, another kind of teaching apprenticeship took shape. He worked on machine learning, deep learning, and natural language processing material for the company’s programs. In a 2025 interview, he singled out fellow educator Luis Serrano. Watching Serrano think through a lesson over lunch altered how Alammar approached his own work. A switch to Apple Keynote helped too: the software made visual iterations and animation quicker. Craft can depend on a colleague’s example and the humble speed of a drawing tool.
From explaining models to building with them
At Cohere, where he is a director and engineering fellow, Alammar’s work moved closer to the people putting language models into applications. His essays for the company have covered prompts, semantic search, classification, and retrieval. He has described the appeal of managed models in plain engineering terms: developers can spend less effort loading and deploying a model and more effort on the problem they mean to solve. The diagrams still matter, but now they point toward decisions a team must make in a working product.
His open source project Ecco extends the curiosity in another direction. It provides ways to inspect and visualize the behavior of transformer language models in notebooks. Where a blog drawing shows a general mechanism, an inspection tool lets a learner poke at a particular model run. Both begin from a stubborn question: what happened between the input and the output?
Alammar is also careful about which AI questions deserve the most airtime. In a 2025 conversation, he said he was more interested in robust systems that solve real needs than in trying to settle a date for humanlike intelligence. He pointed to recommendation engines as already consequential AI systems because they shape what people read and watch. It is a practical position with a civic edge. The model in a feed may be less theatrical than the model in a headline, but it is already choosing what appears on millions of screens.
That stance fits the writing. It asks readers to examine the actual flow of information: what is retrieved, what is generated, what is measured, and what could go wrong when a system is put to use. In his teaching, a picture is a way to make those questions easier to ask. It is not a decorative pause between stretches of prose.
The books get thicker; the question stays small
In 2024, Alammar and Maarten Grootendorst published Hands-On Large Language Models. Grootendorst is known for open source tools including BERTopic; Alammar has praised his documentation for being accessible and visual. Their book moves from the workings of language models to applications and fine-tuning, with color figures built for the text. Alammar has said the writing took about a year and a half, helped by years spent developing a visual language for topics such as attention and embeddings. The work of explaining had become reusable infrastructure.
Their second collaboration, An Illustrated Guide to AI Agents, arrived as an ebook in September 2026. Its subject is wider than a single model: tools, memory, planning, evaluation, and systems in which models act through software. Alammar described more than 300 original figures in announcing the release. For a reader, the shift is significant. Understanding a transformer is a start; understanding how a system uses one while retrieving information, calling tools, and judging its own result requires a larger map.
There is an appealing continuity in that move. He began by drawing so he could check his own comprehension. Now the diagrams help readers check theirs at the scale of full applications. His site has moved newer writing to a Substack newsletter, while the older essays remain available as a record of questions that once felt forbidding and now sit near the center of software culture.
The person outside those diagrams peeks through his about page. He loves science fiction, translated Isaac Asimov’s The Last Question into Arabic, and mentions an unpublished novella with the comic severity of a writer who knows revision is hard. He lists Dune, Berserk, Final Fantasy VII, StarCraft, and Opeth among his interests. They do not form a secret key to the transformer. They do suggest a mind comfortable with elaborate worlds, provided someone takes the time to show how the pieces fit.
That is what Alammar has kept doing. The field supplies new architectures, new products, and new reasons to be confused. He returns with another box, another arrow, another carefully placed word. An explanation is successful when a reader can finally see the next move - and then, perhaps, draw one of their own.