Long before the public learned to ask a machine for a sonnet, a recipe or a suspiciously cheerful email, Durk Kingma was working on a less photogenic problem: how such machines learn. The answer involved probability, calculus and a talent for turning difficult theory into methods other researchers could actually use. One result was the variational autoencoder. Another was Adam, the optimization algorithm that became a familiar setting in deep-learning code. Neither has the glamour of a talking product. Both helped make the products possible.
Kingma is an unusual figure for an age that converts technical work into personal mythology. His own website is spare enough to make understatement look baroque. It lists papers, a brief biography, a portrait, links and honors. At the end comes a dry little farewell: “The end. Want more? Click one of the blue links.” A career that touches the founding of OpenAI, six years at Google Brain and DeepMind, and current work at Anthropic is presented with the fuss of a train timetable.
That restraint directs attention to the right place. Kingma’s influence sits inside the machinery. Researchers recognize his formal name, Diederik P. Kingma, from the top of papers. Friends and colleagues know him by a Frisian nickname, Durk, pronounced like Dirk. The double identity is apt: Diederik belongs on a citation; Durk belongs in a lab. The work travels comfortably between both worlds.
Teaching a network to handle uncertainty
In December 2013, Kingma and his University of Amsterdam adviser, Max Welling, submitted a paper called “Auto-Encoding Variational Bayes.” Its opening question was formidable: how can a system learn efficiently when important hidden variables are difficult to infer and the dataset is large? Their answer joined variational inference, a tool from probabilistic modeling, with neural networks and ordinary gradient-based training.
The practical intuition is gentler than the notation. An ordinary autoencoder learns to compress data and reconstruct it. A variational autoencoder learns a structured probability space in between. Similar examples can occupy nearby regions; new examples can be sampled from that space. The model does not merely file a picture in a tiny drawer. It learns something about the landscape of possible pictures.
A VAE, with the algebra politely removed
example
space
or new sample
The key move was making the hidden probability space trainable with standard stochastic-gradient methods at useful scale.
In 2024, ICLR gave the paper its inaugural Test of Time award, praising the elegance of the method and the way it brought scalable probabilistic inference into deep learning. The honor arrived a decade after publication, which is precisely the point. Research fashions sprint. Infrastructure stays for the cleaning bill.
The optimizer in the engine room
A year after the VAE paper, Kingma and Jimmy Ba introduced Adam. Training a neural network means repeatedly adjusting an enormous collection of parameters to reduce error. The gradients that guide those adjustments can be noisy, sparse or liable to change character as training proceeds. Adam keeps adaptive estimates of their first and second moments, then uses those estimates to set parameter-specific learning rates.
The paper emphasized virtues that sound almost domestic: straightforward implementation, modest memory needs, computational efficiency and little tuning. This was not an algorithm asking to be admired from behind velvet rope. It wanted to work. It proved especially suited to large datasets and models with many parameters, exactly where the field was heading.
“It's kind of surreal have received the ICLR ‘Test of Time’ award two years in a row, this time for the Adam optimizer, but I'll take it!”Durk Kingma, after the 2025 award
ICLR’s 2025 award announcement called Adam one of deep learning’s most widely adopted optimization algorithms. Kingma’s reaction was more personal: gratitude to the conference organizers, pleasure at sharing the award with Ba, and a parenthetical vote for Singapore as a future venue. It was the response of somebody amused by the scale of the afterlife. A method described in 2014 had become so ordinary that countless practitioners selected it from a menu.
From Amsterdam to the institutions of the AI boom
The foundations came first. Kingma worked as a junior research scientist in Yann LeCun’s lab at New York University in 2009 and again in 2012. Between those periods he co-founded Advanza and served as its technical lead; the company was acquired in 2016. The startup is an easily missed detail, but it places practical building beside academic research from the beginning.
He began his doctorate at the University of Amsterdam in 2013 under Welling, with Joris Mooij as assistant supervisor. Summers in 2014 and 2015 took him to DeepMind for collaborations. In 2015, Google awarded him its first European Doctoral Fellowship in Deep Learning. Two years later he defended “Variational Inference and Deep Learning: A New Synthesis” with the university’s cum laude distinction. According to his biography, it was the computer science department’s first such result in three decades.
The surviving record of those doctoral years is charmingly workshop-sized. There are slides from talks in Beijing, Banff, Montreal, Sheffield, Cambridge and New York, plus a small demonstration that generates handwritten digits with a VAE. The itinerary shows an idea being tested in public before it became a textbook fixture. The digit demo is modest by current standards, but that is exactly why it matters. Today’s polished generators descend from a period when making a blurry new numeral was evidence that the underlying mathematics had begun to behave.
The VAE and Adam papers arrive during his Amsterdam doctorate.
He joins OpenAI's founding team and leads its algorithms group.
At Google Brain and DeepMind, he leads generative-model projects across text, images and video.
He researches large-scale machine learning at Anthropic, primarily from the Netherlands.
OpenAI’s original announcement named Kingma among the founding researchers in December 2015. The organization was then a nonprofit research company promising broad publication and collaboration. Kingma led its algorithms team, focused on basic research. His work there included Glow, developed with Prafulla Dhariwal. The 2018 model used reversible transformations and an invertible 1×1 convolution to generate high-resolution images, perform exact inference and expose a useful latent space. Its online demonstration let visitors interpolate between faces and alter attributes. Abstract machinery had learned a party trick.
Kingma moved to Google Brain in 2018, later part of Google DeepMind. His account of those six years is concise: he led projects on generative models for text, image and video. The publication trail fills in the shape, including score-based work and Variational Diffusion Models. The recurring theme was not loyalty to one architecture. It was the search for scalable ways to model how data is generated.
The thesis accumulated honors as the methods spread. In 2019, Kingma received an ELLIS PhD Award and the Gerrit van Dijk Prize at the Dutch Data Science Awards. The prize jury noted that work from the dissertation had already drawn more than 25,000 citations, an extraordinary total so early in a research career. Numbers of that sort can become gaudy. The more interesting fact is what people were citing: a set of techniques they could put to work.
A new lab and an old concern: scale
In October 2024, Kingma announced that he had joined Anthropic. He would work mostly from the Netherlands, he said, while visiting the San Francisco Bay Area as frequently as he could. The arrangement fits a career that has been geographically Dutch and intellectually international: Amsterdam as a base, collaborations and institutions spread across London, New York and California.
“Anthropic's approach to AI development resonates significantly with my own beliefs.”Durk Kingma, October 2024
He described his aim as contributing to the responsible and safe development of transformative AI. It is a carefully bounded public statement, not a manifesto, and it points toward a tension now familiar to the field. The systems have become more powerful, the training runs more expensive and the consequences less confined to research benchmarks. Scalability is no longer only an engineering virtue. It comes with obligations.
In January 2026, Kingma returned to a public lectern in the Netherlands for a talk titled “Generative Models: Past, Present, Future.” The abstract promised a review of techniques developed across his university and industry years, then a discussion of LLM productivity tools and the capabilities likely in the short and medium term. The title could describe his career as neatly as the talk. Past: the methods. Present: the scaled systems. Future: what those systems become in use.
There is a temptation to make every AI career a story about destiny. Kingma’s is better understood as a story about compounding. A reparameterization makes probabilistic models easier to train. An optimizer handles noisy gradients with less fuss. A reversible model opens another route through images. Each contribution solves a technical problem; together they alter what later teams can attempt.
That may be the cleanest definition of quiet leverage. The public sees the answer appearing in a chat window. The engineer sees a training curve. Somewhere underneath, an old paper is still doing its shift.