A research lab is a splendid place to discover that a machine is wasting your afternoon. In 2009, machine-learning researchers at UC Berkeley had algorithms they wanted to run repeatedly over large stores of data. The reigning machinery, MapReduce, was reliable and influential, but its stop-start model was a poor fit for iterative work. The computers could cope. The workflow could not.
Matei Zaharia, then a Berkeley doctoral student, noticed the mismatch. He had already worked with early Hadoop users at Facebook and Yahoo, and he understood both the appeal of distributed computing and the awkwardness that appeared when people asked it to do something beyond its original brief. His response was Apache Spark, a computing engine designed to keep working data available across operations and to support more than one kind of job.
The first Spark report, published in 2010, made a crisp claim: on iterative machine-learning jobs, the system could outperform Hadoop by a factor of ten. It also showed interactive queries over a 39-gigabyte dataset returning in less than a second. Those figures caught attention, naturally. Waiting is a universal language, and ten times less of it is a fluent sales pitch even when nobody is selling anything.
“I first began Spark in 2009 to tackle the machine learning use case.”Matei Zaharia, 2015
A solution with room for strangers
Spark's more consequential design choice was breadth. It was not confined to one clever benchmark. The team shaped it into a general engine with libraries for SQL, streaming, machine learning and graph computation. A user could assemble a complete data pipeline without changing systems at every turn. The abstraction was technical; the courtesy was human.
Zaharia also understood that open-source software lives or dies by what happens after the code compiles. The early team mentored outside contributors, accepted their patches and published free training materials. In March 2012, he wrote that the community had crossed a peculiar threshold: the Berkeley group no longer knew every Spark user personally. Questions and feature requests were arriving from strangers. Better still, unfamiliar uses of the software were sending new research questions back into the lab.
The habit of making room for strangers predates Spark. At the University of Waterloo, Zaharia studied computer science within a Bachelor of Mathematics program and added a minor in combinatorics and optimization. In 2005, his Waterloo team finished fourth at the programming contest world finals in Shanghai, first among North American teams. They solved seven of ten problems and came home with gold medals. Competitive programming rewards speed, but team competition rewards a second skill too: turning private cleverness into a shared result.
His early jobs read like a tour of the places where large-scale computing was being assembled. At IBM in 2005, he worked on a tool for testing web-service performance. At Google in 2007, he added features to an internal code-search system. At Facebook the next summer, he built a fair scheduler for Hadoop MapReduce and a dashboard for monitoring the company's clusters. He later contributed that scheduling work to open-source Hadoop while contracting for Cloudera. None of these assignments was Spark, but each supplied a close look at the mundane negotiations inside a shared computer: whose job runs first, where the data sits, which slow machine holds everyone up, and how an operator can tell what is happening.
Berkeley gave those observations an academic shape. His doctoral work included delay scheduling, a technique for balancing fairness with the advantage of running a task near its data, and LATE, an approach to handling unusually slow tasks in MapReduce. He also helped start Apache Mesos, which let different kinds of applications share a cluster while retaining control of their own scheduling. Spark was therefore less a lightning bolt than a synthesis. Zaharia had spent years studying the small frictions that accumulate when many programs, machines and people must share resources without chaos.
His dissertation title was almost comically direct: An Architecture for Fast and General Data Processing on Large Clusters. “Fast” promised relief; “general” carried the larger ambition. The work received the 2014 ACM Doctoral Dissertation Award. In the same year, a team using Spark set a world record for sorting 100 terabytes of data, completing the task with one-tenth the machines previously required. A lab project had learned how to travel.
Waterloo finishes fourth worldwide at the ICPC finals.
Iterative machine learning exposes MapReduce's awkward edge.
Seven Berkeley collaborators found Databricks.
DSPy, GEPA and agent research make model behavior more programmable.
The company arrived after the community
Databricks was founded in 2013 by Zaharia and six Berkeley collaborators: Ali Ghodsi, Andy Konwinski, Ion Stoica, Patrick Wendell, Reynold Xin and Arsalan Tavakoli-Shiraji. Zaharia became chief technology officer. The order matters. Spark already existed. Its users had already begun to pull it in unforeseen directions. The company was built around a working technical community, not the hope that one might materialize later.
Zaharia did not leave research behind. After completing his Berkeley PhD under Scott Shenker and Stoica, he joined MIT's faculty in 2015, moved to Stanford in 2016 and returned to Berkeley as an associate professor in 2023. The itinerary looks busy because it was. Yet the academic and corporate halves of the career fed each other. A university is good at asking what could exist. A company is excellent at reporting, sometimes loudly, what breaks on Monday morning.
The next projects followed the same rhythm as Spark. Machine-learning teams were juggling experiments, parameters and deployments across incompatible tools, so Zaharia helped create MLflow, an open platform for tracking and managing the model lifecycle. Cloud data lakes were cheap and flexible but short on the guarantees expected from databases, so he co-developed Delta Lake, which brought transactions and dependable table management to object storage.
Spark
Keep working data available so iterative and interactive jobs do not repeatedly begin from cold storage.
MLflow
Give experiments a record, models a lifecycle and teams a common way to reproduce what worked.
Delta Lake
Add transaction guarantees and structure to inexpensive, sprawling cloud data stores.
This is Zaharia's recurring move: find a messy practice, identify the piece of structure it lacks, and make that structure available as a system. There is little romance in phrases such as “experiment tracking” or “transaction guarantees.” That is their charm. Civilization is mostly maintained by features too sensible to star in the keynote.
Now the unreliable thing can talk
Artificial intelligence has given the pattern fresh material. Large language models are fluent, flexible and maddeningly difficult to program with precision. Zaharia's recent Berkeley work includes ColBERT for retrieval, DSPy for expressing and optimizing language-model programs, LOTUS for semantic queries and GEPA for improving prompts, code and other textual parameters through reflection. The names change. The appetite for a better abstraction does not.
In 2026, Zaharia also announced Omnigent, an open-source layer for composing coding agents and custom agents. He and his collaborators are asking practical questions about context, permissions, state, evaluation and cost. Once again, a glamorous application is leaning on an ungainly systems problem. Agents may converse like colleagues, but somebody still has to decide what they may read, what they may change, how they recover and how anyone knows they have done the job correctly.
“People are starting to think a lot more about data as a competitive moat.”Matei Zaharia
The Association for Computing Machinery recognized this long arc in April 2026, naming Zaharia the recipient of its 2025 ACM Prize in Computing. The citation covered distributed data systems and computing infrastructure that enabled machine learning, analytics and AI at global scale. The prize joined an earlier ACM award for his doctoral dissertation, the SIGOPS Mark Weiser Award, an NSF CAREER Award and the U.S. Presidential Early Career Award for Scientists and Engineers.
Awards supply punctuation, though not necessarily a full stop. Zaharia's current publication list is crowded with work on agent reliability, verified attention and caching, semantic queries and prompt evolution. Spark solved the expensive repetition of data work. The current challenge is the expensive unpredictability of model behavior. Both reward the same unfashionable virtues: measurement, fault tolerance, careful interfaces and an honest account of what happens when things go wrong.
The engineer goes for a walk
Zaharia's public persona is more precise than theatrical. In a 2015 interview, asked what he did away from computers and big data, he replied that one is never very far from big data anymore. When he did escape, he said, he liked books, walks and occasionally trying to cook. He had enjoyed Andy Weir's The Martian, a novel in which survival depends on solving one practical problem after another with the equipment at hand. The preference is almost suspiciously on theme.
There is a temptation to tell technology stories as if one person sees the future and everybody else eventually catches up. Zaharia's record suggests a less mystical mechanism. He pays attention to what capable people cannot conveniently do. He works with collaborators. He builds an abstraction, opens it to users and lets their inconveniences sharpen the next question. The future, in this version, is not foreseen. It is debugged in public.
Spark began with researchers waiting for their algorithms. Seventeen years later, many of the systems beneath modern data work carry Zaharia's fingerprints, while his agenda has moved to the next blockage. The machinery is larger now. The principle remains pleasantly modest: if good people are spending their time fighting the tool, fix the tool.
Go deeper