Breaking
The database was solved - until production disagreed100 terabytes · 23 minutes · one stubborn benchmarkPascal first, Spark later

Profile / Systems & second thoughts

Reynold Xin Keeps Making the Complicated Feel Inevitable

The Databricks co-founder helped give big data a humane interface. Now he is returning to a problem his adviser once called solved - and finding it wonderfully unfinished.

The advice was excellent, authoritative and wrong in precisely the way that good advice sometimes is. When Reynold Xin began his doctorate at the University of California, Berkeley, his adviser told him that online transaction-processing databases were solved. They worked. The interesting frontier lay in analytics, where torrents of structured and unstructured data were beginning to make old machinery groan. Xin followed the advice. Sixteen years later, he would open an essay about his newest database work by repeating it, like a detective lovingly preserving the first mistaken clue.

Following that clue led somewhere consequential. At Berkeley’s AMPLab, Xin joined the research effort that became Apache Spark. In 2013, he and six lab colleagues founded Databricks. He helped design the systems by which Spark learned to speak graph, SQL and streaming data without demanding that every user become a specialist in each. Then Databricks became a customer of operational databases. The supposedly finished technology proved clunky, fragile and hard to scale. A solved problem, Xin discovered, is often merely a problem enjoying tenure.

“We realized OLTP databases were far from a solved problem: they were clunky, difficult to scale, and incredibly fragile.”Reynold Xin, on the road to Lakebase

The circle from analytics back to transactions suits him. Xin’s career has not been a procession of unrelated inventions. It is an extended argument about who should carry complexity. His answer is rarely the user.

The boy in the forum

Xin is Canadian and studied Engineering Science at the University of Toronto. Programming arrived earlier, in middle school, by way of Pascal and then PHP. He taught himself and spent time in online forums. In a 2016 interview in Tokyo, he connected those early communities to the work of maintaining open source: long before he was negotiating the direction of Spark, he had learned that software is also a conversation among people who owe one another patience.

Before Berkeley, he worked on advertising infrastructure at Google and distributed databases at IBM. Graduate school brought a string of projects with names that sounded more charming than their technical papers. CrowdDB explored how a database might use people to answer questions computers could not; it received the inaugural Best Demo Award at VLDB in 2011. Shark put fast SQL analytics on Spark and won the SIGMOD Best Demo Award the next year. Success did not make Shark sacred. Its lessons fed Spark SQL, a more general system. In infrastructure, sentimentality is an expensive storage format.

AMPLab was an unusually fertile neighborhood. Matei Zaharia had created Spark; faculty members Ion Stoica and Michael Franklin helped shape the lab; Ali Ghodsi, Andy Konwinski, Arsalan Tavakoli-Shiraji and Patrick Wendell joined Zaharia and Xin in carrying the work into Databricks. The founding team was not assembled by a recruiter matching seven heroic job descriptions. It grew from people who had already argued over code, papers and machines. Their company inherited that collective grammar. Xin’s role within it remained architectural: finding the layer at which several kinds of work could become one product without becoming one muddle.

He eventually completed the doctorate in 2018. Its title, Go with the Flow, is jauntier than most dissertations; its subtitle names graphs, streaming and relational computation over distributed dataflow. The document gathers GraphX, Structured Streaming and Spark SQL into a single academic case. What appears in a product brochure as breadth appears there as a precise systems claim: fine-grained recovery, relational optimization and real-time computation need not live in separate kingdoms. A dissertation usually closes a period of apprenticeship. Xin’s reads more like the specification of the next decade.

GraphX challenged the idea that graph computation required a separate, purpose-built world. DataFrames gave programmers in different languages a common, structured way to work. Project Tungsten reworked Spark’s execution so modern CPUs and memory were used with rather less waste. Structured Streaming let a developer describe an evolving result with the same DataFrame model used for static data. The names change; the instinct holds. Bring more work under an interface that people can understand.

Why “Tungsten”?

Hadoop projects favored animals. Spark favored light. Tungsten, the metal of a lamp filament, suggested software and modern hardware clicking together until the machine lit up.

A trillion little proofs

Abstractions need evidence. In October 2014, Xin led a Databricks team into the Daytona GraySort benchmark, a stern and wonderfully literal test: sort 100 terabytes of data, or one trillion 100-byte records. Their Apache Spark setup finished in 23 minutes on 206 virtual machines. The result set a world record for the category and beat the earlier Hadoop MapReduce record by a wide margin. A research engine, running in the public cloud, had acquired a useful swagger.

100 TBdata sorted
1T100-byte records
23 minelapsed time

The benchmark is easy to admire because its numbers are so obliging. Yet Xin’s durable contribution is less photogenic: he helped transform performance into a product of the system rather than an ordeal repeatedly assigned to its users. DataFrames could be optimized because they described intent instead of dictating every operation. Tungsten could change execution beneath the API. Structured Streaming could update results incrementally while preserving a familiar mental model. The fast path became the path people could actually find.

Reynold Xin presenting on a large conference stage at the 2023 Data and AI Summit
Big room, small type: Xin takes Databricks’ data-warehouse argument to the 2023 Data + AI Summit stage. Photo by Michael Segner.

His public manner reflects this preference. He can discuss registers, write-ahead logs and query optimizers, but he returns to the shape of the user’s problem. When he recently wrote an explanation of Lakebase, he said the difficult part was choosing depth: one draft contained too little detail to land the idea; another contained so much it lost readers who were not storage experts. He published the third. Architecture, in that telling, includes the reader.

Open systems, high bars

There is steel underneath the accessibility. Xin speaks often about first principles and high standards. Databricks resisted pressure to build an on-premises business because its founders believed computing would keep moving to the cloud. The company maintained a demanding hiring bar even when lowering it would have made growth easier. Short-term accommodation, in his account, can become a long-term cultural ratchet.

The severity does not turn into autocracy. Asked whether Spark might be run in the style associated with a singular project dictator, Xin laughed the idea away. Apache’s democratic process mattered; an overbearing leader would drain the community’s energy. He served on Spark’s Project Management Committee and managed the Spark 2.0 release, but described release management as carrying out decisions the community had already made. Authority was useful when it clarified the work, not when it became theatre.

His GitHub repository “Readings in Databases” offers another clue. It is a syllabus of algorithms, relational foundations, classic systems, consensus and modern hardware. The notes are crisp and occasionally opinionated. Proposed additions may wait, he warns, because he needs to read the paper. This is perhaps the most architect-like promise imaginable: the queue will move only after the dependency has been inspected.

The machinery stays difficult. The achievement is allowing everyone else to think about something better.The pattern across Xin’s systems work

Back to the beginning

Xin’s recent work returns to the database that was allegedly finished. Lakebase, Databricks’ serverless Postgres system, separates compute from the log and data services that preserve a database’s state. That makes compute disposable, scaling less ceremonial and branches cheap enough to create quickly. The larger LTAP idea tries to let transactions and analytics meet at the storage layer: one open copy of data, with different engines suited to different jobs. It is vintage Xin. Do not force one specialist to impersonate another. Give them a common foundation.

AI agents make the argument urgent. An agent needs live operational context, permissions, memory and a safe place to act. In a 2026 conversation with fellow Databricks co-founder Matei Zaharia, Xin’s requests for an agent environment were charmingly concrete: he wanted a shell, files he could inspect and Markdown rendered properly. The grand future of software apparently still requires someone to tail the logs.

This pragmatism keeps his story from becoming a fable about prediction. Xin has been right about cloud computing, open systems and unified data tools, but his best work often begins with a correction. Shark gave way to Spark SQL. Analytics led back to transactions. A supposedly solved category reopened. The architect’s advantage is not clairvoyance. It is the willingness to revise the plan without surrendering the principles beneath it.

In 2016, asked about life beyond work, Xin paused. Skiing, he said, was about the only serious pursuit outside programming. He had started young, was good at it, and had spent so many years on Spark that he had not thought much about some distant personal destination. The answer sounds less monastic than kinetic. Skiing and systems design share an unforgiving relationship with edges. Both reward balance, speed and the good sense to choose a line before gravity chooses one for you.

Xin’s line now runs from Pascal forums through Berkeley papers, a cloud sorting record and a company built by seven researchers, back toward the database underneath the applications. It has bent considerably. It has not wandered. Again and again, he arrives at the same question: can hard infrastructure be made broad, open and simple enough that more people can do serious work with it? The answer is never finally yes. For an architect, that is excellent news.