# Reynold Xin

> Reynold Xin is the Canadian computer scientist, engineer, Databricks co-founder and Chief Architect whose work helped turn Apache Spark from a Berkeley research project into widely used data infrastructure. He designed or led major Spark systems including GraphX, DataFrames, Project Tungsten and Structured Streaming, helped set a world record for sorting 100 terabytes in 2014, and now applies the same taste for simple interfaces to databases, data intelligence and AI infrastructure.

- **Role:** Co-founder and Chief Architect at Databricks
- **Organizations:** Databricks, Apache Spark, Apache Arrow, UC Berkeley AMPLab, Google, IBM, Hyperbolic
- **Nationality:** Canadian
- **Education:** PhD in Computer Science, University of California, Berkeley, BASc in Engineering Science, University of Toronto
- **Known for:** Co-founded Databricks in 2013 with six other UC Berkeley AMPLab colleagues., Designed and led development of GraphX, Project Tungsten and Structured Streaming, and co-designed the DataFrame API in Apache Spark., Served as release manager for Apache Spark 2.0.

## Career timeline

- **Before 2010** — Worked on advertising infrastructure at Google and distributed databases at IBM before graduate school.
- **2011** — CrowdDB received the inaugural Best Demo Award at VLDB.
- **2012** — Shark won the Best Demo Award at SIGMOD.
- **2013** — Co-founded Databricks with six fellow UC Berkeley AMPLab researchers and Spark contributors.
- **2014** — GraphX was merged into Apache Spark; Xin led the Databricks team that set a Daytona GraySort record for sorting 100 TB.
- **2015** — Spark DataFrames and Project Tungsten advanced a common data API and a faster execution engine.
- **2016** — Served as release manager for Apache Spark 2.0, which introduced Structured Streaming.
- **2018** — Completed his Berkeley PhD dissertation, Go with the Flow, covering Spark SQL, Structured Streaming and GraphX.
- **2023-2024** — Presented Databricks' data warehousing and Apache Spark road maps at the Data + AI Summit.
- **2025-2026** — Worked on Lakebase and LTAP, arguing for operational and analytical workloads to share open storage while using specialized engines.

## Achievements

- Co-founded Databricks in 2013 with six other UC Berkeley AMPLab colleagues.
- Designed and led development of GraphX, Project Tungsten and Structured Streaming, and co-designed the DataFrame API in Apache Spark.
- Served as release manager for Apache Spark 2.0.
- Led the Databricks team that set the 2014 Daytona GraySort record for sorting 100 TB, completing the benchmark in 23 minutes.
- Won Best Demo Awards at VLDB 2011 for CrowdDB and SIGMOD 2012 for Shark.
- Co-authored highly cited systems papers on Shark, GraphX and Spark SQL.
- Earned a PhD in Computer Science from UC Berkeley with Michael Franklin and Ion Stoica as co-chairs.
- Helped create multiple billion-dollar product lines at Databricks, according to Gold House.

## Latest updates

- **2026-06** — Published an architectural explanation of Lakebase and LTAP, tracing his return to operational databases after years focused on analytics.
- **2026-06** — Joined Matei Zaharia on the Latent Space podcast to discuss Omnigent, agent infrastructure, Lakebase, LTAP and Databricks' product culture.
- **2026-06** — Appeared in Data + AI Summit sessions including the Tuesday keynote and forums on technology and cybersecurity.
- **2025-12** — Co-authored a Lakebase product update covering autoscaling, scale-to-zero and rapid provisioning.
- **2025-01** — Participated as an angel investor in Gumloop's Series A, according to CB Insights.

## Links

- Website: https://rxin.github.io/
- LinkedIn: https://www.linkedin.com/in/rxin
- Twitter/X: https://x.com/rxin
- GitHub: https://github.com/rxin

---

Profile page: https://yespress.io/reynold-xin
Published by YesPress — https://yespress.io
Last updated: 2026-09-26
