Systems notebook: Cuckoo filters, cloud storage, hungry GPUs Bin Fan · VP of Technology · Alluxio founding engineer

Profile / Distributed systems

Bin Fan Has Spent a Career Shortening the Distance Between Data and Work

Before AI teams worried about feeding GPUs, Bin Fan was making key-value systems waste fewer bytes. His career follows one stubborn question across two technological eras: how do you keep expensive machines from waiting on data?

In the family album of computer science, infrastructure people are usually standing at the back. The models pose. The applications grin. Storage holds somebody else's coat. Bin Fan has built a career in that unfashionable corner, where a millisecond is examined like a suspicious receipt and wasted memory is treated as a small moral failure. Long before GPU clusters became coveted company assets, he was asking why useful machines spend so much of their lives waiting.

The question first appeared in peer-to-peer networks, then in flash storage, caches, hash tables and cloud systems. Today it appears whenever a rack of accelerators pauses for data sitting in a remote object store. Fan's titles have changed along the route. He is now VP of Technology and a founding engineer at Alluxio. The question has been remarkably loyal.

The education of a systems appetite

Fan studied computer science at the University of Science and Technology of China, then earned a master's degree at the Chinese University of Hong Kong. His early papers looked at BitTorrent-like file sharing: systems in which the location, availability and appetite of many participants determine whether anything arrives on time. It was a useful apprenticeship in distributed inconvenience.

At Carnegie Mellon, beginning in 2007, he joined the Parallel Data Lab and worked with professor David Andersen. The lab's concerns were concrete: networks, disks, memory, failure, cost. Fan's 2013 dissertation, Algorithmic Engineering Towards More Efficient Key-Value Systems, sounds modest until one remembers how much of the internet depends on key-value lookups. His aim was to make those systems scale while using less memory and keeping throughput high, even when requests arrived unevenly.

30%less memory for small key-value pairs in the MemC3 project description
3×throughput reported for MemC3 on read-heavy workloads
2013Carnegie Mellon computer science Ph.D. completed

The thesis did not merely diagnose hotspots. It assembled fixes: a fast load balancer for skewed traffic, compact structures for remembering what was present, and designs that got more work from each byte. MemC3 reworked the familiar Memcached formula with smarter hashing. SILT pursued persistent lookups on flash with a tiny memory footprint. The names were playful; the measurements were not.

“A data orchestration platform brings your data closer to compute across clusters, regions, clouds, and countries.”Bin Fan, describing the idea in 2019

A bird enters the database

The best-known artifact from this period is the cuckoo filter, developed with Andersen, Michael Kaminsky and Michael Mitzenmacher. Like a Bloom filter, it can answer a useful preliminary question: is an item probably in this set? The cuckoo filter uses compact fingerprints and, crucially, allows deletion. Its name comes through cuckoo hashing from the bird that displaces an egg in a nest. Computer scientists do occasionally permit wildlife into the machine, provided it improves locality of reference.

Published at ACM CoNEXT in 2014, the work became a practical option in the vocabulary of systems design. It also revealed Fan's recurring preference. He is drawn to the middle layer, the place where one carefully designed abstraction can spare the rest of a system from waste.

After completing his doctorate, Fan went to Google in Mountain View to work on next-generation storage infrastructure. An O'Reilly conference biography says he received Google's technical infrastructure award. His old Carnegie Mellon homepage preserves the transition with endearing academic plainness: one line announces that he started at Google in November 2013, followed by a long publication list and course notes. It looks less like personal branding than a desk left neatly arranged.

Bin Fan presenting a diagram about independently scaling compute and data at Data Council
THE MIDDLE LAYER GETS A SLIDE · Fan explains how compute can scale independently from data storage at Data Council. The podium is small; the architecture is not.

The layer between here and there

In 2015, Fan joined Alluxio as a founding engineer. The project had begun as Tachyon in UC Berkeley's AMPLab, built to share data at memory speed among frameworks such as Spark. Fan became a Project Management Committee maintainer and later ran the open-source initiative. He spoke about testing distributed systems, onboarding contributors and the less photogenic obligations of community software: documentation, reviews, releases and the patient conversion of a first contribution into a second.

Alluxio's language evolved with the infrastructure around it. In the big-data period, the problem was often Spark or Presto reaching files in HDFS or S3. As companies split compute from storage and adopted several clouds, Fan described “data orchestration”: a common access layer across scattered storage, with frequently used data cached near the application.

The phrase can sound grander than the plumbing it describes. Fan's explanation is reassuringly physical: bring data closer to compute across clusters, regions, clouds and countries. A cache absorbs repeated trips. A shared namespace hides some of the differences beneath it. Compute and storage can grow independently, leaving engineers fewer copies to move by hand.

Real deployments supplied the stakes. In 2017, Fan and Baidu engineer Haojun Wang presented an analytics workload accelerated by 30 times across hundreds of machines. A 2024 USENIX ATC paper, co-authored by Fan and a large group from Alluxio, Uber and Meta, described three years of production experience with a local cache for petabyte-scale analytical workloads. This was not the tidy world of a laboratory benchmark. It included changing files, metadata pressure, failures and the chronic unpredictability of popular data.

Then the GPUs began to wait

Machine learning made the old storage complaint expensive in a new way. Training consumes enormous datasets. GPUs are costly and scarce. If the network cannot feed them, beautifully parallel hardware performs the digital equivalent of waiting beside an empty conveyor belt. Checkpoints introduce traffic in the other direction, writing a model's state often enough to recover from failure without surrendering hours of work.

Fan's 2022 writing laid out the problem with characteristic practicality. Large training sets may live in several sources. The same data may be reused by many nodes and jobs. Copying everything onto local disks is awkward; reading everything remotely can starve the compute. His recommendation was conditional rather than evangelical: orchestration is useful when training is distributed, datasets are large, the network limits GPU use, or multiple frameworks need the same material. He has written, plainly, that there is no one-size-fits-all approach.

The hardware changed from “wimpy” key-value nodes to coveted GPU clusters. The engineering complaint remained: too much machine, not enough useful work.

By 2025, Fan's talks had moved squarely into multi-cloud AI. With Fireworks AI, he discussed distributing models across more than ten clouds and fifteen regions, where model files can exceed 70 gigabytes and cold starts carry a visible cost. At NVIDIA GTC, he discussed decoupled storage and compute with Solidigm. In February 2026, he joined an Oracle architect for a session on tiered caching and low-latency access on OCI. The nouns are fashionable now. The method is the same one visible in his thesis: locate the bottleneck, give the system a compact intermediary, and measure the result.

Peer-to-peer systems research at the Chinese University of Hong Kong.

CMU doctorate completed; Google storage work begins.

The cuckoo filter paper appears at ACM CoNEXT.

Fan joins Alluxio as a founding engineer.

Petabyte-scale cache research appears at USENIX ATC.

Current talks focus on the data path beneath multi-cloud AI.

Useful seriousness, with one April exception

Public traces of Fan's personality arrive mainly through technical work, and they suggest a preference for specifics. His talks name the interface, workload, latency and tradeoff. His contributor tutorial starts with a task small enough for a newcomer to finish. His GitHub profile uses the handle “apc999” and opens with a terminal prompt. The leading repository is a concurrent hash table library. Even the decoration is on duty.

There is mischief in the record too. On April 1, 2023, Fan announced that “CacheGPT” had joined Alluxio's Project Management Committee. The fictional model reviewed documentation entirely in emojis, patiently answered community questions and proposed next-generation cache designs. The joke landed because open-source maintainers know exactly how miraculous such a colleague would be.

Alluxio now presents Fan as VP of Technology and Founding Engineer. The role stretches from research papers to product architecture, customer workloads and public explanation. His career has accumulated the usual modern systems geography: China, Hong Kong, Pittsburgh, Mountain View, San Mateo; universities, a hyperscaler, a venture-backed company; local memory, remote clouds, global GPU fleets.

Across it all sits an unglamorous principle. Distance has a cost. Abstractions can hide complexity, but they cannot repeal physics. Good infrastructure acknowledges both facts and bargains carefully with each. Fan's work keeps making that bargain in smaller structures and larger systems.

The next generation of AI will be introduced with model names, benchmark charts and astonishing demos. Somewhere beneath the announcement, data will still need to arrive. If it arrives quickly, the middle layer may disappear from view. For an infrastructure engineer, invisibility is often the standing ovation.