There is a particular kind of frustration that follows a successful invention. The thing works. The demonstration is convincing. The paper is published. Then someone who needs the thing tries to use it, and discovers an obstacle course of data preparation, computing choices, unfamiliar tools and costs. Ce Zhang has spent much of his career in that awkward distance between a working idea and a usable one. In 2017, newly arrived at ETH Zürich, he gave the problem an unusually clear name: the gap between the availability and accessibility of machine learning techniques.
It was a professor’s question, but it did not remain in the seminar room. Zhang went on to co-found Together AI in 2022. He is now its founder and chief technology officer, working on the cloud systems that let organizations train, adapt and serve AI models. The tools are different from those of his first research projects. The question that animates them is remarkably familiar: how much specialist labor should stand between a promising model and someone who can put it to work?
His route to that question passed through Peking University, where he did undergraduate study, and the University of Wisconsin-Madison, where he earned his doctorate. He then spent a postdoctoral year at Stanford. Christopher Ré advised him through the Wisconsin and Stanford stages; years later, Ré would be one of his Together AI co-founders. Those connections give the story continuity, but the technical path matters too. Zhang worked where database systems met machine learning, an intersection concerned with the unglamorous business of turning information into something a model can learn from.
A lab that asked to be taught
At ETH, Zhang joined the computer science faculty in September 2016 and led the Data Sciences, Data Systems, and Data Services Lab, known as DS3Lab. He described an ambition to connect systems researchers, machine learning specialists and people in other disciplines. This was more than a nice sentence about collaboration. His group spent Fridays with visitors from other departments, who gave two-hour crash courses in their own fields. A researcher accustomed to optimizing code might instead find themselves learning the vocabulary of astronomy.
The ritual was practical. Zhang wanted students to understand the problem as their collaborators experienced it. A platform that seems elegant to its author can ask too much of someone with a different job. In his ETH interview, he urged aspiring researchers to feel the pain and frustration of the users they hoped to help. He called research fun, but his account of fun included the hard work of learning what another person actually needs.

That outlook helps explain the breadth of his research. DeepDive, a trained data system for building knowledge bases, emerged from his doctoral work. At ETH, his group pursued data management and machine learning systems, and he later wrote about data quality and the systems needed for trustworthy AI. Different papers tackled different bottlenecks; the recurring concern was how to turn difficult technical methods into tools people could use reliably. It is the kind of agenda that rarely fits on a single product box. It can, however, follow a person across institutions.
“My current research interest is to understand—and try to close—this gap between the availability and accessibility of machine learning techniques.”Ce Zhang, 2017
The view from the terrace
Zhang’s recollection of his first ETH visit has a pleasing lack of technical detail. The day before his interview was sunny. He stood on the main building’s terrace and looked across the lake and old town. He called the view breathtaking. He then noticed the people: potential colleagues with complementary interests and students whose enthusiasm impressed him. A computer scientist seeing a beautiful scene is hardly a surprise; remembering both the view and the future collaborators says more about why he chose the place.
An ETH questionnaire offered other glimpses of him. Outside science, he named calligraphy and looking out over rivers, lakes or the ocean. If given a paid year away, he said he would return to graduate school to learn a new field, perhaps psychology, astronomy or biology. Asked what he might do if he were not a scientist, he answered: chef. The answer has a gentle comic precision. A chef is also in the business of combining ingredients, though dinner is generally expected sooner than a research paper.
His lab philosophy sounded less like a competition for cleverness than a promise to students. He said its educational duty was to put them in a position to reach goals they chose for themselves. He also described the group’s research style as systematic and bottom-up. It is a useful pairing: ambition for the student, patience with the problem. The Friday crash courses fit both. They widened the set of questions the group could ask, while obliging its researchers to listen before they designed.
A research question becomes infrastructure
Together AI began in 2022 with Zhang, Vipul Ved Prakash, Christopher Ré and Percy Liang among its founders. Prakash had built a technology company before; Ré and Liang brought long research careers; Zhang arrived with years spent on the intersection of data systems and machine learning. Together’s work centers on infrastructure for open AI models: the compute, training and serving layers beneath an application. Its own description of the company now emphasizes a full stack platform for production AI powered by systems research.
A platform is an ambitious answer to an accessibility problem because it has to confront the whole chain. A model may be downloadable, yet difficult to run economically. A promising experiment may fail when demand arrives at production scale. A developer may gain model choice and inherit a new set of systems decisions. Zhang’s job as CTO sits among those tensions. In public research and talks, he has kept returning to the mechanics: efficient inference, the interaction of data and computation, and how an open model moves from an academic result into an application people depend on.
His current work supplies concrete examples. In January 2026 he was among the authors describing how Together worked with Cursor on low-latency inference, where responses have to arrive inside the rhythm of a programmer’s editor. In March, he co-authored work on cache-aware serving for long prompts, an attempt to keep inference responsive when applications send large, mixed workloads. He also appeared in Together’s NVIDIA GTC 2026 program discussing lessons from production inference. The vocabulary is specialized; the underlying demand is ordinary. People expect software to respond when they need it.
There is a difference in scale between a Friday crash course and a production model service, of course. The link is the insistence that the system be judged where it meets its user. A student who learns what an astronomer needs is less likely to build an elegant dead end. An infrastructure team that understands an application’s latency budget is better placed to make a model usable. Zhang’s career offers those scenes as variations on one method: work backward from the friction.
A foot in both rooms
Zhang has kept an academic appointment alongside his company role. He is a Neubauer Associate Professor of Computer Science and Data Science at the University of Chicago, where his stated research concerns the tension among data, models, computation and infrastructure. He also serves as a co-editor-in-chief of Data-centric Machine Learning Research. At Together AI, the work is exposed to customers, traffic and the price of computing. Each setting asks a different question of an idea, which may be why he has continued to work in both.
The two roles also put different clocks on the same questions. A paper can isolate one variable and explain a result with care. A production service has to answer at once for cost, speed and reliability, even while the models change. Zhang has written that his research aims to make machine learning cost-efficient and trustworthy as well as accessible. Those are qualities that become tangible in a service: how long it takes to respond, what it costs to run, and whether the system behaves consistently enough for someone to build around it. The professor can ask why a system works. The CTO must also ask whether it works on Tuesday.
The recognition along the way is substantial: a SIGMOD Best Paper Award, a SIGMOD Research Highlight Award, a Google Focused Research Award and an ERC Starting Grant. Yet awards are a curious way to measure a career dedicated to access. They record that peers noticed the work. Zhang’s own description of the goal is less ceremonial: make machine learning widely available, cost-efficient and trustworthy for people who want to use it. That is a standard with no neat finish line. Every change in models, hardware and demand can reopen the gap.
The charming detail about the would-be chef remains tempting, but the better image may be that Zurich terrace. Zhang saw an attractive view, then noticed the people with whom he could work. A decade later, the view is an AI cloud and the company building it. The work still depends on seeing beyond the technique to the people waiting on the other side of it. A powerful model may be an achievement. Making it usable is a longer story.