Profile / Chenyu ZhaoGoogle Vertex AI to Fireworks AIEngineering the instant answerProfile / Chenyu ZhaoGoogle Vertex AI to Fireworks AIEngineering the instant answer

People / AI infrastructure

Chenyu Zhao and the Engineering of an Instant Answer

He helped build Google Cloud's Vertex AI, then co-founded Fireworks AI to make open models run at production scale. His through-line is the difficult engineering between an impressive model and an answer that arrives on time.

A person types a question. A response appears a moment later. The interval looks like magic only until the question is multiplied by millions. Then it becomes a problem of machines, placement, memory, routing and price. Chenyu Zhao has spent much of his career in that interval. He helped build Google Cloud's Vertex AI and later co-founded Fireworks AI, where he has led the core infrastructure team. The attractive part of artificial intelligence is what it can say. Zhao's work is about whether it can say it quickly enough, reliably enough and often enough to become a product.

On a Silicon Valley stage in 2025, Zhao made the stakes unusually plain. His talk was called Building AI Agents. A slide declared, “Model is Your Product.” The line is short, but it carries a long engineering bill. If a model is part of a product, its quality cannot be considered apart from the time it takes to respond, the way it improves on a company's own data, or what happens when thousands of people ask for help at once. A polished demo gets a round of applause. A working service gets another request.

Chenyu Zhao presenting a slide that reads Model is Your Product at the 2025 TechEquity AI Summit
Zhao onstage at the 2025 TechEquity AI Summit. The slide makes the product argument in five words. Still from his recorded talk.

An education in scale

Zhao studied electrical engineering and computer science at the University of California, Berkeley, from 2010 to 2013. His public GitHub account still carries evidence of an engineer willing to go down interesting side streets: a Java project for reconstructing a three-dimensional object from a series of two-dimensional images, and a Quake 3 graphics engine, also in Java. These are old public projects, not a map of his private life. They do, however, make an amusing prelude to a career spent persuading computers to do complicated things under real constraints.

After Berkeley came nearly a decade at Google. Zhao rose to senior staff software engineer. He headed a machine learning team behind the company's mobile ads business, working on large-scale modeling, ad auction algorithms and the systems that supported them. Advertising auctions are a stern school for infrastructure: the result must be calculated under pressure, at enormous volume, while cost and delay remain visible. The work did not look like today's conversational AI, but the discipline of serving machine learning inside a busy product crossed over neatly.

At Google Cloud, Zhao helped co-found Vertex AI and led a 50-person engineering team developing AutoML, MLOps and explainable AI solutions. The terms can sound like a conference program, so it helps to translate them. AutoML makes parts of model building easier to automate. MLOps is the work of operating models after they leave an experiment. Explainability gives people tools for understanding what a model is doing. Together, they address the awkward middle ground between a promising algorithm and a system an organization can use.

2013Berkeley studies completed; the Google chapter begins.
2013–22Mobile ads machine learning, then Google Cloud's Vertex AI.
2022Fireworks AI is founded with six fellow co-founders.
2025Zhao speaks about building AI agents at a Silicon Valley summit.
2026Fireworks reports serving more than 40 trillion tokens a day.

A different kind of founding team

In 2022 Zhao joined six other founders to start Fireworks AI. The group mixed experience from Google with engineers who had worked on PyTorch and machine learning infrastructure at Meta. CEO Lin Qiao had led PyTorch at Meta; fellow founders included Benny Chen, Dmytro Dzhulgakov, Dmytro Ivchenko, James Reed and Pawel Garbacki. The combination mattered. PyTorch had helped make model development accessible to researchers and engineers. Fireworks set out to tackle the next practical question: how to run and customize those models for products that need to answer in real time.

Zhao's place in that team is distinctive. He brought experience from a managed AI platform built inside a cloud giant. At Fireworks, his public speaker biography describes him leading the core infrastructure team and building reliable, scalable systems for training and inference. The latter word simply means running a trained model to produce an answer. It sounds tidy until the answer has to come from the right hardware, at the right cost, in the right region, under a changing load. Infrastructure is full of these unglamorous adjectives. Users notice them most sharply when one fails.

The company's first platform release in 2023 put speed, affordability and customization in the same sentence. Developers could run models through an API, fine-tune them for a particular job and deploy the result. That pairing of customization and serving would become central to Fireworks' story. A generic model may know a great deal in the abstract. A business usually needs a model to work with its own vocabulary, documents and workflows. The practical task is to make that specialized behavior available without asking every product team to build an inference stack from scratch.

“Model is Your Product.”Chenyu Zhao, Building AI Agents, 2025

The economics of a quick reply

At AMD's AI Dev Day in 2025, Zhao joined Crusoe's Erwan Menard to discuss open models, optimization and GPU infrastructure. His comment about the cloud was brisk: “We rely on clouds like Crusoe to give us the best price per performance.” There is no romance in that sentence, and that is why it is useful. Inference costs money every time a customer asks for an answer. The platform has to make choices about hardware, software and capacity so an application can afford to keep answering. A beautiful response delivered on an uneconomic machine has a short career.

Those choices become harder as the menu of models expands. Some are fast and light; some are large and demanding; some are tuned to a customer's data. An engineering team must decide how requests are routed, how much work each GPU can handle, when to cache, and what happens if a server disappears. These are industry-wide infrastructure problems, and Fireworks presents them as part of its product. Zhao's publicly described role places him close to that machinery. It is a kind of work where success often looks like a user having no reason to think about the machinery at all.

The scale is now difficult to visualize. In July 2026, Fireworks said it served more than 40 trillion tokens a day and had passed $1 billion in annualized revenue run rate. It announced a $1.505 billion Series D at a $17.5 billion valuation. The company said more than 95 percent of those tokens came from models specialized on customers' proprietary data. Those are Fireworks' reported figures, and they describe the business rather than one engineer's personal achievement. They do show how far the original infrastructure question has traveled since the company's 2022 founding.

2022Fireworks founded
40T+Tokens served daily
95%+From specialized models

Company-reported figures from July 2026. Tokens are units of text processed by AI models.

Growth also changes the geography of an answer. Fireworks acquired Hathora in March 2026 to strengthen global compute orchestration. Hathora's background was low-latency multiplayer game infrastructure, where a few extra milliseconds can turn a smooth experience into a frustrating one. For AI systems, the comparable challenge is routing a request to available compute across regions and hardware types. The acquisition was a company decision; it should not be credited to Zhao personally. It is relevant to the engineering territory his team works in and to Fireworks' shift from a clever service into a large operating system for responses.

The product inside the model

Zhao's 2025 presentation on agents brought those infrastructure questions back to the application. He moved through examples of agents that can take signals, perform actions and receive feedback. His talk connected the broad excitement about agents to the narrower work of training and deploying them. The slide about the model being the product captured the turn: if an agent is expected to do useful work, the model must be chosen and adapted with that work in mind. The product experience depends on the model's behavior and on the system that makes that behavior available.

That idea runs through Fireworks' move toward specialized intelligence. The company describes customers building models around their own data and workflows instead of relying only on general-purpose behavior. Cursor's coding models and Harvey's legal AI appear among the examples in its 2026 announcement. Neither example is Zhao's personal project in the public record. Both show the kinds of applications for which an inference platform becomes consequential: the model must know the task, and the service must keep up with the people using it.

The distance between a researcher releasing a model and a customer depending on it can be vast. A model might need quantization to fit efficiently on hardware, tuning to understand a domain, monitoring to catch poor results and routing to meet a deadline. That distance is where Zhao's career becomes a coherent story. At Google, he helped create tools for building and managing machine learning. At Fireworks, he has helped build the underlying systems that run it. The names and interfaces changed; the stubborn problem of making machine learning usable at scale remained.

A small joke in a large system

There is one small personal detail in Zhao's otherwise technical public presence. His LinkedIn biography tells recruiters that if the person on camera does not look like a hamster, it is not him. It is a dry joke about identity in a world of online interviews, and a welcome reminder that an infrastructure engineer can have a sense of humor. His public posts tend to stay near the work: a note about Fireworks' Llama 3 serving in 2024, a welcome for the Hathora team, and other developments around models and systems.

The more lasting portrait is professional. Zhao's path joins an ads auction, a cloud AI platform and an inference company into one line of inquiry: what must be built so that a model's promise survives contact with actual users? Fireworks' numbers are now enormous, but the unit of experience remains small. One person asks. One answer arrives. Between those two events sit teams, code, data centers and the decisions that make the exchange feel simple. Zhao has spent years working in that hidden interval. The next request is already on its way.