Latest
September 2026: Fireworks introduces a Specialized Intelligence IndexJuly 2026: Fireworks reports more than 40 trillion tokens served dailySeptember 2026: Fireworks introduces a Specialized Intelligence IndexJuly 2026: Fireworks reports more than 40 trillion tokens served daily

The profile / AI infrastructure

Benny Chen Saw the GPU Bill Before the AI Boom

Inside Meta, rising demand for machine learning compute made the future feel less like a prediction than a capacity meeting. Chen left to co-found Fireworks AI in 2022; now his argument is that useful AI must be affordable to run and specific to the job.

The future of artificial intelligence first appeared to Benny Chen as a capacity problem. At Meta, he worked on the infrastructure that served advertising models. The models had to answer quickly and often. The numbers in planning meetings kept rising. GPUs, servers, specialized chips: an internal appetite for machine learning was becoming hard to overlook. Chen later recalled the demand going “up and to the right.” In 2022, before ChatGPT made generative AI a family conversation, he left with a group of engineers to build Fireworks AI.

It is a less cinematic founding scene than a lone programmer at a kitchen table. There is no sudden revelation in a shower. There is a spreadsheet, a compute curve and a question: if one large company needs this much machinery, what happens when everyone else starts building with AI too? For Chen, who had spent years close to model serving, the answer pointed toward infrastructure. The public might meet AI through a chat box. Someone would have to keep the box running, reliably and at a price its owner could bear.

That distinction has shaped his career. Chen is a co-founder of Fireworks, alongside CEO Lin Qiao and a founding team drawn largely from Meta’s PyTorch and machine learning organizations, with Chenyu Zhao bringing experience from Google Vertex AI. Chen’s own specialty was Meta’s ads infrastructure. Advertising is rarely the glamorous chapter in an AI story, but serving predictions at enormous scale is good preparation for the less glamorous question beneath every clever demo: how does it work when millions of people use it?

The bet before the chatbot

Chen studied computer science and engineering at UCLA and worked as a backend engineer before joining Facebook in 2014. There he moved through software engineering roles and became a principal engineer. The work that stayed with him involved serving ads models with GPUs and custom chips. It meant thinking about the machinery, the cost and the next surge in demand, not simply whether a model could produce an impressive answer once.

Fireworks began in 2022 with a different product emphasis from the one it has today. Chen has said the team first built a PyTorch training platform. That made sense to engineers who knew the ecosystem and had seen compute demand grow. Then ChatGPT launched. The question changed. Was generative AI a passing frenzy, or the workload that would make the infrastructure problem everyone’s problem?

The open models available then were hardly an easy sales pitch. Chen has recalled how early Falcon and Llama models struggled to keep a short conversation going. A platform built around those models required faith that the models would improve, and enough engineering patience to make them usable along the way. Looking backward can turn a choice into an inevitability. Chen resists that neatness. “So it was a bet, and I’m glad it paid off,” he said of the shift.

The founders also had the ordinary shortcomings of a young company. Chen has described an early website as “very ugly” because the team had no front-end engineer. The detail is small and useful. Infrastructure founders often know exactly how much time a model spends on a GPU while a visitor still has to squint at the homepage. Early certainty about the market did not grant polish everywhere else.

The expensive art of answering

Chen’s preferred explanation of model economics is almost comically modest: a voice assistant booking a restaurant table. It does not need to know every historical joke or solve every unfamiliar problem. It needs to understand the request, choose a table and complete the booking. Paying for a general model’s full breadth on every such request can become an expensive habit once a product has real users.

That is the case for specialization. A company can start from an open model, tune it around its own tasks and serve it efficiently. The point is not merely to produce cheaper tokens. It is to make the model fit the work: a coding assistant, a search feature, a customer service workflow. Each job has different standards for speed, accuracy and cost. Chen argues that businesses should be able to define those standards themselves rather than inherit the same general answer for every problem.

Fireworks has built the layers underneath that proposition: access to GPUs, model-specific optimizations, a runtime to execute requests, and routing that can send traffic where it needs to go. Chen described those layers in a 2026 infrastructure interview, down to the difficulty of buying the right GPUs in the right places. At scale, the unromantic details become the product. A model that answers beautifully in a test but stalls in production is not much help to the person waiting for it.

Benny Chen, right, with MongoDB's Gregory Maxson at MongoDB.local NYC in 2024
Benny Chen, right, with MongoDB’s Gregory Maxson at MongoDB.local NYC in 2024. The conversation turned on getting AI applications into production. Photo: SiliconANGLE.

The production theme was visible in a 2024 conversation with MongoDB’s Gregory Maxson. At MongoDB.local NYC, Chen spoke about helping developers move from a model experiment toward an application people could actually use. The two companies discussed combining model services with access to business data. For a founder who came from model serving, the attraction is clear: AI becomes valuable when it meets a real workflow and survives the trip from demo to daily use.

Who decides what good means?

Speed and cost lead to another question, one Chen returns to repeatedly: how do you know whether a model is doing its job? A company can count tokens and milliseconds. Judging quality is harder. He offers the example of AI-generated HTML. The code can be syntactically correct and still create an ugly or broken page. A judge reading only the markup will miss what a person sees on screen. Even a screenshot of the first screen may hide a mess farther down.

To train a model for a particular business, someone must turn the business’s sense of quality into a usable evaluation. That work requires product judgment as much as engineering. A support team may know the difference between an answer that technically addresses a question and one that actually resolves the customer’s problem. A coding team may know when a patch passes a narrow test but makes the system worse. The hard part is writing those standards down clearly enough that a machine can learn from them.

“No one model leads on every dimension.”Benny Chen, on evaluating models for specialized work

Chen also argues that training and serving should resemble each other. If a model learns in a tidy simulation and then operates in a very different production setting, the feedback may teach the wrong lesson. Fireworks’ work on reinforcement learning and evaluation tools follows that logic. Teams need ways to judge model behavior in conditions close to the ones their customers actually encounter.

This is where Chen’s view of AI feels unusually grounded. The person with the product idea still has to say what success looks like. The engineer still has to build a fair test. A faster GPU cannot settle a disagreement about taste, usefulness or trust. The machine can optimize for a standard, but a human team has to choose the standard in the first place.

A company grows around the bill

Fireworks’ public milestones show how far the original capacity hunch has traveled. A 2024 funding round put the company’s valuation at $552 million. In October 2025, it announced a $250 million Series C at a $4 billion valuation. In July 2026, the company announced a $1.505 billion Series D at a $17.5 billion valuation and said it was serving more than 40 trillion tokens a day. These are company figures, and they describe scale rather than an individual’s wealth. They also make the old planning meeting easier to understand: the rising curve did not stay inside Meta.

40T+tokens served daily, according to Fireworks’ July 2026 announcementA company measure, not a personal metric

The company says more than 95 percent of those tokens come from models specialized on customer data and particular jobs. Its customers include teams building coding and legal AI products. That mix matters to Chen’s argument. The prize is not simply hosting a catalog of interchangeable models. It is helping a business shape a model around what it knows, then run that model repeatedly without letting the bill swallow the product.

His public role has broadened with the company. He has spoken in interviews about inference economics and hardware choices, and at an AI engineering event about evaluations. In September 2026, as Fireworks introduced a Specialized Intelligence Index for benchmarks drawn from real work, Chen argued that the best model depends on the dimensions a task actually requires. The point is a useful check on any universal leaderboard: a high score is only relevant if it measures the thing you need done.

The person behind the plumbing

Chen’s most revealing anecdote is about how difficult it once was to explain his job at home. He told his parents he pushed ads onto Instagram. They saw the ads, but the machinery behind them remained abstract. Years later, he watched his mother use AI applications for her own work. He had expected infrastructure demand to rise; he had not expected adoption to become so ordinary so quickly.

That surprise keeps the story human. Founders are sometimes described as if they can see the next decade with perfect resolution. Chen saw a pressure building in the systems he knew. He made a decision, adjusted the product when the world changed, and admits that parts of the outcome surprised him. He has also talked about the need for patience when teams customize models. The first useful version takes more than an enthusiastic launch announcement.

Outside the engineering discussion, he has described a small thinking practice: a quiet walk, or 15 to 30 minutes with nothing to do. He likens the wandering thought to turning up a language model’s temperature. The joke works because it admits something the infrastructure can never quite replace. A person still needs room to reconsider a decision before sending another request into the machine.

The company he helped start now operates at a scale that makes its early, awkward website seem distant. Yet the questions that brought Chen out of Meta remain recognizable. How much compute will a useful answer require? Who pays for it? How do we know it is good? The AI industry has accumulated grander claims and larger numbers since 2022. Chen’s story keeps returning to those plain questions, which is perhaps why the capacity meeting makes such a good beginning.