THE PROFILEFrom Kharkiv’s coding contests to the machinery behind open AI modelsDmytro Dzhulgakov · Fireworks AI

People / Engineering / The long game

Dmytro Dzhulgakov Learned to Race Code. Now He Races AI Models.

Before he worked on the machinery behind generative AI, Dmytro Dzhulgakov was a Kharkiv student who wrote code for an Intel 8080 and competed against the world’s fastest programmers. At Fireworks AI, the contest has changed, but the clock still matters.

At eight, Dmytro Dzhulgakov wrote a program in Assembler for an Intel 8080. The chip belonged to an earlier era of computing, but the attraction was already modern: give a machine a precise set of instructions and see what happens. He later recalled that he had been obsessed with Lego as a child and thought that might explain his interest in engineering. The toys became code. The code became contests. The contests became a career spent measuring what machines can do under pressure.

In 2010, he was an applied mathematics student at the National Technical University “Kharkiv Polytechnic Institute.” He lived in Kharkiv, a city he described as one of Ukraine’s educational and scientific centers. In 2010, he sketched a technical childhood without trying to turn it into mythology: Assembler at eight, then Pascal and Delphi, then C++ at sixteen. Outside programming, he liked travel and photography. The competition, he noted with a light touch, meant a trip to Las Vegas.

That small detail matters. The man now identified with GPU throughput and model serving once had a very tangible prize for solving hard problems quickly: a journey across the world. Competitive programming gave his curiosity an audience and a stopwatch. It also taught a habit that still appears in his public comments: a result is only as good as the conditions under which it was measured.

2006IOI silver medal
6Google Code Jam finals
2ICPC World Finals

Kharkiv, with a clock running

The record is unusually concrete. Dzhulgakov won a silver medal at the 2006 International Olympiad in Informatics. He reached six Google Code Jam finals between 2008 and 2015, finishing fifth in 2009 and ninth in 2012 and 2014. His university teams reached the ICPC World Finals in 2010 and 2012. He also made the Topcoder Open algorithm final. These are competitions where elegant ideas meet unforgiving execution: the program has to work, and it has to work before time expires.

Contest rankings are easy to read as a tally of victories. Their more interesting legacy may be the kind of judgment they cultivate. A competitor has to decide which problem to attack, which shortcut is safe, and when an apparent solution is hiding a bad assumption. Years later, Dzhulgakov would discuss AI services in similarly practical terms. How long does a response take? How much hardware does it consume? Does a benchmark reflect the customer’s actual work? The scale changed from one program to a distributed service. The impulse to test the answer survived.

“I wrote my first program in Assembler at the age of 8 for the Intel 8080 chip.”
Dzhulgakov, recalling his first program in 2010

His formal studies gave the puzzle instinct a broader base. He described applied mathematics as a mix of mathematical theory, algorithms, and practical programming. The description sounds almost like a map of his later career. The theory matters because models and systems need sound ideas. The algorithms matter because resources are finite. The practical programming matters because customers never run a proof on a whiteboard.

When research had to survive production

At Meta, Dzhulgakov worked where machine learning left the lab and entered products used at enormous scale. He became a PyTorch core maintainer, part of the small group responsible for the framework’s technical direction. Public PyTorch governance lists included him among its core maintainers in 2022; a 2026 list describes him as emeritus. His own professional descriptions connect the work to Facebook’s advertising systems and to making PyTorch useful beyond research experiments.

Two papers put some shape around that work. A 2018 study of deep learning inference in Facebook data centers examined the performance and hardware implications of running models in production. A 2019 paper introduced a deep learning recommendation model and a way to divide its work across machines. Dzhulgakov appears among the authors of both. Neither is the sort of contribution that comes with a dramatic product unveiling. Both concern the persistent engineering question beneath machine learning: what happens when the model meets real capacity, real traffic, and real limits?

He could also be found answering ordinary developer questions in public PyTorch forums. In one 2020 exchange, a user could run a quantized model on a computer but received an error on Android. Dzhulgakov identified the likely mismatch and suggested ways to return data the Java binding could handle. It is a small scene, but a telling one: a core maintainer working through a specific error message with someone trying to ship an application.

Dmytro Dzhulgakov speaking in an AI Engineer World's Fair video image
At the AI Engineer World’s Fair, the former contest programmer made a case for tailoring open models to the jobs they actually perform.

A familiar name sits near his in the Meta and PyTorch story: Lin Qiao, who led PyTorch at Meta and later became Fireworks AI’s chief executive. Fireworks brought together several engineers from PyTorch, Meta, and Google. It was founded in 2022, just as generative models were making the leap from impressive demonstrations to products that would need to run every hour of the day. Dzhulgakov became a co-founder and CTO.

2006

Silver medal at the International Olympiad in Informatics.

2010–2012

Topcoder and ICPC finals while studying applied mathematics in Kharkiv.

2018–2019

Research on production inference and recommendation systems at Meta.

2022

Fireworks AI begins with a team drawn from large-scale AI infrastructure.

2024–2026

Public talks and posts concentrate on serving, specialization, cost, and quality.

The problem after the demo

At a 2024 startup panel, Dzhulgakov described Fireworks as a developer platform for open generative models across text, image, and other forms. A team might start with a base model or customize one. Either way, somebody has to serve it, scale it, and keep its response time and cost within reason. Fireworks takes on that infrastructure work. His explanation was plain enough for a customer to recognize the problem: the interesting model is not necessarily an application yet.

He also described two audiences for a developer product. The person using the tool notices the small details. The business buying it cares about the end use case. A good infrastructure company has to satisfy both. That observation adds a human dimension to what could otherwise sound like a conversation about GPU kernels. Engineers have to enjoy the tool; organizations have to see the result. Neither audience can be waved away with a benchmark chart.

Dzhulgakov’s talks return to specialization. A customer-support assistant, a coding tool, and a financial research workflow have different demands. Some can wait for a batch result; others need a response in the rhythm of a conversation. A model trained or tuned for one job may be more useful than a general model that carries capabilities the application never calls upon. But the model is only one part of the system. Retrieval, outside APIs, structured outputs, caching, and the serving stack all affect what a user experiences.

The application path he keeps describing
Real task
Suitable model
Serving setup
Measured result
A compact reading of his public talks: match the model and infrastructure to the workload, then judge the whole application.

This is where his competition background becomes a useful metaphor, with one important change. In a contest, everybody receives the same problem and the same clock. In production AI, every customer brings a different clock. A voice interface may care intensely about the first fraction of a second. A batch operation may value throughput. An agent that calls tools repeatedly may make dozens of requests before it finishes one user task. The engineering problem is to find the constraint that actually governs each case.

His recent posts make that concrete. In June 2026, he cautioned that a single speed measurement can mislead, because one point in a benchmark may wander and may not match a customer’s workload. In July, he wrote about how repeated cached prompts shape the cost of agent tasks. If an agent makes fifty tool calls, a growing conversation can be billed again and again. The price printed next to a model name is only a starting point; the pattern of use decides the bill.

The launch that waited

In August 2026, Dzhulgakov offered an unusually detailed account of a delayed Fireworks launch. The company was preparing to serve a new model, GLM-5.3-Flash, and found a benchmark difference it could not explain. On reasoning-heavy tests, some open-source serving setups produced roughly twice as much thinking as the model maker’s API, even while scores looked similar. Longer reasoning can mean more tokens, more waiting, and potentially different behavior when a limit is reached.

Fireworks held the public release, opened a private preview for some customers, investigated, and reran benchmarks as more information became available. It launched after the results made sense. In his telling, the delay was a quality decision. The episode is modest as corporate drama, but revealing as engineering practice. A number that looks good can still conceal a question. If the question affects what customers will receive, the clock should wait.

“Quality trumps hype.”
Dzhulgakov, explaining the August 2026 launch decision

His work now extends beyond Fireworks. Dnipro VC lists him as an AI advisor in San Francisco, a public connection to a network of founders with Ukrainian ties. It is one more line in a career that began in Kharkiv and now operates across the global AI industry. He still speaks and writes as an engineer: curious about models, attentive to what can break, and willing to ask whether an impressive measurement survives a more realistic test.

The arc from Lego to the Intel 8080 to PyTorch and Fireworks has an appealing neatness, but its real substance is less tidy. It is years of code that had to run, contests that could be lost to a hidden flaw, and services that people will judge by what happens after the demo. Dzhulgakov’s arena is larger now. The task remains familiar: understand the problem, respect the clock, and make the answer work.