Jiang Daxin asked ChatGPT how old it was. This sounds like the sort of question one asks to fill a silence at a dinner party, but Jiang had spent years making machines converse. He knew the traps. Older chatbots often fetched an answer or followed a scripted rule. When he asked the follow-up, how old will you be next year, he was testing whether the system could work out what “next year” meant and carry the calculation through. It answered. Jiang’s first reaction, he later recalled, was suspicion: surely it was cheating. Then he went back to the papers.
At the time, he had a very good job to go back to. Sixteen years at Microsoft had taken him from research to a global vice presidency. His work touched Bing, Cortana, Azure and Microsoft 365. He knew what it took to make a clever idea survive contact with an actual product. Yet the little exchange on a screen made the existing map feel incomplete. By the 2023 Lunar New Year, he had decided to leave and build a company of his own.
One university friend had given him an image for such decisions. Many winds blow through the world, the friend said, but one may come to your doorstep only once or twice in a lifetime. Jiang was more cautious than the advice suggested. He had watched other technological enthusiasms come and go without wanting to lead a company around them. This one made him reach for the research and then for a team.
The ten minute distance
The walk from Microsoft’s Beijing office to StepFun’s early office was about ten minutes. It is a neat piece of geography for a change that had taken a working lifetime. Jiang entered the University of Science and Technology of China in 1993 and earned his bachelor’s degree in 1998. He later completed a computer science doctorate at the University at Buffalo. In 2007 he joined Microsoft Research Asia, where his interests in search, language and data moved steadily closer to products used at scale.
By 2017 he had become a Microsoft global partner, serving as deputy dean and chief scientist in its Asia internet engineering organization. In March 2023 he became a corporate vice president. The title lasted only a short while before he left. A reader under an early news report wondered why someone doing so well at Microsoft would start again. Jiang liked the question enough to put it on a presentation slide. It was funny because it was sensible.
He had an answer, though not one that would fit neatly on the slide. At a large company, a person can drive a technology a very long way inside an established strategy. Jiang wanted to decide what to build from the beginning. He had also learned something from that first ChatGPT test: reading about a capability is different from making it work. He could explain the papers. He wanted the experience of building the machine.
“I think running with the door closed might let us run more freely.”Jiang Daxin, in a 2024 interview; translated from Chinese
An iron triangle before a product
His first problem was people. A large model needs more than a theory of language. Jiang described the founding disciplines as an “iron triangle”: algorithms, data and systems. He could lead algorithms. Jiao Binxing, a search ranking leader from Microsoft, brought knowledge of data. Zhu Yibo, who had worked on large computing systems, took on the machinery that lets a model train at scale. They joined around a goal ambitious enough for Jiang to say aloud: pursuing general artificial intelligence.
The triangle also reveals what Jiang carried from his previous life. Search had taught his team to sort the useful internet from the noisy one. Data needed cleaning, ranking and a sense of which sources deserved trust. Systems work required a different kind of patience, one concerned with memory, communication and a fleet of computers that will occasionally fail. The glamorous part of AI is the answer appearing on a screen. The less glamorous part is arranging that the answer can appear again tomorrow.
StepFun began training its first large language model, Step-1, on July 1, 2023. The team completed that initial training run by late August, according to Jiang. A multimodal model, Step-1V, followed in November. Yet the company remained comparatively quiet. Jiang said the roadmap was clear and the team had no need to gather attention before it was ready to show its work. In March 2024 it presented a preview of Step-2. The doors had opened, at least a little.
The staircase on the door
Visitors to Jiang’s office noticed a small, hand-drawn symbol on the door. It was his sketch for the company logo, shaped by the step function, an early building block in neural networks. Seen as a line, the function rises in a sudden stair. It gave StepFun both a name and a visual joke for a company obsessed with what the next step might be.
Jiang has long argued that the next steps would involve more than text. Human interaction mixes language, sound and images without asking permission to switch modes. StepFun built models across text, audio, images and video, and explored systems that could both understand an input and generate an output. In a 2025 discussion, Jiang described this combination of understanding and generation as an important unsolved problem in visual AI. The ambition sounds abstract until a device must listen, see, respond and continue a task without dropping the thread.

At a February 2025 event, he stood on a bright blue stage to talk about agents, software that can carry out tasks across several steps. This was a considerable change of scene for a man who described himself as an “extreme introvert” in an earlier interview. The stage did not turn him into a showman. His recurring subjects were still the mechanics: perception, reasoning, models and their use in phones, cars and other devices.
By July, StepFun was introducing Step 3 and working with chip makers on how models run. The company presented a model and chip alliance with Chinese infrastructure firms. Jiang’s point was practical. A model that performs well in a research setting must also work economically on the hardware available to its users. This is where training, deployment and product decisions meet, and where a founder trained in both research and engineering has to keep several clocks running at once.
When the window opens
Jiang offered a blunt description of technical advantage at that July event. Technology, he said, creates a window of opportunity rather than an everlasting moat. A company can use that window to win a first user. If the application is useful enough, the user returns, the product improves and a stronger position may develop. The researcher who left to build models was talking about a decidedly unromantic test: whether somebody chooses the thing twice.
That view also helps explain StepFun’s move toward the places where AI meets daily activity. Voice inside a car is a different challenge from a polished chatbot demonstration. An assistant on a phone has to respond quickly and cope with the limits of the device. A video model has to do more than produce a lovely isolated clip. Each application places a price on delay, reliability and cost. Jiang’s old product experience becomes relevant again every time a model leaves the lab.
“Technology is a window of opportunity, not a permanent moat.”Jiang Daxin, in a 2025 interview; translated from Chinese
The company itself kept changing. In January 2026, StepFun appointed Yin Qi, the former Megvii chief executive, as chairman and announced a 5 billion yuan Series B+ funding round. Jiang remained chief executive. The new arrangement placed an experienced device and computer vision entrepreneur alongside the founder as StepFun sharpened its interest in AI that works in cars and other physical settings. For Jiang, it was another version of the founding triangle: combine people whose strengths solve different parts of the same problem.
In February 2026, StepFun released Step 3.5 Flash as an open model for reasoning and agent tasks. Its published design activates only a portion of its total parameters for each token, an attempt to deliver useful capability without paying the full computational cost at every turn. The release was a long way from the quiet first training run in 2023, but the core question had stayed familiar. Could the team make a model that could do something, and could it make that model practical enough to use?
The question after next year
Jiang was named an ACM Distinguished Member in 2025, a formal nod to the research career that preceded StepFun. A university alumni notice marked the path from his undergraduate years to work on Bing, Cortana and Azure. Still, the most vivid picture of him may be the small one: a technically trained executive testing a chatbot with a childlike question, refusing to take the surprising answer at face value, then rereading the papers to work out what had happened.
There is another image, too: the drawn staircase on an office door. It suits a career in which each apparent leap followed years of preparation. Search led to data. Data led to language models. Models led to sound, vision and applications. The company name promises a step; Jiang’s own story suggests that taking one requires knowing exactly what is underfoot. The wind reached his door in 2023. What he built after opening it has been an exercise in seeing how far that first step can carry a team.