On his personal website, Yang Zhilin has left a remarkably economical account of his present occupation: “I am working on a startup.” The page offers a dense list of papers, collaborators and conference appearances, then that modest sentence. It is the sort of understatement available only to someone whose startup has become rather difficult to overlook. Moonshot AI, founded in Beijing in 2023, makes Kimi, a family of AI systems that began by making long documents feel less like a burden for a chatbot. The product carries Yang’s nickname. The company’s Chinese name carries a Pink Floyd reference. Between those two names sits a career spent trying to give machines a longer memory.
Yang’s route to Kimi was unusually direct in one respect. Long before anyone was measuring an assistant’s context window in hundreds of thousands of tokens, he was publishing research on what happens when language models lose the thread. His 2019 papers included Transformer-XL, which tackled the fixed-length limit that forced models to forget earlier text, and XLNet, another influential approach to learning from language. These are technical documents, but their human premise is easy to grasp. A conversation changes meaning when the listener can remember what came before.
A drummer takes the long route into code
Yang was born in Shantou, in Guangdong province, in 1992. At Tsinghua University he began in thermal engineering and moved to computer science. Accounts of that turn include an unlikely nudge: a Haruki Murakami novel with a programmer whose work appealed to him. The anecdote survives because it punctures the usual origin myth. A future AI founder did not need a neat childhood prophecy. Fiction helped him imagine what writing software might feel like.
Tsinghua also gave him two things that remained in his story. He studied with Jie Tang, who later became associated with another Chinese AI company, Z.ai. And he played drums in a campus band, Splay, with fellow students. Yang’s fondness for Pink Floyd is more than a convenient profile detail. Moonshot’s Chinese name, Yue Zhi An Mian, refers to the dark side of the moon. The album title made its way from a record sleeve into a company name. Silicon Valley has produced stranger naming stories, though few with such a good rhythm section.
The musical and technical parts of his life should not be made to explain each other too neatly. A drummer counts time; a language model uses context. The tempting metaphor is also too easy. What the overlap does show is Yang’s appetite for work that has structure and invention in the same frame. At Tsinghua, his academic work was already moving from networks toward language. By 2015 he had a bachelor’s degree and a place at Carnegie Mellon University, where a much longer question awaited him.
The problem of a machine that forgets
At Carnegie Mellon, Yang worked with Ruslan Salakhutdinov and William W. Cohen, earning a PhD in language and information technology in 2019. He also spent time with Meta AI and Google Brain. The names suggest proximity to powerful research groups, but the published work says more. Yang co-authored HotpotQA, a dataset designed around questions that require reasoning across several pieces of information. He also helped write papers on language models, including Transformer-XL and XLNet. Across them runs a concern with how a system carries information forward and uses it.
A short context window is like reading a novel while someone quietly tears away every earlier page. A system may respond fluently to the sentence in front of it yet fail to use the chapter that gives that sentence its meaning. Transformer-XL proposed a way to reuse representations from earlier segments, letting the model reach beyond a single fixed block of text. XLNet addressed pretraining from a different angle. Yang was an equal contributor on both papers, working among collaborators whose names recur throughout modern machine learning. The research was collective, as good research usually is, and his later company would make its own collective work visible.
There was a first startup before Moonshot. Yang co-founded Recurrent AI while still pursuing his doctorate; it applied language technology to business conversations. It was a bridge between the laboratory and the messy world of customers, sales calls and tools people must actually use. He returned to China after graduating. Moonshot would arrive in 2023, after the public release of ChatGPT had made large language models a commercial race as well as a scientific one. Yang had the background to join that race, and a specific view of where a smaller entrant could begin.
The first bet was a longer conversation
Moonshot’s first assistant, Kimi, arrived in October 2023 with long text as its distinctive feature. At the time, the company emphasized its ability to process 200,000 Chinese characters. The figure mattered because it turned an abstract research interest into a user action: hand the assistant a long paper, a collection of documents or a crowded set of notes, then ask for help with the whole thing. Yang had argued that context was more than a larger input box. He saw it as a path toward systems that could understand a person’s instructions and history over time.
That ambition becomes clearer in his interviews. He compared long context with memory in a computer and connected it to personalization. An assistant that sees only today’s prompt starts from scratch. One that can use relevant past interactions might respond with a better sense of the task and the person giving it. There is an obvious difficulty: holding more text does not guarantee finding the important sentence, interpreting it correctly or earning the trust needed for someone to share it. Yang recognized the distinction. “You can’t just look at the numbers,” he said when discussing context windows.
For a time, the question was whether Kimi’s long-document appeal could become a durable company. Moonshot needed capital and computing power, as every frontier model lab does. It also needed a reason for people to return after the first impressive upload. Yang’s public description of an AI product put weight on personalized interaction rather than a single demonstration. That is a harder target than a benchmark score. It asks a model to become useful in the unglamorous middle of work, where documents contradict one another, instructions change, and the right answer may be buried on page 47.

From an assistant to open weights
Kimi’s story widened in 2025. Moonshot released Kimi K2, an open-weight model, making the model’s parameters available to developers under its terms. Kimi K2 Thinking followed later that year. In January 2026 came Kimi K2.5, with stronger visual and agent-style capabilities; Yang spoke at NVIDIA GTC in March about how the team scaled it. The company’s research calendar began to look less like a sequence of chatbot updates and more like a set of experiments in what models can do over long tasks.
Kimi K3, announced in July 2026, gave that progression a striking number: 2.8 trillion parameters. It also brought native vision and a one-million-token context window. Moonshot released the weights and technical report later that month. A figure this large is useful as a signpost, though the company’s own technical account does not reduce its work to size. It describes new attention mechanisms and other changes meant to use computing power more effectively. Bigger models can still be forgetful or careless. The old question remains whether information survives the journey from input to answer.
The open-weight turn changed who could take part in the experiment. Researchers and developers outside Moonshot could inspect, run and build on Kimi models, subject to licenses and the formidable hardware required for the largest releases. That choice placed Yang’s work in a global conversation about access to capable AI. It also made claims about performance easier to test beyond the company’s own demonstrations. A public model meets a wider variety of users, and users have a gift for finding tasks no release note anticipated.
What a longer memory is for
The temptation in AI profiles is to turn every technical milestone into destiny. Yang’s path is more interesting without that shortcut. He began with research questions about language and memory. He worked at two large research organizations, co-founded a business software company, and then built Moonshot with Tsinghua peers. The assistant they named for him has changed shape repeatedly. None of those turns requires the conclusion that the problem is solved. Each is a new attempt to learn what a model should keep, what it should forget and what it should do with what remains.
There is also a human-scale question inside the engineering one. A useful memory is selective. People do not want a machine to recite every line they have ever written; they want it to notice the constraint that matters now. Yang’s 2024 remarks about trust acknowledge the other side of that exchange. A system can only use personal context people choose to give it, and it has to reward that choice with something more useful than generic familiarity. That is a product challenge as much as a model challenge.
The name Moonshot invites grand claims. Yang’s better image was smaller and stranger: moving toward snow-covered mountains without knowing what lies inside them. It captures the uncertainty behind a field that often advertises certainty. His website’s spare sentence works in the same way. “I am working on a startup” leaves out the scale, the funding and the headlines. It keeps the verb. For Yang Zhilin, the work has always been a question about the distance between seeing more and understanding more. Kimi has made that question public, and the answer is still being written.