This was my first interview with Yang Zhilin, conducted in early 2024 and published on March 1, 2024—exactly the first anniversary of Kimi's founding. At the time, Kimi had only 80 people, working out of their first, somewhat run-down office. There was no logo at the entrance. Only a white piano standing guard by the door.
This article generated quite a stir in China's tech community at the time.
Back then, I was still a print journalist, so this interview exists only as text and an audio podcast.
As we can see, many of the views expressed in this article have since been borne out.
Just rereading these words from two years ago, you really can't help but marvel at how dramatically the world has changed!
(This translation was generated by Kimi K3.)
Yang Zhilin: “If everyone thinks you are normal—if your dream is one that anyone could have—it adds nothing to the sum total of humanity’s dreams.”
By Zhang Xiaojun
Just one year ago, AI scientist Yang Zhilin did a precise calculation in Silicon Valley. He realized that if he decided to launch a foundation-model startup aimed at AGI, he would need to raise more than $100 million within the next few months.
Yet that was merely a ticket to the game. A year later, that figure had multiplied thirteenfold.
For foundation-model companies, competition is less a scientific contest than, first and foremost, a brutal contest of money. With investors holding their purse strings tight, you have to stay ahead of your rivals in raising more money, buying more GPUs, and grabbing more talent.
“It requires a concentration of talent and a concentration of capital,” says Yang Zhilin, founder and CEO of Moonshot AI, the foundation-model company established on March 1, 2023.
Over the past year, Chinese foundation-model companies have seemed to live on a tense, constricted edge of survival. On the surface, each of them holds hefty sums of cash. But on one hand, they must immediately pour freshly raised money into extremely costly research to chase OpenAI—first catching up to GPT-3.5, and before GPT-4 is even reached, along comes Sora. On the other hand, they must race nonstop to find viable real-world use cases, to validate for themselves that they are companies, not research institutes that only devour capital. And that is not enough: for every one of these ventures, whether the exit is an IPO or an acquisition, the way out remains anything but clear.
Among the founders of China’s foundation-model companies, Yang Zhilin is the youngest, born in 1992. The industry describes him as a staunch AGI believer and a founder with rare technical charisma. Much of his academic and professional record is tied to general-purpose AI, and his papers have been cited more than 22,000 times.
In mid-2023, China’s tech community turned abruptly from euphoria to chill on foundation models, and accelerating real-world deployment became the pragmatic mainstream melody. This inevitably left foundation-model CEOs torn violently between ideal and reality. In a Chinese AI ecosystem where everyone chants PMF (product/market fit) and everyone chants commercialization, this founder—an AI researcher by training—is in no particular hurry.
With 80 people, Moonshot AI has the smallest headcount among China’s leading foundation-model companies. Unlike his rivals, Yang did not opt for the safer B2B business or seek deployment in verticals such as healthcare or gaming. He built one—and only one—consumer product: the AI assistant Kimi, which accepts inputs of up to 200,000 Chinese characters. Kimi is also Yang Zhilin’s English name.
Yang prefers to see his company as a system that combines science, engineering, and business. You might picture it this way: above the human world, he is erecting an AI laboratory bench—with one hand he runs experiments, and with the other he brings cutting-edge technology down into the real world, discovering applications through interaction with people and delivering those applications into consumers’ hands. Ideally, the former burns through billions and tens of billions of dollars of capital; the latter earns that money back hundreds or thousands of times over. However you hear it, it sounds as thrilling—and as perilous—as walking a tightrope.
“AI is not about what PMF I can find in the next year or two; it’s about how to change the world over the next ten to twenty years,” he says.
Such abstract, idealistic thinking makes one sweat nervously on his behalf: can a young AI scientist carve out room to survive in a realist China?
In February 2024, Moonshot AI closed a large funding round against the market tide. It is understood that the company raised a Series B of more than $1 billion at a $1.5 billion pre-money valuation, led by Alibaba with follow-on participation from Monolith Management, Xiaohongshu, and others. Upon completion of the deal, Moonshot AI’s post-money valuation stood at roughly $2.5 billion—making it, at this stage, the highest-valued unicorn in China’s foundation-model race. (The company declined to respond to or comment on the matter.)
In the midst of this third funding round, we sat down with Yang Zhilin to talk about his first year of entrepreneurship—a cross-section, in miniature, of a year in which Chinese foundation-model companies raced ahead from the starting line.
His company did not set up in Sohu Network Plaza in Beijing, the hub where foundation-model companies cluster. For a company with total funding of about RMB 9 billion, this office in the Liangzi Xinzuo building looks crude and run-down. There is not even a company logo at the entrance—only a white piano standing guard by the door.
The meeting room sits in a corner; with its small windows it is dark inside, and the heater hums as it blows warm air against the winter cold. In the dim light, Yang describes how the past year has felt to him: “It’s a bit like driving down a road with a range of snow mountains stretching out ahead. You don’t know what’s inside them. You just keep walking forward, one step at a time.”
Below is the full interview with Yang Zhilin. (For readability, the author has made some textual edits.)

This photo was taken in early 2024 at Kimi’s first office in Beijing. They’ve moved out of this location now. No logos lined the entrance — only a white piano stood quietly by the door.
**
Part 1
Standing at the Beginning
“You Have to Ride the Wave”
**
Zhang Xiaojun: How have you been lately?
Yang Zhilin: Busy—there’s a lot going on. But I’m still excited. We’re standing at the very beginning of an industry, and there is enormous room for imagination.
Zhang Xiaojun: When I came in just now, I saw a pure white piano at your company’s entrance.
Yang Zhilin: There’s a Pink Floyd album sitting on it, too. I have no idea who put them there—I suddenly noticed them a couple of days ago and haven’t had a chance to ask. (Pink Floyd is the British rock band that released the album The Dark Side of the Moon.)
Zhang Xiaojun: On the day ChatGPT was released in November 2022, what were you doing?
Yang Zhilin: I was already preparing for this—recruiting people, building a team, exchanging new ideas. Seeing ChatGPT was thrilling. Three to five years earlier—even in 2021—it would have been inconceivable. That kind of higher-order reasoning had been very hard to achieve.
I sensed that many variables were about to shift in the market: capital on one side, talent on the other—the core factors of production for AI. If those variables fell into place, it would become possible to build a proper company to do this—an organization built for AGI could go from 0 to 1. That was a major epiphany. An independent company made more sense, but it wasn’t something you could do the moment you wanted to; ChatGPT jolted the variables and brought the factors of production together. You have to ride the wave.
Zhang Xiaojun: After you decided to found an AGI company, what preparations did you make? How did you assemble the two factors of production—capital and talent?
Yang Zhilin: It was a winding process. ChatGPT took time to diffuse. Some people learned of it early, some late; some doubted at first, then were shocked, then became believers. Finding people and finding money were tightly bound to timing.
We began focusing on our first funding round in February 2023. Had we delayed to April, we basically would have had no chance. But doing it in December 2022 or January 2023 wouldn’t have worked either—the pandemic was still on, and people hadn’t processed it yet. So the real window was just one month.
One night in the United States, I did a precise calculation. When I finished, I concluded we needed to raise at least $100 million within a few months. Many in the market hadn’t started fundraising yet, and many didn’t believe you could necessarily raise that much. But it turned out to be possible—even more than that.
The talent market started moving, too. Inspired by ChatGPT, many people had this realization in March or April 2023: this is the only thing worth doing in the next decade. You have to reach out to the right people at the right time. A year or two earlier, talent would not have clustered to this degree. Back then, more people were doing traditional AI or AI-adjacent businesses—none of it was general-purpose AI.
Zhang Xiaojun: To sum up: February was the window for fundraising, and March and April were the window for hiring?
Yang Zhilin: More or less.
Zhang Xiaojun: That night in the U.S.—where were you when you did this math? How exactly did you calculate it?
Yang Zhilin: From late 2022 into early 2023, I spent a month or two in the U.S., talking to people. I did it where I was living. You work out how many FLOPs you need, the training cost, inference, and the user numbers.
Zhang Xiaojun: At that moment, what was the prevailing mood in Silicon Valley?
Yang Zhilin: The product began picking up many early adopters, concentrated in the tech circle. We were in that circle ourselves, so we felt it more keenly. At the big Silicon Valley companies, people have to write performance reviews every six months, and many started writing them with ChatGPT. Some people whose writing was usually not that professional turned in reviews written with ChatGPT, and everyone sounded dead serious.
Undercurrents were stirring. Many people were thinking about their next job or about starting a company. Quite a few friends who talked with us later went off to found startups. And there was intense FOMO—fear of missing out. Nobody could sleep. Whether it was midnight, 1 a.m., or 2 a.m., if you reached out, people were always there. A bit anxious, a bit FOMO, and very excited.
Zhang Xiaojun: The night you calculated you needed to raise $100 million—how late did you stay up?
Yang Zhilin: It was fine—the calculation itself didn’t take long. But afterward, I couldn’t tell too many people. If I had, no one would have believed it could be done.
Part 2
Technical Lineage
“Free Yourself from Endless Carving”
Zhang Xiaojun: When the venture capital world talks about you, they say, “The founder is brilliant, has technical charisma, and the team is full of technical stars.” So before we discuss your foundation-model venture, I’d like to start with your academic background. You studied computer science at Tsinghua as an undergraduate and earned your PhD at Carnegie Mellon’s School of Computer Science. Has AI always been your focus?
Yang Zhilin: I was born in 1992 and started my undergraduate degree in 2011. From my sophomore year to now—more than a decade—I’ve been in this field. At first I explored more divergently, looking around everywhere; I did some work related to graphs and to multimodality. In 2017, I converged on language models. At the time I felt language models were a relatively important problem; later I came to feel it was the only important problem.
Zhang Xiaojun: In 2017, how did the AI industry generally understand language models, and how did that understanding evolve?
Yang Zhilin: Back then it was a model used to rank speech-recognition results. (Laughs.) After a segment of speech was recognized, you’d get many candidate results, and you’d use the language model to see which one had the highest probability and output the most likely one. Its applications were very limited.
But you come to realize it’s a fundamental problem, because you are modeling the probabilities of the world. Language is limited, but it’s a projection of the world; in theory, if you make the token space—the space of all possible tokens—large enough, you can build a general world model. How everything in the world arises and develops can be assigned a probability. Every problem can be reduced to how to estimate probabilities.
Zhang Xiaojun: Your academic mentors are very prominent: your PhD advisors were Ruslan Salakhutdinov, head of AI at Apple, and William W. Cohen, chief scientist of Google AI. Both straddle industry and academia.
Yang Zhilin: In previous years, industry and academia came together more, but the trend is now shifting: more valuable breakthroughs will happen in industry. That’s an inevitable law of development. It starts with exploratory research and gradually shifts into a more mature industrialization process. That doesn’t mean research is unnecessary during industrialization—only that pure research will struggle to produce valuable breakthroughs.
Zhang Xiaojun: What did you learn from these renowned mentors?
Yang Zhilin: I learned the most at Google, where I interned for a long time. I began working on Transformer-based language models in late 2018. My biggest learning was freeing myself from endless “carving”—the obsessive refinement of surface details. That was crucial.
You should look at what the big direction is, the big gradient. When ten roads lie before you, the average person worries about how to brake for a pedestrian ahead on this one road—short-term details. But which of the ten roads to take is what matters most.
This field previously had exactly that problem. For example, on a dataset of only one or two million tokens, you’d look at how to push perplexity lower, how to push loss lower, how to improve accuracy—and you’d fall into endless carving. People invented many bizarre architectures; these were carving tricks. After carving, you might do better on that kind of dataset, but you miss the essence of the problem.
The essence is analyzing what the field is missing. What is the first principle? Why can the scaling law serve as a first principle? You only need to find a structure that satisfies two conditions: first, it is sufficiently general; second, it is scalable. General means you can model all problems within this framework; scalable means that as long as you pour in enough compute, it keeps getting better.
This is the thinking I learned at Google: if something can be explained by something more fundamental, you shouldn’t over-carve at the upper layers. There’s an important line I strongly agree with: if you can solve a problem with scale, don’t solve it with a new algorithm. The greatest value of a new algorithm is in how it lets you scale better. When you free yourself from carving, you can see much more.
Zhang Xiaojun: Was Google also a follower of the scaling law back then? How did it implement first-principles thinking?
Yang Zhilin: Many such ideas already existed there, but Google didn’t implement them especially well. It had this way of thinking, but it couldn’t organize itself into a true moonshot. It was more like: here are five people pursuing my first principles, and over there five people pursuing theirs. There was nothing top-down.
Zhang Xiaojun: During your PhD, you published papers in collaboration with Turing Award winners Yann LeCun and Yoshua Bengio—and you were first author on those papers. How did those collaborations come about? What I mean is: they’re Turing Award laureates and they weren’t your advisors—what did you rely on to attract them?
Yang Zhilin: Academia is very open. As long as you have a good idea and a meaningful problem, it’s fine. What two brains—or n brains—produce is more than one brain alone. This applies when developing AGI, too. An important strategy in AI is called “ensembling”—using multiple different models or methods and combining their predictions for better performance. It’s essentially doing the same thing: when you have diverse viewpoints, you can spark many new things. Collaboration is hugely beneficial.
Zhang Xiaojun: Would you first have an idea and then ask them whether they were interested?
Yang Zhilin: That’s roughly how it went.
Zhang Xiaojun: Which is harder: winning over academic heavyweights in research, or winning over capital heavyweights in fundraising? What are the similarities?
Yang Zhilin: “Winning over” isn’t a good phrase—the essence behind it is cooperation. Cooperation means both sides win, because mutual benefit is the precondition for cooperation. So there’s really no difference: you need to offer others unique value.
Zhang Xiaojun: How do you earn their trust? What do you think your gift is?
Yang Zhilin: There’s no particular gift—just working hard.
Part 3
The Old System No Longer Works
“AGI Needs a New Kind of Organization”
**
Zhang Xiaojun: You just said “more valuable breakthroughs will happen in industry”—does that include startups and the giants’ AI labs?
Yang Zhilin: Labs are history. Google Brain used to be the biggest AI lab in industry, but it was a research organization embedded inside a big company. That kind of organization can explore new ideas, but it’s very hard for it to produce a great system—it could produce the Transformer, but it couldn’t produce ChatGPT.
The way development now evolves is that you’re building an enormous system, which requires new algorithms, solid engineering, and even a lot of product and commercialization work. It’s like the early 2000s: you couldn’t research information retrieval in a lab; it had to live in the real world, as a huge system, a product with users—like Google. So research and education systems will shift their function toward primarily cultivating talent.
Zhang Xiaojun: How would you describe this new form of system? Is OpenAI its prototype?
Yang Zhilin: It’s the most mature organization of this kind today, and it’s still gradually evolving.
Zhang Xiaojun: So it can be understood as an organization established for humanity’s grand scientific goals?
Yang Zhilin: I want to emphasize: it is not pure science—it’s a combination of science, engineering, and business. It has to be a commercial organization, a company, not a research institute. But this company is built from zero to one, because AGI needs a new kind of organization. First, the mode of production differs from the internet era; second, it shifts from pure research to a combination of research, engineering, product, and business.
At its core, it should be a moonshot program, with a great deal of top-down planning—yet within that planning there is room for innovation, because not all the technology is predetermined. Bottom-up elements exist within a top-down framework. Such an organization didn’t exist before, but the organization must adapt to the technology, because technology determines the mode of production; if they don’t match, you can’t produce effectively. We believe it will very likely need to be redesigned from scratch.
Zhang Xiaojun: During last year’s OpenAI boardroom coup, one option for Sam Altman was to join Microsoft and lead a new Microsoft AI team. What is the essential difference between that and being CEO of OpenAI?
Yang Zhilin: You’d have to grow a new organization inside an old culture—and that is extremely difficult.
Zhang Xiaojun: You want to build “China’s OpenAI”—can we put it that way?
Yang Zhilin: Not quite accurate. We don’t want to be China’s anything, and we don’t necessarily want to be OpenAI.
First, real AGI will definitely be global. There is no such thing—at least not long-term—as an AGI company confined to some regional market because of market-protection mechanisms. Globalization, AGI, and having a product with a very large user base: these three are ultimately necessary conditions.
Second, should it be OpenAI? If you look at 2017–2018, OpenAI had a terrible reputation. When people in our circle looked for jobs, they generally considered places like Google. Many people who talked with Ilya Sutskever, OpenAI’s chief scientist, came away thinking the man was crazy and far too full of himself—OpenAI was either madmen or scammers. But they committed very early, found the non-consensus, and found what is now the only first principle that works in AI: scaling through next-token prediction.
I believe there will be a company greater than OpenAI. A truly great company can combine technological idealism with a great product, co-creating with its users—AGI will ultimately be something produced by co-working with all of its users. So it’s not only about technology; it also requires pragmatism and real-world pursuits—ultimately, a perfect combination of the two.
Still, we should learn from OpenAI’s technological idealism. If everyone thinks you’re normal—if your dream is one that anyone could have—it adds nothing to the sum total of humanity’s dreams.
Part 4
The Moonshot’s First Step Is “Long Context”—What’s the Second?
“Two Big Milestones Are Coming Next”
**
Zhang Xiaojun: Back to the moment you decided to start the company—did you launch the first funding round immediately after returning to China?
Yang Zhilin: It began in the U.S. in February (last year), some of it remotely. In the end, domestic investors made up the majority.
Zhang Xiaojun: Did the first round raise $100 million?
Yang Zhilin: The first round wasn’t that much; later rounds exceeded that figure. We completed two rounds in 2023, totaling nearly RMB 2 billion.
This is now the third round. We haven’t formally announced the financing, so I can’t comment at this time.
Zhang Xiaojun: Some people say that since the second half of 2023, no one has been willing to invest in foundation-model companies anymore. Are they wrong?
Yang Zhilin: There still are. You can indeed see the shift in sentiment, but it’s not that no one is investing—at least for now, there’s quite a lot of investment interest in the market.
Zhang Xiaojun: Besides capital and people, what other key decisions did you make in 2023?
Yang Zhilin: Deciding what to do. That’s the advantage of companies like ours—having a technical vision for decisions at the highest level.
We do long context. That requires judgment about the future: you need to know what is fundamental and where things are heading next. Again, it’s first principles—the process of “de-carving.” If you focus on carving, you can only look at what OpenAI has already done and figure out how to reproduce it.
You’ll find that doing lossless long-text compression in Kimi gives the product a unique experience. When you read English-language papers, it helps you understand them remarkably well. Using Claude or GPT-4 today, you won’t necessarily do as well; this required laying the groundwork in advance. We worked on it for over half a year. That’s very different from spotting a long-context trend today, hastily assembling two teams, and developing it at maximum speed.
Of course, the marathon has only just begun; more differentiation will come, and that requires you to anticipate in advance what counts as “a non-consensus that holds true.”
Zhang Xiaojun: In what month was this decision made?
Yang Zhilin: February or March—it was decided as soon as the company was founded.
Zhang Xiaojun: Why is long context the first step of the moonshot?
Yang Zhilin: Because it’s fundamental. It is the new computer’s memory.
The old computer’s memory grew by several orders of magnitude over the past few decades, and the same thing will happen with the new computer. It can solve many of today’s problems. For example, current multimodal architectures still need a tokenizer, but with a losslessly compressed long context, you don’t need one—you can put the raw input in directly. Taken further, it’s the foundation for making the new computing paradigm more general.
The old computer could represent everything with 0s and 1s; everything could be digitized. But today’s new computer can’t yet—there isn’t enough context, so it isn’t that general. To become a general world model, you need long context.
Second, it enables personalization. AI’s core value is personalized interaction; the value ultimately lands on personalization, and AGI will be more personalized than the previous generation of recommendation engines.
But personalization isn’t achieved through fine-tuning—it’s achieved by supporting very long context. Your entire history with the machine is context, and that context defines the personalization process. It cannot be replicated, and it makes for more direct dialogue—dialogue that generates information.
Zhang Xiaojun: How much room is there to scale this up?
Yang Zhilin: Enormous. On one hand, expanding the window itself still has a long way to go—several orders of magnitude.
On the other hand, you can’t only expand the window, and you can’t just look at the number; whether the window is a few million tokens or several billion today is meaningless in itself. You have to look at the reasoning ability it enables within that window, the faithfulness—fidelity to the original information—and the instruction-following ability. You shouldn’t chase a single metric; you have to combine metrics with capabilities.
If these two dimensions keep improving, you can do a great deal. It could follow an instruction tens of thousands of words long, and the instruction itself could define many agents—highly personalized.
Zhang Xiaojun: Are the technologies behind long context and catching up with GPT-4 reusable for each other? Are they the same thing?
Yang Zhilin: I don’t think so. It’s more about adding a new dimension—a dimension GPT-4 doesn’t have.
Zhang Xiaojun: Many people say the leading Chinese foundation-model companies are all doing roughly the same thing—chasing GPT-3.5 in 2023, chasing GPT-4 in 2024. Do you agree?
Yang Zhilin: Improving general capabilities certainly has key milestones, so that statement is right to a degree—as a latecomer, you inevitably go through a catching-up process. But it’s also one-sided. Beyond general capabilities, there is a lot of space to develop distinctive capabilities and reach state-of-the-art in certain directions. Long context is one. DALL-E 3’s image generation is thoroughly outclassed by Midjourney V6. So you have to work on both fronts.
Zhang Xiaojun: What proportion of time and resources goes to general capabilities versus new dimensions?
Yang Zhilin: They have to be combined. A new dimension can’t exist apart from general capabilities, so it’s hard to give a direct ratio. But sufficient investment is required to do the new dimension well.
Zhang Xiaojun: Will all these new dimensions be carried by Kimi?
Yang Zhilin: Kimi is certainly a very important product for us, and we’ll have some other attempts as well.
Zhang Xiaojun: What do you make of the comment by Li Guangmi, founder of Shixiang, that the technological distinctiveness of Chinese foundation-model companies is still not very high today?
Yang Zhilin: I think it’s fine—we’ve already produced quite a lot of differentiation today. It’s a matter of time; this year you should see more dimensions. Last year, everyone was just putting up the scaffolding and getting things running first.
Zhang Xiaojun: If the moonshot’s first step is long context, what’s the second?
Yang Zhilin: There will be two big milestones ahead. First, a truly unified world model—one that unifies all the different modalities, a truly scalable and general architecture.
Second, enabling AI to keep evolving without human data input.
Zhang Xiaojun: How long will it take to reach these two milestones?
Yang Zhilin: Two to three years—possibly faster.
Zhang Xiaojun: So three years from now, we’ll already be looking at a world completely different from today’s.
Yang Zhilin: At the current pace of development, yes. The technology is now in a budding, fast-growing stage.
Zhang Xiaojun: Can you imagine what will exist three years from now?
Yang Zhilin: There will be a certain degree of AGI. Many of the things we do today, AI will also be able to do—even better than us. But the key is how we use it.
Zhang Xiaojun: And for you—for Moonshot AI as a company—what’s the second step?
Yang Zhilin: We will go do those two things. All the remaining problems are derived from these two factors. The reasoning and agents people talk about today are byproducts of solving these two problems. Some additional carving is needed, but there’s no fundamental blocker.
Zhang Xiaojun: Will you go all in on catching up with GPT-4?
Yang Zhilin: GPT-4 is a necessary stop on the road to AGI. The key is not to be satisfied with merely matching GPT-4. First, you have to ask what the real non-consensus is now: beyond GPT-4, what’s next? What should GPT-5 and GPT-6 look like? Second, you have to see which distinctive capabilities you have within that—and that matters more.
Zhang Xiaojun: Other foundation-model companies publish their model capabilities and rankings. You don’t seem to have done that?
Yang Zhilin: Chasing leaderboard rankings means very little. The best leaderboard is the users—you should let users vote. Many leaderboards have problems.
Zhang Xiaojun: Is being the fastest to reach GPT-4 among China’s foundation-model companies your goal? Does fast versus slow make a difference?
Yang Zhilin: Definitely. Over a long enough time frame, everyone will eventually get there. But it depends on how long your lead or lag is. A gap of six months or more is meaningful—and it also depends on what you can do with that window.
Zhang Xiaojun: When do you expect to reach GPT-4?
Yang Zhilin: It should be quite soon, but I can’t disclose the specific timing publicly.
Zhang Xiaojun: Will you be the fastest?
Yang Zhilin: That has to be assessed dynamically—but we have a real chance.
Zhang Xiaojun: After launching Kimi, what is your North Star metric?
Yang Zhilin: Today it’s about making the product better and adding more dimensions. For example, we shouldn’t just be fighting tooth and nail over a search use case—search will later be only a small fraction of this product’s value; the product should have a much bigger increment. Being 10% or 20% better than a traditional search engine isn’t worth much—only something truly disruptive deserves the three letters “AGI.”
The unique value is your incremental intelligence. You have to hold onto this point: intelligence is always the core incremental value. If only 10%–20% of your product’s core value comes from AI, it doesn’t hold up.
Part 5
I’m Not at All Anxious About Commercialization
“User Scaling and Model Scaling Need to Happen at the Same Time”
**
Zhang Xiaojun: Mid-2023 was a huge watershed—the market turned from frenzy to chill very quickly. How did you perceive it?
Yang Zhilin: I don’t fully agree with that characterization—we did complete a funding round in the second half of the year. And new things kept coming out. Today’s model capabilities were unimaginable at the end of last year. The user numbers and revenue of more and more AI companies kept rising. It has continuously proven its value.
Zhang Xiaojun: For you, what felt different between the first and second halves of the year?
Yang Zhilin: Not much changed. Variables certainly exist, but you return to first principles—how to give users a good product. Ultimately, we must satisfy user needs, not win a race. We are not a company built for competition.
Zhang Xiaojun: The industry believes a notable difference between the first and second halves of 2023 was a shift of focus: the first half was more about AGI; the second half turned to how to land applications and commercialize. Did you make that shift?
Yang Zhilin: Of course I’m going to do AGI—it’s the only meaningful thing for the next decade. But that doesn’t mean we don’t build applications. Or rather, it shouldn’t be defined as an “application.”
“Application” makes it sound like you have a technology and you’re looking for somewhere to use it, with a commercial loop closed. But “application” isn’t the accurate word. It and AGI complement each other. It is both the means to achieve AGI and the purpose of achieving it. “Application” sounds more like a purpose: I want to make it useful. You have to combine Eastern and Western philosophies—you have to make money, and you have to have ideals.
Today, users help us discover scenarios we never considered. Someone uses it to screen résumés—something we never thought of when designing the product, but it naturally works. User input, in turn, makes the model better. Why is Midjourney so good? It scaled on the user side—user scaling and model scaling must happen at the same time. Conversely, if you only focus on applications and ignore the iteration of model capabilities—ignore AGI—your contribution will be limited.
Zhang Xiaojun: Zhu Xiaohu, managing partner at GSR Ventures, only invests in foundation-model applications. One of his views: the hardest core problem is PMF for AIGC—if ten people can’t find PMF, a hundred people won’t either; it has nothing to do with headcount or cost, so don’t burn money on it. He says, “Train on LLaMA for two or three months and you can at least reach the level of the top 30 humans—it can replace people immediately.” What do you think of his view?
Yang Zhilin: AI is not about what PMF I can find in the next year or two; it’s about how to change the world over the next ten to twenty years—these are two different ways of thinking.
We are staunch long-termists. When AGI or something stronger is achieved, everything today will be rewritten. PMF is certainly important, but if you rush to find PMF, you’ll very likely be hit by another “dimensionality-reduction strike”—being crushed by a higher-dimensional technology. That has happened too many times. In the past, many people built customer-service and dialogue systems, doing slot filling—some were companies of decent scale. But they were all wiped out by a higher-dimensional blow. It was painful.
That’s not to say the approach never works. Suppose you find a scenario today where current technology suffices, where the 0-to-1 incremental value is enormous and the 1-to-n space isn’t that big—that scenario is fine. Midjourney is like that, or copywriting generation—relatively simple tasks with very visible 0-to-1 effects. Those are opportunities for the applications-only camp. But the biggest opportunity isn’t there. If your premise is commercialization, you can’t think about it apart from AGI. If I only build applications now—fine, but in a year you could be crushed.
Zhang Xiaojun: You could quietly upgrade the underlying model, couldn’t you?
Yang Zhilin: But that approach can never become bigger than the model itself. Technology is the only new variable of this era; the other variables haven’t changed. Returning to first principles, AGI is the core of everything. From that, we deduced: a super app definitely requires the strongest technical capabilities.
Zhang Xiaojun: Can you use open-source models? (The latest news is that Google announced the open-source model Gemma.)
Yang Zhilin: Open source lags behind closed source—that’s also a fact.
Zhang Xiaojun: Might the lag be only temporary?
Yang Zhilin: It doesn’t look that way so far.
Zhang Xiaojun: Why can’t open source catch up with closed source?
Yang Zhilin: Because open-source development works differently now. In the past, everyone could contribute to open source; today, open source itself is still centralized. Many open-source contributions probably haven’t been validated by compute. Closed source enjoys concentrations of talent and capital, so in the end closed source will definitely be better—it’s a consolidation.
If I had a leading model today, open-sourcing it would very likely be irrational. It’s the laggards who might do that instead, or open-source a small model—to stir things up; after all, if you don’t open-source it, it has no value anyway.
Zhang Xiaojun: How do you push back against the anxiety in China? People say that a foundation-model company that doesn’t quickly produce commercial scenarios and products that meet investor expectations will struggle to raise its next round.
Yang Zhilin: You need a balance between the long term and the short term. Having no users and no revenue at all definitely won’t work.
As we’ve seen, going from GPT-3.5 to GPT-4 unlocked many applications; from GPT-4 to GPT-4.5 and then GPT-5, it will very likely keep unlocking more—even exponentially more. The so-called “Moore’s law of scenarios” means the number of usable scenarios rises exponentially over time. We need to improve model capabilities while finding more scenarios—that kind of balance.
It’s a spiral. It depends on how much of your investment goes to the short term and how much to the long term. You pursue the long term on the condition that you can survive. The long term absolutely cannot be abandoned, or you’ll miss the entire era. Drawing conclusions today is truly too early.
Zhang Xiaojun: Do you agree with the “two-wheel drive” idea put forward by Wang Huiwen, co-founder of Meituan and founder of Light Year?
Yang Zhilin: That’s a good question. To a degree, the logic holds. But how you actually execute makes a huge difference. Can you truly make some “non-consensus bets with favorable odds”?
Zhang Xiaojun: As I understand it, their two-wheel drive also requires quickly finding new application scenarios; otherwise, there’s no way for the technology to land.
Yang Zhilin: It still comes down to the difference between model scaling and user scaling.
Zhang Xiaojun: In China, besides you, who else takes the model-scaling approach?
Yang Zhilin: That’s not for me to judge.
Zhang Xiaojun: Most people probably take the user-scaling approach. Or can we put it this way: is this the difference between the academic camp and the commercialization camp?
Yang Zhilin: We are not the academic camp. The academic camp definitely doesn’t work.
Zhang Xiaojun: Many foundation-model companies commercialize through B2B—after all, B2B offers more certainty. Do you?
Yang Zhilin: We don’t. From day one, we decided to go B2C.
It depends on what you want. If you know something isn’t what you want, you won’t get FOMO—because even if you got it, it wouldn’t mean much.
Zhang Xiaojun: Have you been anxious over the past year?
Yang Zhilin: More excitement and exhilaration. Because I’ve thought about this for a very long time. We were probably among the earliest who wanted to explore the dark side of the moon. Today you find that you’re really building a rocket, and every day you’re discussing what fuel to add to make it go faster—and how to keep it from blowing up.
Zhang Xiaojun: To sum up the “non-consensus bets with favorable odds” you’ve made—besides B2C and long context, are there others?
Yang Zhilin: More are in the works; I hope to share them with everyone soon.
Zhang Xiaojun: China’s previous generation of entrepreneurs tasted success with applications and scenarios, so they focus more on product, users, and the data flywheel. Can the new generation of AI entrepreneurs you represent stand for a new future?
Yang Zhilin: We care deeply about users, too. Users are our ultimate goal, but it’s also a process of co-creation. The biggest difference is that this time it will be more technology-driven. It’s still the horse-carriage-versus-car question: we’re now in the leap from horse carriages to cars, and we should focus as much as possible on how to give users a car.
Zhang Xiaojun: Do you feel lonely?
Yang Zhilin: Ha ha ha… That’s an interesting question. I think I’m fine, because I still have dozens—nearly a hundred—people fighting alongside me.
Part 6
Before We’ve Even Caught Up with GPT-4, Sora Arrives
“Right Now It’s Like the GPT-3.5 Moment for Video Generation”
**
Zhang Xiaojun: Sora’s sudden appearance this year—how much of it was within your expectations and how much was not?
Yang Zhilin: That generative AI could achieve this effect was within expectations; what was unexpected was the timing—it came earlier than we had estimated. It also reflects how fast AI is developing now: a lot of the dividends of scaling haven’t been fully harvested yet.
Zhang Xiaojun: Last year, the industry judged that foundation models in 2024 would inevitably compete hard on multimodal narratives, and that video generation quality would improve as rapidly as text-to-image did in 2023. Did Sora’s technical capabilities exceed, meet, or fall short of your expectations?
Yang Zhilin: It solved many previously difficult problems. For example, maintaining consistency of generation over a relatively long time window—that’s the key point, and a huge improvement.
Zhang Xiaojun: What does it mean for the global industry landscape? What new narratives will foundation models see in 2024?
Yang Zhilin: First, near-term application value: it can further improve efficiency in production processes, and of course we hope for more extensions built on current capabilities. Second, combining with other modalities. It is itself a model of the world; with that knowledge, it’s an excellent complement to existing text. On that basis, there is quite a lot of room and opportunity, whether in agents or in connecting with the physical world.
Zhang Xiaojun: Overall, how do you assess Sora?
Yang Zhilin: We had been planning a similar direction ourselves and had worked on it for a while. Directionally, it wasn’t much of a surprise—it’s more about the technical details.
Zhang Xiaojun: What technical details are worth learning from?
Yang Zhilin: OpenAI hasn’t fully explained many of them either. They described the broad strokes; there are key details you have to infer from its outputs or from available information, combined with our own earlier experiments. At least for us, it adds more data points to the development process—more data input.
Zhang Xiaojun: Compared with text generation, what were the main bottlenecks in video generation? What solutions can you see OpenAI found this time?
Yang Zhilin: The main bottleneck—the core is still data: how do you fit the data at scale? That hadn’t been validated before, especially when the motion is complex and the generated result is photorealistic. Under those conditions, being able to scale—that’s what it solved this time.
What remains unsolved includes, for example, the need for a unified architecture. DiT is still not a very general architecture. Modeling the marginal probability of purely visual signals can be done very well, but how do you generalize that into a universal new computer? You still need a more unified architecture—there’s still room there.
Zhang Xiaojun: Have you read OpenAI’s Sora report, “Video generation models as world simulators”? What key points in it deserve highlighting?
Yang Zhilin: I have. Given the current competitive situation, they definitely wouldn’t write down the most important points. But it’s still worth learning from. This was essentially paid content—things you might otherwise have to spend money on many experiments to learn. Now you can know some of it without paying for the experiments, and form a rough understanding.
Zhang Xiaojun: What key signals did you extract from it?
Yang Zhilin: That this thing is scalable to a degree. In addition, it gives a fairly concrete account of how the architecture is built. But it’s also possible that different architectures don’t make such an essential difference on this problem.
Zhang Xiaojun: Do you agree with that line of theirs—“Scaling video generation models is a promising path towards building general-purpose simulators of the physical world”?
Yang Zhilin: I strongly agree. These two things optimize the same objective function—there’s not much doubt about that.
Zhang Xiaojun: What do you think of Yann LeCun once again speaking out against generative AI? His view: “Modeling the world by generating pixels is wasteful and doomed to fail. Generation happens to work for text because text is discrete, with a finite number of symbols. In that case, dealing with uncertainty in prediction is easy; dealing with predictive uncertainty in high-dimensional continuous sensory inputs is intractable.”
Yang Zhilin: I now think that when you model the marginal probability of video, the essence is lossless compression—no essential difference from next-token prediction in language models. As long as you compress well enough, you can explain whatever in this world is explainable.
But there’s also important work yet to be done: how does it combine with the capabilities that have already been compressed? You can think of it as two different kinds of compression. One compresses the raw world—that’s what video models do. The other compresses the behaviors humans produce, because human behavior has passed through the human brain—the only thing in the world that produces intelligence. You can think of video models as doing the first kind and text models the second, though video models also contain some of the second kind: some human-made videos contain the intelligence of their creators. Ultimately, it will probably be a mix—you need to learn from different angles through both approaches, and both contribute to the growth of intelligence.
So generation may not be the goal; it is merely the compression function. If you compress well enough, the generation will be very good in the end. Conversely, if a model itself cannot generate, is it still possible to compress extremely well? That’s doubtful. It’s possible that generating very well is a necessary condition for compressing very well.
Zhang Xiaojun: Sora and last year’s ChatGPT are two different milestones. Which is bigger?
Yang Zhilin: Both are very important. Right now it’s a bit like the GPT-3.5 moment for video generation—a step-function improvement. The model is still relatively small, and it’s foreseeable that there will be bigger models, which is a guaranteed improvement in capability.
Zhang Xiaojun: Some people also say that, for multimodality, Google Gemini’s breakthrough matters more.
Yang Zhilin: Gemini follows the GPT-4V line and incorporates that understanding as well. Both are important; the final step is putting all these things into the same model, and that hasn’t been solved yet.
Zhang Xiaojun: Why is putting them in the same model so hard?
Yang Zhilin: Nobody knows how yet. There is still no validated architecture.
Zhang Xiaojun: What will Sora + GPT produce?
Yang Zhilin: Sora can be applied to video production right away, but if combined with language models, it could connect the digital world and the physical world. You could also complete tasks more end-to-end, because your modeling of the world is now better than before—it can even be used to improve your understanding of multimodal inputs. So in the end you can switch quite fluidly between modalities.
To sum up: you understand the world better; you can do more end-to-end tasks in the digital world; and you can even build a bridge to the physical world to complete tasks there. This is the starting point. Autonomous driving, for example, or some household chores—these are, in theory, all instances of connecting with the physical world. So the breakthrough in the digital world is certain, but it also holds the potential of a path to the physical.
Zhang Xiaojun: What does Sora mean for Chinese foundation-model companies? What’s the right response?
Yang Zhilin: It doesn’t change much. This was always a direction of certainty.
Zhang Xiaojun: Chinese foundation models haven’t caught up with GPT-4 yet, and now Sora has arrived. What do you think? The two worlds seem to be drifting further and further apart—do you feel anxious?
Yang Zhilin: It’s simply an objective fact. But the actual gap may still be narrowing—that’s the law of technological development.
Zhang Xiaojun: What do you mean? That the technology curve is steep at first and then gradually flattens?
Yang Zhilin: Yes. I’m not really surprised—OpenAI has been working on next-generation models all along. Objectively, the gap will persist for a while, and gaps between different Chinese companies will persist for a while too—this is a period of technological explosion.
But in another two or three years, it’s possible that China’s top companies can do more of the foundational work well here—including technical infrastructure, talent reserves, and the sedimentation of organizational culture. With that honing, they’ll be more likely to lead in certain respects—but it will take some patience.
Zhang Xiaojun: Could China and the U.S. end up with completely different AI technology ecosystems?
Yang Zhilin: The ecosystems could differ, if you look at it from a product and commercialization angle. But from a technology angle, general capabilities won’t follow completely different technical routes—the basic general capabilities will definitely be similar. Because the space of AGI is vast, differentiation on top of general capabilities is more likely to happen.
Zhang Xiaojun: There’s a long-running debate in Silicon Valley: “one model rules all” versus “many specialized (smaller) models”—one general model for all kinds of tasks, or many specialized smaller models for specific tasks. What’s your view?
Yang Zhilin: I take the first view.
Zhang Xiaojun: On this point, will China and the U.S. diverge greatly?
Yang Zhilin: I don’t think so, ultimately.
Part 7
I Accept That Failure Is a Possibility
“It Has Already Changed My Life”
**
Zhang Xiaojun: Foundation-model entrepreneurship is a rather peculiar creature in China: you’ve raised so much money, yet it seems a large chunk of it goes toward scientific experiments. How do you persuade investors to open their wallets under these circumstances?
Yang Zhilin: No differently than in the U.S. The money we’ve raised today isn’t even that much. So we still have a lot to learn from OpenAI.
Zhang Xiaojun: I’d like to know: how much more money does it take to reach GPT-4? How much to reach Sora?
Yang Zhilin: Neither GPT-4 nor Sora requires that much. The money now is more about reserving for the next generation—or the generation after that—of models, for frontier exploration.
Zhang Xiaojun: Chinese foundation-model startups have taken the giants’ money, but the giants are also training their own models. How do you view the relationship between foundation-model startups and the giants?
Yang Zhilin: There’s both competition and cooperation. The giants and the startups have different first priorities. Look at any big tech company today: its first priority differs from an AGI company’s first priority. Your first priority shapes your actions and results, and ultimately defines the different relationships within the ecosystem.
Zhang Xiaojun: Why do the giants spread smaller investments across multiple foundation-model companies instead of betting heavily on one?
Yang Zhilin: It’s a matter of stage. There will be more consolidation going forward—and fewer companies.
Zhang Xiaojun: Some say the endgame for foundation-model companies is being acquired by a giant. Do you agree?
Yang Zhilin: Not necessarily, I think. But they may well have very deep partnerships.
Zhang Xiaojun: For example, how might they cooperate?
Yang Zhilin: OpenAI and Microsoft are the classic model of cooperation. Much of it can be referenced, and some of it can be improved.
Zhang Xiaojun: Over the past year, where did the twists and turns of entrepreneurship show up for you?
Yang Zhilin: There were many external variables—capital, talent, GPUs, product, R&D, technology. There were highlight moments, and there were difficulties to overcome. Take GPUs.
There was a lot of back and forth: supply was tight for a while, then it improved. The most extreme stretch saw prices change daily—a machine might cost 260 one day, 340 the next, then fall back two days later. It was a constantly moving situation. You had to watch it closely. Prices kept changing, so strategy kept changing: which channels to use, whether to buy or rent—there were many different options.
Zhang Xiaojun: What drove these fluctuations?
Yang Zhilin: Geopolitical reasons; production itself comes in batches; and market sentiment plays a role. We observed many companies starting to return GPUs, realizing they didn’t necessarily need to train this model. As market sentiment and people’s decisions shifted, supply and demand shifted with them. The good news is that overall supply has improved enormously lately. My personal judgment is that for at least the next one to two years, GPUs won’t be a major bottleneck.
Zhang Xiaojun: You seem to think constantly about organization. How have you gone about team building?
Yang Zhilin: Our approach to hiring evolved. AGI talent is extremely limited worldwide, and people with relevant experience are rare. Our earliest hiring profile focused on finding geniuses with directly relevant skills. That proved very successful. People who had previously “performed surgery” on models and had firsthand experience training ultra-large-scale models could build things very quickly. The Kimi launch included—the capital efficiency and organizational efficiency were actually very high.
Zhang Xiaojun: How much did it cost?
Yang Zhilin: A rather small number—compared with many other expenses, it was doing big things with small money. For a long time we were at 30–40 people. Now we’re at 80. We pursue talent density.
The talent profile changed later. In the earliest days we hired geniuses, believing their ceiling was high—a company’s ceiling is determined by the ceiling of its people. Later we rounded out the team with more dimensions of talent—people on the product-operations side, leader types, people who can take things to the extreme. Now it’s a more complete, resilient team that can fight.
Zhang Xiaojun: After a year of foundation-model entrepreneurship in China, how do you assess the milestone results so far?
Yang Zhilin: We’ve built a rocket prototype and are now test-firing it. We’ve assembled a team, figured out some of the fuel formulas, and can more or less see an embryonic PMF. You could say we’ve taken the first step of the moonshot.
Zhang Xiaojun: What do you think of Yann LeCun’s position? He’s not optimistic about the current technical route—he believes self-supervised language models cannot acquire true knowledge of the world, and that as models scale, the probability of errors—machine hallucinations—will only grow. He has proposed the idea of a “world model.”
Yang Zhilin: There’s no fundamental bottleneck. When the token space is large enough, it becomes a new kind of computer that solves general problems—and that is a general world model.
An important point behind his statement: everyone can see the current limitations. But the solution doesn’t necessarily require an entirely new framework. The only thing that works in AI is next-token prediction plus the scaling law. As long as the tokens are complete enough, everything is doable. Of course, the problems he points out do exist today—but you solve them by making the token space very general. That’s all.
Zhang Xiaojun: So he’s magnifying the limitations.
Yang Zhilin: I think so. The underlying first principle is sound—it’s just that some small technical problems remain unsolved.
Zhang Xiaojun: What do you think of Geoffrey Hinton, the godfather of deep learning, repeatedly calling attention to AI safety?
Yang Zhilin: His focus on safety actually shows he has enormous confidence in the coming improvement of technical capabilities. The two are opposites.
Zhang Xiaojun: How do you solve the hallucination problem?
Yang Zhilin: Still the scaling law—it’s just that what you scale is something different.
Zhang Xiaojun: How likely is it that, in the end, the scaling law turns out to be a dead end?
Yang Zhilin: The probability is approximately zero.
Zhang Xiaojun: What do you think of the view of your CMU alumnus Qi Lu: OpenAI will definitely be bigger than Google—it’s just a question of whether by one, five, or ten times?
Yang Zhilin: The most successful AGI company of the future will definitely be bigger than every company today—of that there’s no doubt. Ultimately, it could be a matter of double or triple GDP. It may not be OpenAI; it could be another company. But there will certainly be such a company.
Zhang Xiaojun: If you happened to become the CEO of this AI empire, what would you do to protect humanity?
Yang Zhilin: Thinking about that question now still lacks some preconditions. But we would certainly be willing to cooperate with and learn from different actors in society, including putting more safety measures into the models.
Zhang Xiaojun: What are your goals for 2024?
Yang Zhilin: First, technical breakthroughs—we should now be able to build a model far better than in 2023. Second, users and product—I hope for more users at scale and stronger retention.
Zhang Xiaojun: What are your predictions for the global foundation-model industry in 2024?
Yang Zhilin: More capabilities will appear this year, but the landscape won’t look much different from today—the top few players will still lead. On capability, there should be some fairly big breakthroughs in the second half of the year, many coming from OpenAI; it definitely has a next-generation model—maybe 4.5, maybe 5. That feels like a high-probability event. Video generation models can definitely keep scaling.
Zhang Xiaojun: And your predictions for China’s foundation-model industry in 2024?
Yang Zhilin: First, you’ll see new, distinctive capabilities emerge. Chinese models—because of earlier investment and having the right teams—will achieve world-leading capabilities in certain dimensions. Second, products with much larger user bases will appear—that’s highly probable. Third, there will be further consolidation and divergence in route choices.
Zhang Xiaojun: In starting this company, what’s the one thing you fear most?
Yang Zhilin: Nothing much, really—you just have to charge forward fearlessly.
Zhang Xiaojun: Anything you’d like to say to your peers?
Yang Zhilin: Let’s keep at it together.
Zhang Xiaojun: Name one question about the foundation-model industry that you don’t yet know the answer to but most want to know.
Yang Zhilin: I don’t know what the ceiling of AGI looks like—what kind of company it will produce, and what kind of products that company will create. That’s what I most want to know right now.
Zhang Xiaojun: As AGI keeps developing like this, what’s the one thing you’d least want to see?
Yang Zhilin: I’m fairly optimistic about it. It can take human civilization to the next stage.
Zhang Xiaojun: Has anyone ever said you’re too much of an idealist?
Yang Zhilin: We’re very down-to-earth, too. We’ve actually built some real things—we’re not just talking.
Zhang Xiaojun: If the money you’ve raised today were the last money you’d ever get, how would you spend it?
Yang Zhilin: I hope that never happens, because we’ll need a lot more money in the future.
Zhang Xiaojun: If you don’t make it, would you consider yourself a failure?
Yang Zhilin: It wouldn’t matter that much—I accept that failure is a possibility.
This endeavor has already completely changed my life, and I am full of gratitude.
(End)





