👀Liang Wenfeng on AGI, Compute, and Why DeepSeek Stays Open Source
Full translation of DeekSeek Founder's four-hour investor meeting: the AGI roadmap, the US-China gap, Huawei chips, and open source.
Hangzhou-based AI startup DeepSeek recently raised over RMB 50 billion (~$7.4 billion) in its first-ever financing round at a pre-money valuation of over $50 billion. And the comapany is reportedly in talks to raise another $1.5 billion at a $71 billion to $74 billion valuation ahead of a planned 2027 IPO.
At a recent investor meeting, DeepSeek CEO Liang Wenfeng laid out in detail DeepSeek’s organizational culture, open-source strategy, technology roadmap, and his personal views on today’s AI competitive landscape.
Multiple Chinese media outlets shared their versions of this four-hour investor meeting transcript, and I found Tencent Technology's version to be the most complete, including 118 remarks. Below is the full translation done with AI and fact-checked by me. You can find the original Chinese link here. Grace Shao and Geopolitechs each translated different versions, so I’ve attached them here for reference.
01 Vision and Restraint
When we first started this company, the original intent was never about how much money I would ultimately make, about going to the capital markets, about listing, anything like that. The first few dozen people never thought that way at all. If someone had thought that way, they wouldn’t have come.
We’re doing this with enormous goodwill toward the world. We think it’s useful for humanity, and that it’s a matter beyond money. Our original intent, our vision, and the vision we’ve held on to until now, is not framed around maximizing commercial gain.
Managing a large company doesn’t run on your rules and regulations, it runs on vision. Vision isn’t a slogan hanging on the wall. Vision is how you act, not how you talk, it’s how you actually operate.
We don’t have an organization, we’re vision-driven, organized around a vision. We don’t operate on the basis of “I have to hit some KPI”—there’s no performance review. There’s only the vision.
This vision isn’t even written down, it has never been put into words, nothing has ever been written. The vision lives in how we do things, in our attitude toward the world.
We don’t have many other advantages. We have no special powers, we’re not richer than anyone else, and it isn’t the case that our people are better than other companies’. Really, no. When we founded this company two years ago, we didn’t have much money, we didn’t have many chips, we had no name recognition, no ability to rally people. We were just a group of very ordinary people.
The more restrained you are, the easier it may be to succeed—or at least that’s what’s been borne out so far, that’s what makes sense so far. Otherwise there’s no way to explain why we were able to pull this off: we had no weapons, we started from a very low base, we had very few resources, and our people are really just a random group of ordinary people.
This AI thing is too big, the interests at stake are too big. We are very restrained, because as long as we can pull it off, the payoff in the end will be enormous. Even a small slice of it is enormous. So right now there’s no need at all to think about which slice of the upside to take, or how to take it. I think there’s no need to think about that at all, because the upside is already large enough.
Last Spring Festival we suddenly had a lot of users, but we didn’t set out to retain those users, or to monetize them, or to grab the commercial upside and cash in on the user base. We didn’t fight for users, we didn’t try to make money, but we worked very hard to figure out how to serve those users well.
We would never have the thought that I want to build the next super app, that I want to compete with someone, that I want to become the next ByteDance or the next Tencent. No such thought at all. I think the AGI opportunity ahead is enormous, the AGI opportunity ahead is always enormous.
Restraint is a strategy. It lies in the fact that sometimes you can give things up in exchange for more of something else. Not open-sourcing works the same way—you could see it as pressure on us, or you could see it as us giving up margin.
My understanding of this restraint is that, over the long run, it increases our probability of achieving AGI. When I consider something, I have no doubt that AGI will have enormous commercial value. On that basis, my first consideration isn’t how to add a bit more share, how to grab a bit more share. My first consideration is how to increase the probability that I can pull it off.
We’ve always been very restrained, unwilling to become the adversary of any internet giant or any small player. I hope I can empower them, or that I can help everyone do this, that I can help everyone accomplish this.
I think we held that attitude before, and we didn’t actually get any less because of it. Open-sourcing, our goodwill, the help we gave others—none of it caused us to get any less. If anything it may have been a plus. That looks counterintuitive, but that’s really how it is.
AGI is our goal, but we’ve been commercializing the whole time, which is why we have consumer users and why we have B2B revenue. Judging from experience, that strategy has worked.
02 The AGI Roadmap
If you can describe a problem very clearly and give it complete context and instructions, it already surpasses humans. But there’s a definition here, a precondition: you give it complete context, you give it complete instructions.
AI can’t replace your employees. But if AI had the ability to learn continuously, then like your employees, it could come to the company and learn for two months—and then it could replace everyone in the world. So we’re one step away from continual learning.
We can think of AI’s development as a staircase. The step taken last year was chain of thought. Because we discovered that with chain of thought, you can push intelligence to a higher level.
This year’s step is agents, because we’ve found that with an agent approach, it can handle more, its range of capability is broader, and its intelligence ceiling is higher. Agents have to use CoT, and CoT has to use the step before it, which is language models. So not one step was wasted.
After agents, the problem we think needs to be solved is continual learning—how to let a model keep learning, rather than requiring you to give it one heavy round of training. It should be able to do relatively long-horizon continual learning, the way a person does.
After continual learning, we may reach a singularity. That singularity is that once the model can learn continuously, it can already do everything humans can do. It can develop its own version, it can do its own research, and then develop its own next version, developing better AI models.
This singularity isn’t really a singularity, it’s a gradual process too. That process may be a fairly long gradual change, not a sudden break. But out of habit we all tend to think of it as a singularity.
This is our conjecture, this is what we think the timeline should be: first solve the learning problem, then reach the intelligence singularity, the self-iterating singularity, and only after that embodied intelligence. Once you get to embodied intelligence, it enters the physical world, it can do your housework, it can care for you in old age.
If we solve continual learning first, then the self-iteration singularity, then embodied intelligence, the road gets easy. Because later on you can use the earlier technology to help develop the later technology.
We only work on the main line toward AGI. The AI field is very broad, and there’s a lot we think isn’t on that main line—3D, video generation, for example. I don’t think those have much to do with the main line of intelligence, so we won’t do them.
When video generation first came out it was very hot, as if it were something you had to do, as if you weren’t an AI company if you didn’t. That struck me as strange, because if you actually think it through, it has nothing to do with the intelligence roadmap.
Commercially it’s a good business, commercially it is a good business. But it has nothing to do with intelligence. We won’t do something because it’s a good business. We’ll only do it because it’s on the intelligence roadmap.
By our judgment, world models and intelligence aren’t the most important thing at this stage. What matters most is AI training, and how to solve continual learning after training. That’s our company’s judgment—of course, every company’s judgment differs.
Right now we fairly strongly believe the narrative that AI can accelerate AI research. That is, it isn’t linear, because you can use AI to accelerate your own research, so further out it may become nonlinear.
I think we definitely have to get into embodiment eventually. Because for an ordinary person, what they need isn’t a computer, right? An ordinary person—food, drink, entertainment, clothing, housing, transport—doesn’t need a computer. What they need is embodied intelligence to solve concrete human needs.
What do we hope AGI can do? Help me iterate the next version of the model, exactly that. And once we have embodiment, what we want it to do is the same: let it iterate the next version of embodiment, let it build the next version of the robot.
The core capability of the next generation of models has to be continual learning—only then does it deserve to be called next-generation. Before that, what we can do is lower costs, improve quality, and increase speed. But for a big breakthrough, it has to have continual learning.
Today’s agents are limited in capability because they can’t learn continuously, they can’t learn continuously in an effective way. If continual learning gets solved first, AI’s capabilities will be very strong, and it will hugely improve the efficiency of our own research.
Get continual learning first and general intelligence may come easily, easily built on top of it. That’s the outcome I’d rather see, because we’d expend less effort, we’d have it easy. Otherwise, building general intelligence by hand right now is tiring and painful, it’s data-intensive and labor-intensive, and the return on effort isn’t good.
03 Team and Talent
What our earlier experience taught me is that the AGI vision is very powerful. The talent advantage isn’t that my people are smarter than his. It’s how I organize these people, how I motivate them, and then how they collaborate.
Putting smart people together doesn’t mean they’ll naturally collaborate, or that they’ll naturally chase a goal with passion and see it through. So you need a vision.
Our biggest core interest is maintaining the stability of the team. That’s our biggest core interest, arguably the only one. As long as I can keep the team stable, I will definitely pull it off, definitely achieve AGI. It’s that simple.
Money is certainly not a problem, resources are not a problem, all the other elements are easy to obtain. For us there’s only one core interest, only one thing we can’t compromise on: we must keep the team stable.
That’s also a very big challenge for us, or rather, I think it’s the biggest risk. Of course, that risk has been substantially defused by our recent financing, because the options everyone received are fairly substantial, the amounts are fairly large.
On team stability, as long as the most important employees, the longest-tenured employees, are stable, then the others aren’t likely to leave. Even with fewer options and less income, they won’t leave, because they aren’t in it purely for the money. Everyone wants to work in an environment where AGI can actually be achieved.
Everything else is a matter of time. At worst it costs us half a year or a year, but it won’t mean we can’t do it. We definitely won’t be short of money, definitely won’t be short of resources. None of that is lacking.
Our gap with the US is mainly in resources. In terms of people, the gap isn’t large. On people there’s almost no gap, because it’s the same pool of people, and they may well be Chinese. When Chinese people go abroad, some stay in China, some stay overseas, some go overseas—it isn’t that the smart ones go abroad. That’s not the case.
Talent isn’t the bottleneck, resources are the biggest bottleneck. Resources first affect talent development, because with less compute we get fewer chances to run experiments, so our talent overall lags the US. The talent gap is fundamentally a compute gap.
The AI talent shortage is also a phase, and we’ve already seen it ease substantially. Because there really is no shortage of AI people—every company will train people up fast. Training people is quick.
There are somewhat too many companies building models in China, still too many. In the US there may be just three. China has far too many outfits building foundation models. In the end you certainly don’t need that many people doing foundation models, it will definitely consolidate.
Our company’s management really runs on two lines: one top-down, one bottom-up. Bottom-up means each person does what they themselves want to do, on their own, with nobody managing them and no KPIs.
Generally we want employees to have half their time unassigned—they do whatever they want. That’s the scope for research, letting them explore on their own, pursuing whatever they think is important, with no requirements set in advance.
We generally don’t work overtime much. There are two reasons for that. First, doing research requires a relatively relaxed environment. If you push too hard there’s no way to do research, because it depends on your own interest, on you thinking about these problems in your ordinary time. So it has to be a relatively relaxed environment for exploration to be possible.
Second, we’re very focused. Being very focused means we have very few things to do. So I don’t have that much to do, and I don’t need to work overtime. This follows the same thread as the restraint I mentioned earlier.
Our company as a whole is built on consensus. It isn’t that I decide everything alone. I have to seek consensus. My authority within the company, and my influence within the company, are built on consensus.
This decision mechanism is really a consensus-seeking mechanism. It isn’t that I can push something through—it has to be consensus before I can push it through, and only then will I push it.
As headcount grows, we’ll make adjustments. In fact we have to make that adjustment right away, because I’m already making it. Without that adjustment, a lot of things can’t move forward. There really are many departments that ought to have an org structure.
04 Compute and Resources
How many chips do we need? Right now, obviously the more the better. Within what we can bear, more chips is unquestionably better. So our current strategy is to buy as many chips as we can at a reasonable price.
In practice, spending that much money is very hard—you can’t buy that many chips, they’re hard to buy, and prices are high. And you can’t just pay any price, you still have to make sure the price is reasonable. If we can spend RMB 20 billion this year, our procurement department will have had a spectacular year.
The biggest gap between us and the US is in resources. On compute, part of it is that chips simply can’t be bought in China, and part is that our capital investment is smaller than in the US. We invest much less capital, and salaries are a very small share of that. You see the hundred-million-dollar pay packages they offer, but when you do the math, talent salaries are still a small share. The bulk is compute.
Every difference we see—differences in talent, differences in model capability, differences in applications—can be attributed to differences in compute resources.
Our gap with the US may be about 12 months behind, maybe 12 to 18 months behind, or 6 to 12 months. Put simply, we’re two years behind the US, and we did it with one-twentieth of their compute.
The narrative is: one to two years behind, but using one-twentieth of the compute. Going forward we want to rewrite that narrative to: we use some fraction of their compute, but we compress the time gap further, down to 6 months, down to 3 months. I think that’s a goal.
Scaling—we believe in scaling. Bigger scale definitely means better results, and it unlocks more capabilities. What stops us from scaling is compute. It isn’t that we don’t want to scale, it’s that we don’t have enough compute to do it.
We train models of this size not because I think this size is enough, but because this is the amount of resources I happen to have. I work backward from my resources to how large a model I can accept and can train. That’s how the number comes out. It isn’t that this model size is enough.
When Silicon Valley says scaling has hit its limits, that’s for Silicon Valley. For Chinese players, we’re still far from that, we haven’t scaled to anywhere near that level. And scaling here includes scaling of data, scaling of model size, and training cost.
05 Domestic Chips and Ecosystem
Nvidia’s CUDA moat is being dismantled fast. Part of that is that we now have AI, and with AI, building up an ecosystem is far easier than before, because AI can write code.
The compute-chip market is already bigger than the gaming-chip market, so there’s no reason for the two to stay coupled. The trend now is that they’ll be decoupled going forward. Which means dedicated chips—whether from Huawei or from Nvidia itself—will all be dedicated chips from now on, not the things we had before.
There’s a historic opportunity right now for domestic AI chip substitution. We believe that within the next year we’ll see one thing proven out: that there is absolutely no problem with the domestic chip ecosystem. People previously thought there was a problem, that it couldn’t be used, that it was hard to use. But within a year I think we can flip that perception, or facts will flip it.
Domestic AI chips have no problems in hardware or ecosystem. The only problem is insufficient production capacity. Adapting to domestic cards presents no obstacle, and Nvidia can’t stop it. In a normal commercial environment, where I could buy Nvidia cards, domestic substitution would be quite hard. But when Nvidia cards can’t be bought, everyone has no choice, everyone has to go domestic.
When V3 was trained it still used Nvidia cards, but it no longer used Nvidia’s ecosystem. V3 used Nvidia cards but not the Nvidia ecosystem. Instead we first wrote a high-level compiler called TileLang, and built everything else on top of the TileLang ecosystem. So we already barely depend on Nvidia’s ecosystem.
I’m fairly optimistic about domestic compute. On this point I think Nvidia is digging its own grave. Huawei’s supernodes, Huawei’s 950 supernode, can fully substitute for Nvidia’s GB200 and GB300 on both performance and price.
Four Huawei cards equal one Nvidia card.
On our chip gap with the US, I believe there will no longer be a gap on the ecosystem side, but on chips it’s four times plus two years.
Right now we mainly work with Huawei. Huawei does its own adaptation, but we participate in the ecosystem ourselves, we get deeply involved on the Huawei side. Huawei’s problem is still insufficient capacity.
I don’t really believe that five years from now we’ll still be stuck on capacity. Right now we’re certainly stuck on capacity—this year, next year, the year after, I think we may still be stuck on it. But five years out, I think not necessarily. I’m fairly optimistic.
06 Competitive Landscape and Industry Judgments
The gap that ultimately separates each company’s models should be a composite one. Comparing model quality only means something if you compare at the same cost, because when you compare two cars you compare cars in the same price bracket.
Anthropic is now ahead of OpenAI—is that durable? I don’t think it’s durable, it’s certainly a phase. OpenAI and Google will most likely keep trading places going forward.
In the global division of labor in AI, the role Chinese companies are quite likely to play is still the largest producer. By ordinary logic, our production capacity is the largest, including chips—our capacity in chips may be the largest, and we have the most electricity.
Chinese players will make the product as cheap as possible, and then compete on quality. After all, for many goods today there’s no big difference between Chinese-made and American-made. AI in the future may be the same, but Chinese-made AI may be cheaper. That cheapness may be systematically lower, the same way Chinese-provided services in other industries are cheaper.
The final gap should come down to three things: cost, time, and user experience. Beyond that, there probably isn’t much of a gap.
Cost is certainly one difference—I think cost is the number one difference. The second is time: when you can get there. A few months earlier or later makes a difference.
OpenAI from the beginning thought it really could monopolize the world, but in reality it will meet many, many challengers. It will face challenges, so it won’t have it easy. The US will face challenges, and in the future it may also face challenges from China, because Chinese players are willing to take less in return for providing the service.
Those who take more will be beaten by those who take less. In fact you don’t even have to actually take more—if your vision is to take more, you’ll be beaten by whoever’s vision is to take less. Nobody has actually taken any money yet, it’s only a vision. If your vision is to take more, you’ve already lost, you’ll face greater difficulty.
For us, it isn’t about capturing the most profit, or pricing to maximize returns. It’s about earning a reasonable return. That’s the explanation. I believe in this, I’m not looking for a justification for it, because there’s no need for one.
I think in many aspects of experience, we may be able to do better than the US. On product capability we may not be worse than the US. Costs should also be lower than the US, so China will still be competitive.
Cost is easy to understand—they don’t have to do it, so they don’t develop the capability. They certainly don’t take it as seriously as we do. We can treat it as extremely important, but for them it’s unimportant.
For LLMs, maybe two large companies and two small ones is already enough. There are only two differences: time and cost. So no one is going to earn outsized profits, I don’t think there’ll be outsized profits. Whoever controls costs well earns a bit more, whoever controls costs poorly earns a bit less. That’s all.
07 Model R&D and Technology
Maybe half the people at our company think OpenAI is better on any given day. Anthropic does have a first-mover advantage, but that advantage should disappear soon. It isn’t an advantage it can hold long-term. All three of them are formidable, and among the three, its efficiency is the highest—the cost it spends, the money it burns, should be the least.
On multimodal, we’ve always been working on it. For products it’s very important, for consumer-facing products it’s very important. But for the ceiling of intelligence it’s a component, it isn’t the main line itself.
We will likely ship the relevant models—V4 and subsequent versions of V4 will support native multimodality. But for us multimodality is a component of intelligence, we don’t treat it as intelligence itself.
All I can say is that with the scaling of language models, I don’t yet see a ceiling. Neither our level of intelligence nor the level achieved in the US shows a ceiling yet.
Internally, a lot of us think this way: first it has to be useful to us, first it’s for our own use. That’s the fastest route to AGI. When it’s good for us, that probably means it’s good for others too. But first we have to make sure it’s good for us.
For the models we build, the first goal isn’t that everyone finds them good to use, it’s that we find them good to use. First it has to be useful to us. Once it’s useful to us, I’ll be faster when developing the next version of the model.
We call this “drawing lots.” The bar is low, anyone can try, but who draws something out of it—I don’t know whether that comes down to talent or something else. So there’s no need for us to allocate resources here. The difference between us and other companies is just that we spend time discussing the question, we think about it, and we treat it as important.
08 Commercialization and Pricing
Our API pricing reflects a reasonable profit—roughly, we go to the market and buy a batch of equipment and recover the cost in ten months. I think that’s a reasonable profit.
If we were maximizing profit, we should set prices higher. Because in this price range, demand is inelastic. If I raised prices by half again, or doubled them, token consumption wouldn’t differ much.
With one of our models, we were worried at first that demand would be too high, so we set the price relatively high, and the team wasn’t very happy. Later I brought the price back down, cut it to a quarter, and everyone was happy.
The ceiling on the B2B business should still be demand. Against the backdrop of this generation of AGI/AI technology, B2B demand should be limited. It will grow fast, but it isn’t infinite. In the end it’s constrained by demand, not by compute.
As things stand now, I think we can do it—we can have both. Say this year I have a few hundred million dollars in B2B revenue, plus our consumer user base—that in itself is already a commercial foundation. If next year we have B2B revenue and that demand grows further, the company isn’t far from net profit, it may already be net profitable.
In the worst case, just selling API could probably support a listed company. If there’s no further technical progress and our technology freezes here, then in the end we go all in on selling API and doing those services well. I think that would be enough.
Looking at things as they stand, I think the most sensible approach is to go all in on general-purpose agents, and to put other agents at lower priority, including finance agents and doctor agents. Coding has to come first, because coding agents can do a lot, and there are many vertical agents. At this stage, we think coding agents are the most important.
I think low cost is first of all a result. Our models really have been moving in a lower-cost direction architecturally, and that relates to our vision. We still have many algorithmic approaches left, and costs can go lower still.
There’s another reason costs keep going down: the lower the cost, the larger the model I can train, the larger the model I can afford. On the same compute, with limited compute, higher computational efficiency means I can afford a larger model.
09 Open Source Strategy
I think we will open-source, and our strongest model will probably be open-sourced too. Because I don’t see any benefit to being closed-source, I don’t see a necessary benefit. ByteDance’s models are closed-source—what benefit does it get? I don’t see any benefit.
Even if the model is open-sourced and you tell everyone everything, the bar is still very high. For others to actually use it, the bar is still very high. It’s hard for them to put it to use. And beyond that, for them to get costs very low is hard, very hard. It isn’t that easy.
Open-sourcing doesn’t affect revenue. Open source, I think, has no impact whatsoever on our business model.
And I’m not worried about others deploying our models and competing with us—not worried at all. We actually hope they’ll deploy them. We give the open-source community as much help as we can, helping everyone get our models deployed.
When we deal with the outside world, our attitude is: we only work on the main line toward AGI. In dealing with the outside world, we’re very willing to assist and help anyone, even our competitors, including Alibaba, Zhipu (Z.ai), and Moonshot AI, to do better. Because we don’t lose anything—we were open-source anyway.
Is the open-source model we release the same as the model we deploy ourselves? It’s the same. We won’t open-source a weaker model and then use a better one for our own deployment. We won’t do that, it’s the same model.
10 Data and Post-Training
Data is probably equal to half the model. And before that comes the labeling problem. On data labeling, this ties back to our capital investment. With our capital investment structure, we can’t support the cost of that much high-quality data labeling, because the cost is very high.
The cost of data labeling in the US and in China isn’t much different. Labeling data in China has no cost advantage, especially on high-end data, which makes it hard for us to invest in labeling the way the US does. This path is hard in China, because labeling data is simply too expensive. Whether we outsource it or do it ourselves, it’s painful.
Right now we’re basically walking on two legs. It isn’t that we can’t label at all—it’s that some labeling is cheap and some is expensive. We do the cheap parts first.
You could also say that right now half the company is labeling data. Half of our core researchers, our most important people, are labeling data. We’re concentrated on labeling data. Solving the AI problem at this stage comes down to labeling data.
The bottleneck on high-quality data labeling, I think, is time—it needs time. Because for OpenAI, for players abroad, for Anthropic, they all started earlier, and they have more capital and more chips.
Hallucination in LLMs has a significant effect on user experience. There’s a way to address hallucination, but it’s a long-running proposition. Hallucination can be seen as something solvable through better post-training, a problem that can be solved and improved.
11 Organization and Company Positioning
First, we have no model to imitate. Every step comes from our actual situation, from seeking truth from facts, making decisions based on real conditions and finding what we should do. So it’s a product of its time, or a reflection of real circumstances. It isn’t the result of imitation.
We’re clear that we have to commercialize. In the end we still have to survive. We are, after all, a company—the government isn’t going to give me a cent.
Fundamentally we’re still a company. It’s just that we make trade-offs about which money to earn, when to earn it, how much to earn, and what to earn it from. Many companies become great because they have a pursuit beyond profit. In the end that pursuit doesn’t hurt their commercialization—it actually lets them commercialize better.
As for partners, our financing was carefully selected. First, I think interests should be relatively aligned: those whose interests align most with ours, who bear us the least hostility, or who most want us to succeed. Not everyone wants us to succeed, because we do harm a lot of other people’s interests.
AI right now doesn’t lack taste or intuition. What it lacks is the ability to learn continuously. AI’s taste and intuition are fine. Ask it to write an article—its taste and intuition, I think, are fine.
We want to do just one piece. I think AI is a big thing, and it doesn’t need me to... I’ll do just one piece. If we stay focused, and I believe the business upside here is already large enough—if the AI era produces many trillion-scale companies, I think we’ll be one of them.
We’d like to support more people, but we don’t have that much bandwidth. We have the intent, and there’d be no conflict of interest, but whether we’ve actually done it is another matter. At least there’s no conflict of interest here, and we hope for win-win cooperation.




