⚔️Out of the Trough: How Huawei's Ascend Chips Climbed Back
A translation of a five-hour interview transcript between Zhang Xiaojun and Liao Heng, a Huawei Fellow and chief semiconductor scientist, on Ascend, US sanctions, and the τ Law.
Huawei’s semiconductor unit HiSilicon has been quiet since U.S. sanctions hit in May 2019. It can no longer work with global fabs like TSMC or Samsung that rely on U.S. technology, nor, like the rest of China’s chip sector, access EUV lithography machines from ASML, or EDA and other design software from global vendors. Huawei has taken on the role of making China’s chip supply independent and free from U.S. restrictions.
But since last year, Huawei has been more visible in the media as its Ascend AI chip series began delivering real results. In 2025, the company reportedly shipped 600,000 to 700,000 chips, and plans to double that in 2026. The next-generation 950 and 960 superpod will come online in 2026 and 2027 respectively.
Even so, it was a surprise that Liao Heng, a Huawei Fellow and chief semiconductor scientist, sat down for a five-hour interview with Zhang Xiaojun, a rising Chinese podcaster and journalist known for long-form, high-quality conversations. Liao is an incredible speaker and storyteller. His explanations of semiconductors and his analogies won’t put you to sleep.
There’s not much exclusive information here. It’s about looking back on the journey and sharing the mindset, not laying out a roadmap or chip specs. Even so, the conversation is a rare window into how Huawei has pulled its AI chip business up out of the trough. A few of my personal takeaways:
When HiSilicon was sanctioned in 2019, most of Huawei’s engineers and scientists didn’t leave. They stayed to take on the challenge.
The chip industry is fragmented, with each player handling its own piece. The sanctions forced Huawei to study every layer of the problem and move toward a vertically integrated model.
Open-source models actually help Huawei design chips like the 950, because the team gets a better understanding of what’s happening inside. Working closely with DeepSeek and other LLM clients, they can optimize chip designs faster than ever.
The Ascend 960, 970, and 980 are all already in the development pipeline.
In 2025, China newly deployed a bit over 1 gigawatt of data centers, while the U.S. may have added 7 to 8 gigawatts, roughly a one-fifth to one-seventh ratio.
While Chinese companies still lag behind U.S. frontier labs on compute, Chinese universities and national research labs actually have far more compute than their U.S. counterparts.
Liao said he isn’t betting against the U.S. but betting long on China. That mindset represents a lot of Chinese researchers and talent: they don’t necessarily want to build AI to defeat the U.S., but they hope AI can help China grow and prosper.
Below is a full AI-enabled translation of the interview. You can find the Chinese transcript here and watch the video here. Given the transcript is too long, you can jump directly to the History of Ascend part if you are only interested in Huawei’s AI chips.
Zhang Xiaojun: Hello everyone, I’m Xiaojun. Today our guest is Dr. Liao Heng, Huawei Fellow and Chief Scientist for semiconductors. This should be the first time, since Huawei went through the ordeal of 2020, that a Huawei executive has come out to talk about how Huawei’s Ascend chips climbed step by step out of that low point. At the same time, this is also my personal first time learning about the chip and semiconductor industry. So on one hand we’ll talk about Ascend’s story and China’s choices within it, and on the other hand we’ll take a broader view of the history and overall landscape of the global semiconductor and chip industry.
What follows is my interview with Dr. Liao Heng. What changed in his mindset between designing the 910 and the 950—the two generations of chips before and after the export ban?
Liao Heng: The change in mindset—I think when I try to put it into words, it’s a bit like this: try to imagine, if you were Dong Cunrui about to throw yourself on the explosives, right? Normal people really can’t imagine what his state of mind was, or what the soldiers at Shangganling felt. Or there’s an interview I once saw of a Chinese soldier during World War II—it was also a story on one of Huawei’s propaganda posters. A reporter asked him what he wanted after the war ended. He said he didn’t want anything, because his parents were already dead. There was no need to think about those things anymore.
So I’d say this so-called change in mindset doesn’t really carry that much meaning anymore. It’s just that when you face a difficulty, you want to solve it.
Zhang Xiaojun: Today Dr. Liao asked one thing of me—he didn’t want the focus to be on him personally. But he has witnessed more than thirty, even forty years of global semiconductor industry history. So I really want to start from that angle—to see, through your eyes, this journey through semiconductor industry history. Setting personal experience aside, if we just talk about the semiconductor industry itself—over the past three or four decades, how would you divide it into eras? What’s the core question at the heart of each one?
Liao Heng: That’s a fairly long topic. But before we get into it, let me describe how this interview came about. Huawei’s HiSilicon has rarely appeared in the media in the past. This time I was fortunate enough to receive this invitation. I did think it over for a bit before deciding to accept—I think there were roughly three reasons that led me to make up my mind.
The first is that this is a chance to look back on the engineers’ story, and maybe leave some insight for the rest of the world.
The second reason: about six months ago I gave a lecture at Tsinghua, in Professor Lu Youyou’s Computer Organization course in the computer science department. After giving that lecture, I felt a deep sense of frustration—and Professor Lu, who was teaching the course, felt the same helplessness afterward when we talked. It’s because in this era, young students are much more inclined to chase hot topics—AI algorithms, model training, even inference acceleration. Computer Organization is a hard course to begin with, and the term project requires designing a complete CPU—everyone considers it a “gate of hell” course—so it’s very hard to get students interested in hardware, in computer processor hardware. I felt both extremely surprised and also a kind of anxiety: if people are no longer willing to study this, will the field lose its future talent pipeline? Could we end up with a generational gap in the workforce?
And the third reason: we felt that people’s understanding of this industry—even our own subordinates or colleagues—is lacking. They’re buried in hard work every day and rarely get the chance to see a fuller picture of the industry, to understand how it all came to be. Because of that, it’s easy to feel lost or to waver.
So I thought that if there were an opportunity—not just for peers, or young students, or my own colleagues—to present a relatively complete view spanning several decades, a developmental perspective, the patterns of the industry—I felt that would be a pretty good opportunity. So these three motivations together are what helped me make up my mind to accept this interview.
As for the topic you raised—the general arc of the semiconductor industry over the past thirty years—I started my undergraduate degree in ‘87, then went to the US around ‘96, and joined a semiconductor company in ‘97. Naturally, during my studies—whether as an undergraduate or graduate student—the topics I worked on were related to processors.
If I try to summarize this industry over the past thirty years, there have been some very dramatic ups and downs. I think if you look at it vertically, there are actually two main threads woven together that form the overall arc of chips, or hardware.
One thread is obviously the processor—the CPU. That started around ‘86, or maybe a bit earlier than the IBM PC, around ‘84–’85. From 1984–85 the IBM PC appeared, and things gradually moved from a desktop office terminal into enterprise IT servers, and then later into the entire global infrastructure, driven by digitization—humanity moving into a digital world. These are clearly the most important main thread, centered on the processor. Of course, this thread arrived at a new wave starting around 2005–2006—AI. We’ll probably expand on that in more detail later. That’s one thread—the one centered on the processor.
Then there’s a second, very important thread: the chips brought about by communications infrastructure. Because over the past thirty years, humanity has moved from a non-connected—a physically-connected—world into a world where everybody is connected, every device is connected. That’s the trajectory of the development of the internet.
This trajectory represents another vertical thread. Globally, we went from having no internet, no broadband, no cell numbers, to a state where every person and every device is connected together. And of course this also required a huge amount of chips, a huge amount of infrastructure. And this infrastructure went from very slow dial-up modems on phone lines, to broadband, to fiber optics—that’s the limited bandwidth, and behind it, the trunk lines between cities, between countries. That is what built the entire internet infrastructure.
Of course, with the rise of mobile networks—and later things like Starlink—this wireless infrastructure represents a second wave of the internet. And actually, we’re currently in the middle of that second wave—mobile internet. On one hand it has driven massive infrastructure buildout, but on the device side, the most representative example is the smartphone—it has basically become something we’re accompanied by for six, seven hours a day or more. It has become a necessary part of our lives. This is also a massive, enormous pillar of the semiconductor industry, because at one point the smartphone alone might have consumed 60–70% of all semiconductors, if you count everything that goes into phones.
So those are the two vertical threads. But if we look at it horizontally, organized by time, there’s also a horizontal thread. What we see, organized chronologically, is that the semiconductor industry went through an extremely fervent, rising, prosperous phase. Then it rapidly entered a period of decline—even became a “sunset industry.” The most representative symbol of that sunset period was—probably up until around 2016, before this wave of AI—for a full ten years, maybe even fifteen, Silicon Valley venture capital in the US made essentially zero investment into chip startups. Because everyone believed it was already a finished, sunset industry.
And the most representative case is Broadcom, which invented what’s known as the “Broadcom model”—acquiring relatively mature, profitable semiconductor companies, merging and restructuring them, cutting costs to improve operating efficiency. That is entirely a harvesting model for a sunset industry.
Of course, after 2015–2016, the AI wave arrived. Right now we might be in what feels like an internet bubble, or the second wave of mobile internet—and this AI wave is really the third wave in recent years. This third wave, in terms of the enthusiasm it has stirred up, the amount of capital involved, and how much people expect from the industry, may actually surpass the previous two waves.
So we have to think about a question: why did the previous wave—after the internet—go through ten to fifteen years of this “sunset” period, roughly from 2005 to 2015? From 2005 to 2015, for almost a full decade, VC made essentially no investment, because nobody believed chips had a future.
I think—and this isn’t necessarily politically correct—I think it’s precisely the extreme success of certain sectors that caused the decline of another layer. I don’t know, maybe this is a bit counterintuitive, right?
Zhang Xiaojun: When you say that, I do find it counterintuitive.
Liao Heng: The logic behind this is actually fairly simple. Of course now China has started walking its own industrial model and a self-contained cycle. But if we go back ten years, globally—in tech—it was still largely the American model leading the world’s trends. This American model has one striking feature: when an important, groundbreaking invention or new business gets backing from capital markets, it can rapidly establish a near-monopoly advantage.
Take Google, or Meta—Facebook in social, Google in search—including Windows before that, in the desktop operating system space, and Apple’s iPhone. They each rapidly—in maybe three or four years—established a monopoly-like advantage in their respective fields. That advantage isn’t just technological—more importantly, it’s capital, infrastructure, and user base. People get used to using something and don’t easily switch platforms.
That kind of monopoly advantage actually has an extremely negative effect on innovation. Here’s why: if a customer is the sole purchaser of a certain category of product in the world, then as their supplier, you’re in a pretty miserable position—well, “miserable” might not be quite right, it’s just what the laws of supply and demand naturally dictate. When a customer accounts for the overwhelming majority of purchase volume, they gain enormous pricing power, and they will inevitably squeeze the supplier’s margins down to almost nothing. Second, they will dominate the direction of demand and technology.
This dominance may look completely reasonable on the surface. But let me use a small example to illustrate why total obedience to the customer has no future: in chip design, from when we make the core architecture or product-definition decisions to when the product is actually in end-users’ hands, it takes at least three years. Chip design alone takes over a year. Manufacturing now takes another 9–10 months. Then you have to build it into a complete system, then test it—so three years is the minimum, maybe two and a half if you’re fast, maybe four if you’re slow.
So humans often lack a shared calendar—that’s the wrong way to put it. My personal calendar lives in 2030, but my customer lives in 2026—there’s a four-year gap there. If I could only blindly follow my customer, I would be lagging by four years. So this so-called dominance over technology, and the power to define the future—if it’s placed entirely with one dominant, overwhelmingly advantaged party—will inevitably suppress new innovative forces. Because even if you have the right idea, you don’t get a chance to act on it.
This whole chain of reasoning is why, in the wave before AI, the vast majority of semiconductor companies rapidly fell into decline. One part of it was difficult business economics—no profit to be made. The other part was that even if you had ideas about the future, it was very hard to get the chance to actually build them.
Fortunately, at least in this AI era—starting from hardware, maybe around 2016–17, with the continuous progress of DNNs following things like AlexNet, roughly starting in 2016–17 up to today—we haven’t yet fallen into that kind of unfortunate monopoly, or single-player dominance.
I mentioned that period of decline earlier, and there’s a very interesting phenomenon there. It’s that because this field is extremely active, and progress every month, every quarter, has exceeded everyone’s imagination—and of course China is also now an important participant in the world, a real player in the game—right now, at least among the people we can observe, it’s a galaxy of stars, an extremely vibrant field. Nobody can say Sam Altman, or Anthropic, is the ultimate winner, because every day our peers—especially, perhaps, the extremely smart, ambitious young people around Wudaokou, or in Hangzhou—keep bringing us surprises. This keeps monopoly from ever forming. That’s why you see so much firepower present at every layer of this industry.
Zhang Xiaojun: The “dominant players” you’re referring to—are those companies like Google, these giants?
Liao Heng: I think China also has plenty of dominant players. If you shop online, you’ll immediately recognize who the dominant players are. If you scroll short videos, you also know who the dominant players are on the internet-application side. They represent massive scale, massive purchasing power, a massive consumer base—and that also supports enormous infrastructure.
Zhang Xiaojun: I said it’s counterintuitive because in my mental image, upper-layer applications and underlying chip development should reinforce each other. I never expected that once monopolists form at the top, it would actually push the chip industry into a period of decline.
Liao Heng: Then let me ask you a question—you’re also someone who understands this industry—when you think of Google, what kind of company do you think it is? It’s a search company. It’s an advertising company. Even today, the overwhelming majority of its revenue still comes from search-related advertising. I think there’s a striking feature of the tech industry: almost every entity we can recognize made its name with an extremely strong “first move.” Usually it means they invented something unprecedented, gave humanity some wonderful experience, or some capability that people can’t live without, that they enabled.
But it’s very hard for an enterprise to keep reinventing itself, isn’t it? OK, I had my first killer move, and it built me a monopoly position—but do I still have the ability to reinvent myself, to innovate again, to invent something even greater the next time? There are some clever people who can keep breaking through their own DNA, their own limits, and keep bringing surprises to humanity. But for most organizations, the ability to reconstruct and reinvent themselves—to create an entirely new version of themselves—is limited.
If you don’t believe me, run through all the companies you know well in your head, and you’ll find they each have advantages in certain areas—but isn’t an advantage in one place often precisely a weakness in a new area? So today, in AI, the companies making the most breakthroughs are often not the “big names” we’re familiar with. The previous generation’s big names, sure—they have resource advantages, they want to reinvent themselves, right? And we really do hope they reinvent themselves, because after all they have advantages in talent, resources, organization, execution.
But what I want to say is, the world is a strange place—it’s often not the wealthiest kid who ends up the most successful. Think back on your classmates—the ones from the best-off families are often not the ones who end up making the greatest professional or social contributions.
Zhang Xiaojun: That’s often not the case, right? Now—you entered Tsinghua in 1987, and in the “youth class.” What was the computer science department teaching back then?
Liao Heng: I don’t think the curriculum was that different, honestly. Growing up, or during my early years, the machine I used was an Apple II clone. By the time I got to college, the IBM PC had just appeared. But we were still using floppy disks—hard drives were something new. Musk hadn’t started a company yet—was Musk even born yet? He must have been born, but hadn’t started his career yet.
So back then, if we talk about the atmosphere at school, there were a few notable differences. First, the faculty back then were still fairly old-school—they didn’t emphasize publications. The research the professors did was about actually building the machine—at minimum a prototype. You couldn’t graduate just by writing a paper—you had to design something and actually get a rough circuit board built. That’s interesting, because it represents a certain research culture.
Back then, China’s own industry, its enterprises, had very poor research and development capability. And if you go back even earlier than my time, that gap would be even bigger. So universities, or research institutes, were actually playing the role that R&D departments play today. Because enterprises had no R&D, even things like the computers used by the pioneers of China’s “Two Bombs, One Satellite” program—basically everything they used was built by universities and research institutes. That was my teachers’ generation.
By my generation, that remaining culture was still around, so I was right at a transitional period. After my generation, the research system—including the system for training and graduating students—moved to being publication-driven, right, more like the American model, but even more intensified: everyone had to publish at top conferences as the first priority. That was also very much in step with the times.
Because by the time I left China, a company like Huawei might have had only a few thousand employees—what they were doing was still fairly rudimentary; they hadn’t yet formed the ability to organize thousands or tens of thousands of people to develop something extremely complex in an organized way. That system was still forming.
But now, if you fast forward to today, many Chinese enterprises—large or small—have very strong capability in developing things, in knowing how to organize people, organize resources, build relatively mature HR systems to accomplish this kind of teamwork and organizational work. That capability is actually very strong now—even leading the world in some ways. So universities no longer need to shoulder that kind of development-type work.
But going back to that earlier era—it was actually really interesting, because besides going to school, a huge amount of our time was spent hanging around Zhongguancun. It was very much like today’s startup scene—you’d find that there were a lot of needs in the world, and we’d try to develop a product, or at least a prototype, to try to meet those needs.
Zhang Xiaojun: What did you work on back then?
Liao Heng: Countless failed things—laser printers, VCD players, electronic dictionaries—all ancient stuff now. People today probably don’t even know what an “electronic dictionary” is. But back then, enterprise R&D capability was very weak, so university students effectively filled that gap, developing things on behalf of these small companies. Since I was a competition student, I was decent at programming, decent at software.
At the time I had a best friend named Wu Zhao, also a Tsinghua classmate—we were close in age, less than a year apart. He was a hardware competition student, so for many years I really wanted to be like him—to be able to build circuit boards, to design FPGAs at the time—I really wanted to understand what was actually inside a processor, what was going on inside a chip, right? Because chips now have hundreds of billions of transistors—an extremely complex system.
So once I felt I was decent at software, I put in enormous effort trying to break through that layer, that boundary—it felt like the difference between upstairs and downstairs. That was my longing at the time. So later on, I developed a strong preference for wanting to understand what was going on in that “other layer” of the world. And that’s how my career later ended up being, without exception, mainly about hardware.
Zhang Xiaojun: So at school you leaned more toward software, right?
Liao Heng: My starting point was a programming-competition background.
Zhang Xiaojun: How old were the “youth class” students back then?
Liao Heng: Generally you had to be under 15 to qualify for the youth class. I was 14 at the time. Wow, so young! I don’t think that’s such a big deal—most people who end up there got there through some fortuitous combination of circumstances. Getting admitted purely through your own effort is still quite hard, so it’s not really worth making a fuss over.
What I just mentioned about crossing between different “layers”—maybe we can talk more about that later. On the way here you also showed me that Jensen Huang “five-layer cake” analogy, right? In our minds, it might be more like an 18-tier pagoda, 18 stories of a pagoda.
Zhang Xiaojun: We’ll come back to that in detail later. Going back to ‘87—as I was studying this chip history, I found it really interesting, because TSMC was also founded in ‘87. Morris Chang was 56 that year. TSMC pioneered the wafer foundry model, which later rewrote the entire division of labor across the global chip industry—because before, it was vertically integrated, and afterward it split into design companies as one category and foundries as another.
Liao Heng: By the time I entered the industry, TSMC was already fairly powerful. It hadn’t yet formed today’s kind of overwhelming dominance, but it was already quite strong.
Zhang Xiaojun: How did this model form? What’s the root cause? Why did it fit that era so well?
Liao Heng: I think the most fundamental reason is economics—this is really an economic principle. A wafer fab’s capital expenditure keeps growing larger—every generation, I’m not sure exactly, but at a rate close to Moore’s Law. A process node that used to require maybe ten mask layers now requires a hundred. A machine that used to cost a million dollars, or ten million, now costs a hundred million or a billion.
You could say that fab, foundry manufacturing, is a business with huge capital investment and an extremely long R&D cycle—it’s like the sub-basements I’ll describe later, floors five and six—it’s a business with extremely high risk, poor ROI, and a very long investment payback period. And as the process node evolves, it gets worse. So at that point, as a design company, you don’t have that kind of capital capability.
It’s like—everyone at home needs to use air conditioning, needs electricity—you’re not going to build your own power plant just because you need electricity, because the expense of building that plant is too enormous. So it was Morris Chang who first recognized this—that everyone wants to share this infrastructure publicly. You can find many interviews with him on Bilibili where he reflects on this history. But he was the first to recognize this, and to explicitly turn it into a business model—that was a great contribution to humanity.
Of course this contribution has also allowed countless companies to build success on top of this model. I can say this: without the fabless model, even a giant like Google wouldn’t have been able to build the TPU. Because if they had to build their own fab, that alone might take three years. Developing the process node might take another three years. Getting the fab to run normally would also take a long time. And their scale might not even support the economics needed for a leading-edge fab.
But once this model was established, it led to everything that followed. Frankly, my entire career, up until the last five years, took place under this model. This model has existed and has its own rationale, but it’s not “everything.”
Zhang Xiaojun: AMD founder Jerry Sanders also famously said “real men have fabs.” But in the end, even he spun off his own manufacturing division.
Liao Heng: This model has its merits, but it doesn’t mean everything in the future has to follow this model. Because what we call rationality has certain preconditions. Once those preconditions change, you have no choice but to adapt.
Zhang Xiaojun: What was the precondition at the time?
Liao Heng: The precondition at the time was global liberalization—”the world is flat”—everyone divided up labor according to their own specialty. I wanted to source the best capability available at each layer, worldwide. At HiSilicon, before 2019, almost 95%, even 99%, of our designs were built on the world’s best suppliers—I would go down to the floor, pick the best fifth floor, sixth floor, to build my products. But this precondition was clearly overturned by competition between nations, geopolitical factors, and so on. That’s the first thing.
Second, when something develops too rapidly, it goes through a period of specialization that can’t necessarily meet requirements. For example, the extreme memory shortage we’ve seen these past six months—DRAM prices have gone up by ten-plus times. Vertical division of labor to some extent can’t respond well to that kind of overly rapid change—it’s like the stock market. When you get an explosive crash, a “Black Friday,” nobody can react in time. Or when you get a spike in demand, normal operations can’t keep up with the pace of the surge.
Zhang Xiaojun: When did you realize you might work in the chip and semiconductor industry for your whole career?
Liao Heng: It was already at Tsinghua that I especially wanted to do this—because of that longing I mentioned, that longing for understanding “downstairs.” And I felt “upstairs” was more comfortable, the view was better, you could see farther. If you’re building an application, as long as you have the right concept, you might in half a year reach tens of millions of daily active users. Most super-apps came about that way. But “downstairs” represents a more patient, long-distance-runner kind of player.
Zhang Xiaojun: Are you that kind of long-distance runner?
Liao Heng: I don’t know. Throughout my career, most of my colleagues would say, “You’re not really a chip engineer—you’re a wolf wearing a chip engineer’s sheep skin.” Because I’m not really doing pure chip work—I also do software. When I’m with my software colleagues, they’d say, “you’re the guy who does chips, downstairs.” Or sometimes they’d say I’m doing algorithms.
So I think this kind of multi-faceted character is something chip work needs—you need to know those things, or at least be able to hold your own talking with people in other domains.
Zhang Xiaojun: I’m curious—in ‘96 you went to the US, initially to do a postdoc. At that moment, what was your vision of the future of technology? What kind of mood was everyone in?
Liao Heng: That mood is a bit complicated to describe. I held two thoughts at the time. First, I felt I had to go—because even now, when I interview many excellent candidates who already got their PhDs at top schools, proving their ability is not inferior to anyone else’s, many of them still feel they have to go do a stint at MIT—otherwise they’ll feel inferior. I had that same mentality—because America, as the technology leader, had been that way from—I’d say—1945 to the present, so many years. There’s a kind of regret people feel: if you haven’t experienced it, you’ll never know, you’ll always feel like you’re missing something.
That’s the first layer. The second layer is: I felt at the time that the startup-like things I mentioned before, all those messy product-like things I made, none of them ever sold a million units, or had countless people talking about them enthusiastically. What I built never really went to market. At the time, I genuinely didn’t know what had gone wrong. I didn’t know. But after working for a year or two, I quickly found the answer.
But when I went, I still carried that confusion—why did the things I made never turn out good enough, or never make it to market? Actually the answer is quite simple. After working for about a year and a half, I quickly understood what the missing piece was.
Even to make an electric kettle, you have to ensure that even if one component fails, it won’t keep dry-burning forever—otherwise it could cause a fire, even kill someone. To build an electric kettle, you have to pass safety certification—even if water splashes on the side, it shouldn’t electrocute you. If the relay fails, it can’t keep dry-heating—it has to be inherently safe, must not feel dangerous. That whole process—for something to be usable, you not only need a smart-enough head to imagine how to design a kettle, or how to design a SpaceX rocket, you also need to reliably execute it, and you need enough verification and testing, before it can finally meet the standard of being usable by the general public in the market.
Back when I was a student, I hadn’t realized this at all—I thought if I had a clever idea and built it, that was it. But actually, we can only say we made it to a prototype. You can’t imagine—something like the phone you’re using might require tens of thousands of person-years of R&D per generation, just so it doesn’t do things like go black-screen the moment it gets hot, or shut off the moment you take it to a ski resort in the northeast, or fail to charge properly. There are countless things that require a rigorous process—a strict process. Only after going through that kind of process, that cycle, can a product be polished to the point of being usable.
I recall our autonomous driving effort—it was incubated within our team initially, but the whole process took a full seven years—from having a prototype that could drive around our own campus, to finally launching and selling to the first paying customer, a long seven years.
So as a student, I had zero awareness of this. I naively assumed that because I was smart, or good at programming, being able to build a circuit board meant I could make products.
There’s another related, interesting topic—I realized after working for a year that if you want to organize—never mind ten thousand people—even just twenty people to make a product, you can’t have every single one of them be a top competition-level genius. You need ordinary people, responsible workers, whose intelligence and skills are just the product of average education and your hiring pipeline, to be able to work effectively, and to combine their efforts. You can’t pick nothing but top talent for every role—and if everyone were a genius, that would actually create problems too. If you have twenty geniuses, you definitely can’t manage them—they’ll generate all sorts of conflicts with each other. That’s another aspect related to a company or team.
So everyone has their own way of contributing, and this is actually very healthy, whether for a company or a society—otherwise only the strong survive. If someone is just slightly better than others, or has a slight advantage in one area, they’d crowd out everyone else’s room to survive. But actually it’s the opposite—you need a large number of responsible, conscientious people for things to ultimately get done.
Zhang Xiaojun: In ‘96 you went to the US, to Princeton to do your postdoc. What was your first impression of America?
Liao Heng: I’d actually already been to America once before that—back in ‘87. By ‘96, nine more years had passed since then. The first time, I felt America was a fantasyland—because I was only 14 at the time, and everything felt like a dream, since the gap between China and the US was so enormous. Silly as it sounds—when I got back, I told people, including my future wife, “America is amazing—they even have air conditioning running on the street.” Actually, when I revisited that memory later, trying to figure out why I’d had such a stupid impression, it turns out we’d gone to a shopping mall near Stanford, in Palo Alto. And because China didn’t have shopping malls at the time—we only had farmers’ markets, stores lining the streets—a shopping mall felt like a street with a glass roof built over it. So my mistaken impression at the time was that America was so advanced that it even had air conditioning blowing on the street.
But by ‘96, I was already an adult. First, I was disappointed—because looking at the world as an adult versus as a child is very different, and I felt a deep sense of not belonging, because Princeton has a very elite, American aristocratic feel to it, a different character than other schools. Later I would come to appreciate many valuable, excellent things about Princeton. But at the time, I only felt deeply that this was a place for the American elite to thrive—and I didn’t belong to that class. So I wanted to quickly leave that place.
Zhang Xiaojun: So you only stayed a year before leaving right away.
Liao Heng: Right. And there was a second disappointment: I originally thought I might do a postdoc and maybe have a shot at a faculty position. But by then I fully understood that if I wanted a faculty position, I’d need to redo my PhD—because Tsinghua’s brand wasn’t strong enough at the time. Though today the world is different—Tsinghua’s academic standing, its integration with the world, has drastically closed that gap, even taken the lead in some areas. But back then, I still carried a lot of inferiority—I won’t say it was pure inferiority; maybe it was a mix of arrogance and inferiority.
Zhang Xiaojun: Did you think about going back to China at the time?
Liao Heng: No, I didn’t think about it at the time. I think that was a wrong understanding—I thought going back wasn’t an option, so I could only stay and endure it. I didn’t want to feel that sense of defeat. But later events proved that way of thinking was very foolish. There was no way around it—when you’re young you make some foolish mistakes; I don’t think it even rises to the level of a proper “mistake”—I’d just call it an error.
Looking back now, at classmates my age, or ten years younger, who pursued this career in China—even the “average” ones, the responsible, hardworking, not-lying-flat kind—they’ve done very well. Why? Because a person’s growth is partly about themselves, but a much bigger part is the broader environment. If the whole environment is rising, then naturally you’re on an escalator that’s rising—that shared progress. In America, from when I first went in ‘87, or again in ‘96, progress was slow—even in decline, in some ways. But China, over these past thirty years, has progressed very quickly. Of course, I was fortunate enough to come back for the second half of my career and take part in some of that.
Zhang Xiaojun: As I was researching this chip history, I found that ‘97 also had a big event in the chip industry—in April ‘97 Nvidia launched its third-generation chip, the NV3, and only then, after the failure of its first two generations, did it finally get a firm footing.
Liao Heng: Did you notice that company at the time? No, I didn’t—I think back then they were a nobody. Actually, I think if we’re talking about chip history in ‘97, there was a much more dramatic drama unfolding—but it wasn’t in the thread you just described. That was a peripheral thread, not that important at the time. What was actually the main event playing out then was a very important moment in CPU history.
Because at the time, some of my officemates at Princeton—people working on processors—most yearned for a company called DEC—Digital Equipment Corporation. DEC may have vanished from history by now, but at the time it was second only to IBM. DEC made mini and mid-range computers, based in Boston. This was in DEC’s twilight period. Digital had a flagship program called Alpha—this processor represented the state of the art in CPU design at the time. So everyone working on processors wanted to join that team—it was like today wanting to join Nvidia’s Rubin or Feynman program—that represented the world’s highest level at the time.
Let me back up a bit—this ties into what I said earlier about how once a monopoly forms, it suppresses innovation. The CPU originally started out single-core, single-issue—executing one instruction per cycle, sometimes taking many cycles per instruction. Then gradually RISC processors emerged—reduced instruction set computers—and the inventors of RISC are still very active today: Professor David Patterson and Professor John Hennessy, who later became president of Stanford. They led the effort to simplify processors, making each instruction simple, so one instruction could be executed per cycle.
Then in the ‘80s—maybe ‘84, ‘85 to ‘87, ‘89, ‘90—multi-issue emerged. That means a single cycle could execute multiple instructions—multicore hadn’t appeared yet. At the time there was a very important competition—one that remains perhaps the most critical topic in processors even today: if a processor can execute multiple instructions per cycle, how do you schedule it—how do you generate the software, compile the code, so that multiple instructions execute per cycle, making it faster?
Back then, DOS plus Windows had already formed a kind of monopoly on the desktop CPU. But in the server space, it was still a hundred flowers blooming. The hottest companies at the time were Sun Microsystems, SGI, and a whole series of others—including Digital, still around, running various flavors of Unix. That was the golden age of processor development, lasting roughly ten years. And the biggest debate at the time was about static-scheduling processors—VLIW—versus superscalar.
VLIW means the compiler pre-arranges instructions, and every instruction issued actually represents multiple instructions executing simultaneously. Superscalar, on the other hand, is entirely dynamic—you don’t pre-arrange which instruction executes in which cycle; the processor itself fetches instructions and places them into a scheduling window, dynamically scheduling them.
Why do I bring this up specifically? Because this exact topic reappears later in our AI processor story. Today, all AI processors are statically scheduled—including the world-leading Nvidia GPUs—they’re actually still following the VLIW path, not the superscalar path.
That debate went on for a full decade, because software and IT systems hadn’t yet formed a monopoly—so besides Sun and SGI, there was also HP, IBM, Digital—a “warring states” situation, essentially. It was like China’s Spring and Autumn / Warring States period, with a hundred schools of thought each vigorously trying to prove they were right. But to fast-forward the story: it ended in an epic-scale disaster. First, DEC went bankrupt and was acquired by Intel. When they sold off their assets, the Alpha processor was acquired by Intel—for very little money, probably.
DEC’s biggest asset was actually a search engine called AltaVista—it predated Google by about six or seven years. Search engines back then—AltaVista pioneered that space; the only problem was they didn’t know how to make money from it. Although it was sold off for a lot of money as part of the asset liquidation, nobody ever figured out how to turn a search engine into a profitable business. Technologically, it was a pioneer. Later, smarter people invented better search algorithms and found Eric Schmidt, found the business model—so Google became a giant, and AltaVista disappeared into history.
What I want to say is, that debate still holds deep instructive value for us today—many lessons to learn from. And as we continue moving forward with the evolution of AI processors, we can very likely draw effective direction from this history.
Zhang Xiaojun: What was your view at the time?
Liao Heng: In that debate, my own view didn’t really matter, because at Princeton I was also working on VLIW. I had never built a chip before—never done a full-scale chip that actually needed to be taped out and manufactured. So as a so-called PhD student, I had no real judgment—very naive. To have real judgment, you need real understanding—real awareness of reality. You have to know things like “it’s hot today, I should wear short sleeves”—a sense of the world around you, of the environment. When it’s cold, you need to sense it. Back then I had none of that kind of experience.
Zhang Xiaojun: When you felt Princeton wasn’t quite the right fit, you chose to go into industry.
Liao Heng: I basically just went and found whoever would take me—whoever wanted me, I went. I didn’t have that many options—we all showed up with something like $1,000, which felt like being a rich person already. Plenty of classmates went with just $100 in their pocket—not even enough to cover a taxi from the airport to school. So back then, we were pretty pragmatic—first thing was to survive. In ‘97 you joined PMC-Sierra —
Zhang Xiaojun:—which had just been through a restructuring. It was a merger between a Canadian company and a Silicon Valley company, and you joined right at that merger point.
Liao Heng: Right, and at first it wasn’t really my active choice—they chose me. I think it was actually a small company—at its peak it may not have exceeded two thousand people. To give you a sense of scale—the department I run today has more than five thousand people, and its revenue is already a hundred times theirs. So PMC might’ve been a small entity, but it went through a dramatic, wave-riding kind of trajectory typical of this industry. It happened to ride the first wave of the internet, because it made transport-network chips—key components in the internet’s backbone infrastructure, metro-area networks, the long-haul transmission networks between cities. That’s exactly what they made. So that wave was much like today’s capital market frenzy—right, back then PMC’s market cap at its peak was, at one point, the largest of any company in Canada. Or the second largest—with the money from being number two, they could have easily bought the Bank of Montreal. Wow. But that kind of illusion quickly evaporated—good years, bad years—I started in ‘87, went in ‘97, and then in 2000 the whole thing burst.
Because you may not know this—you look young, maybe you weren’t even born yet—but in 2000 there was a fantastic—I lived through my first wave of this kind of bubble-and-burst cycle. At the time, the internet had become the world’s number one hot topic; people believed the internet would change everything. So on the application side, Microsoft went all-out to crush Netscape’s browser with IE—because that entry point, that portal—just like today, every internet application is fighting over the entry point for users—back then it was the same: everyone believed the browser was the number one gateway, because everything had to go through the browser, so whoever controlled the browser had the advantage.
At the infrastructure level, everyone believed that everything in the world would eventually flow through the internet—so people frantically, globally, kept overbuilding trunk lines and access networks, so that every city would be covered by the internet, and the fiber between cities would have enough bandwidth. So at the time, PMC’s stock price shot up just like Nvidia’s or Cambricon’s today.
But then, pretty quickly, you’d find—I’m not saying this is the case here—there’s a critical economic measurement: whoever invested has to be able to earn it back. This internet bubble lasted roughly three years before people quickly realized that the investments—for example, the US is still, to this day, using fiber that was laid down back in 2000, unused. Because they suddenly overbuilt so much, and then found there was no demand to lease it, no one renting—it becomes impossible to monetize.
I think this particular problem doesn’t necessarily exist in China’s AI industry today, but it did exist back then in the American internet industry. That was when we went through an internet bubble birth—but later, internet companies actually came back, right? Then in 2005, application-layer internet companies took off again—application companies took off—but companies doing infrastructure never came back up, never recovered.
Zhang Xiaojun: So why didn’t you think about switching jobs, changing companies?
Liao Heng: I think it’s a question of whether you consider it “abandoning” the people you know, versus being abandoned by them—whether you’re willing to abandon them. I never consciously interrogated myself about it that explicitly, but I just felt I still wanted to find another product that could give these chip colleagues, or myself, a chance to build something new and bring in new revenue.
Zhang Xiaojun: Did you find it eventually?
Liao Heng: Later we moved into the IT sector—began doing storage-related work, storage products. And our customer base shifted too—from telecom companies like Cisco, Ericsson, Alcatel, to IT companies like HP, IBM, Dell, EMC, NetApp. So I got the chance to learn what the telecom industry was like versus the IT industry. The IT industry used to be a very lucrative business—companies like EMC, IBM sold equipment with very high margins. But this traditional IT OEM model—companies like HP—started declining once hyperscalers emerged.
Why? Because hyperscalers—super-large internet companies with enormous infrastructure needs—in the global market today, if you look at the procurement volume for servers, or “AI infrastructure” broadly, this category we call OTT (Over The Top), or internationally “hyperscaler,” or super-large internet enterprises—their combined purchase volume is probably over half of the global total, maybe 60–70%, even 80%.
So when this kind of enterprise has such enormous purchasing power, it puts tremendous pressure on the “brand name” OEMs. Because those traditional brand-name OEMs used to have the advantage of being able to make relatively high-quality products—but they served customers ranging from the Fortune 500, expanding down to the Fortune 1000, 2000, or globally—tens of millions of large, medium, small, and giant companies. So they needed a wide-coverage sales-and-service capability—if you’re selling servers into China, you’d need service points in nearly every county, because that county might have a thousand companies as your customers, each only buying two, three, ten servers—but if something breaks, you need someone to show up in person to fix it.
But hyperscalers are fundamentally different from that. First, there are very few of them—globally maybe seven or so in the UK, in the US maybe seven “sisters” or so, China maybe seven or ten too—very concentrated. Second, their geographic location is also very concentrated—all their machines sit in giant data centers; you don’t need to set up a service point in some county in Shaanxi. Third, when they’re buying 70% of what you produce, their pricing power is very strong.
This is another factor I mentioned earlier that drives industry transformation.
Zhang Xiaojun: The last year you were at PMC before you came back to join Huawei—2016—that’s also the year PMC was acquired, for $2.5 billion. You lived through that whole cycle—which was exactly the “sunset” period I described. It probably represented the moment American tech fully gave up on, was thoroughly disillusioned with, the semiconductor sector—because first Wall Street abandoned them; their P/E ratios might have averaged only 3 to 5. What does that mean? If your annual profit is a billion dollars, your market cap is only three billion. For reference: today if you look at Chinese AI chip companies, their P/E ratios might be 600 to 800. First, capital markets had completely abandoned this industry—because there was nothing new, nothing to get excited about.
That’s when the “Broadcom model” appeared—Hock Tan, a genius business leader, invented this pattern of “the small swallowing the large”: go private, borrow a large sum of money, acquire a company ten times your size at a low price, then restructure it, eliminate all redundant departments and personnel, cut unprofitable product lines—and your financials suddenly look great. This kind of restructuring/consolidation approach belongs to the terminal phase of an industry—essentially some people liquidating the industry.
Of course, what you mentioned about Nvidia—it injected an entirely new hope into the industry. So those of us in the field genuinely respect the contribution Nvidia has made.
Zhang Xiaojun: That sunset really was long—from 2000, three years after you joined the company, more than a decade of a slow sunset. What did it feel like living through that? I think most people would have left long before then.
Liao Heng: I think, first, there was fear. Every day, colleagues would gather to eat lunch, heat up their lunchboxes in the microwave, and everyone would be discussing when the next layoffs would come, who it would be. But for me—my wife gave me a lot of support at the time, or convinced me not to give up. Because I felt every situation has two sides. First, it gave me plenty of opportunity to think and observe—to think about what kind of technology, what kind of product definition, actually has a chance to be monetized, or to plan something that combines technology, business, and vision for the industry.
And because I had so much failure experience, so many attempts—and also because, although I was never laid off, right up until the very last day—it also gave me a kind of observer’s role. I visited these IT giants many times—IBM, HP—dealt with them long-term, going to Texas one or two times a month, and to IBM’s North Carolina site. So I learned a great deal, including from the gray-haired engineers at EMC in Boston—I learned from them what this industry was really about, which gave me a much more historical perspective. I also saw many failures—I got to know what kind of things would quickly trigger the alarm bells, what kind of things were simply doomed to fail.
Zhang Xiaojun: During that long sunset period, what were the biggest lessons you learned? Looking back at this experience today, it’s probably very valuable—because without it, maybe there’d be no story of you at Huawei later.
Liao Heng: I think what I learned was—when I was interviewing for the Huawei job, my future manager asked me, “What can you bring to the table?” Because Huawei wanted to enter a particular field at the time, one that was new to them, so they wanted to hire someone who already had experience in the field—what they call a “míngbáirén,” someone who understands. I said I don’t have that much success experience—I have a lot of failure experience, so I can tell you what won’t work. Of course I sometimes misjudge things—my judgment isn’t always right—but I know a lot about how not to fail, so we have to work hard to avoid these failure factors.
Zhang Xiaojun: Can you name a few? Maybe three cases where you could immediately tell something wouldn’t work?
Liao Heng: For example, at first we wanted to build an ARM server—a CPU to replace x86 processors. My first instinct: this will fail—this is the flagship product of what’s now our “Turing” business unit today. At the time, I kept telling my leadership: absolutely do not do this, there’s no market for it, no customer demand. Of course, let me be humble here and admit I was completely wrong. Because CPU supply was plentiful, and while the x86 ecosystem—AMD wasn’t as prominent at the time, but Intel’s ecosystem was something everyone was used to. To change people’s habits, you usually need several times the advantage—you either need to be a third of the price, or three times the performance. At the time we didn’t have that kind of value advantage—that was my logic. But that logic turned out to be wrong—events later proved I was quietly wrong.
So at the time, my advice was: we absolutely need to find where the customer is. At the time, I thought the only possible customer base was—because China’s public cloud didn’t have significant scale yet. Cloud companies had already been founded, but hadn’t yet formed a good positive cycle. So I said the only opportunity is in the US—we need to immediately send people to Microsoft and Amazon. So we rushed to Seattle. But unfortunately—well, actually, the action itself wasn’t wrong; we did almost get to the point of deploying this ARM CPU on Microsoft’s or Amazon’s cloud. That was around 2017—around ‘16, ‘17, that’s when it started coming together—we had a real chance to enter large-scale, mainstream cloud provider deployment as a CPU supplier. But then the US–China tech war broke out, and that story ended right there. That’s just how things were at that time.
But when this story picked back up again after the tech war, we now see that ARM servers have indeed—some predictions even say within two or three years ARM might overtake x86—in what we call the cloud/data-center market. Because these enterprises, with all their different people, kept working hard, and even Nvidia itself has now made its own ARM CPU. In China too, we’re deploying millions of units annually now.
The reason isn’t purely a “value” argument—not purely “I have a three-times competitive advantage”—there are a lot of macro political factors. For example, our critical infrastructure in China cannot risk something like what happened with Iran’s power stations and infrastructure being infiltrated via IT—we must have our own processors, and our products have also improved rapidly. Even if there’s not a threefold advantage, they’re not worse—they’re fully usable. So the whole of society became much more ready for this, and the products matured a lot too. But that whole process took another ten years. And how many decades does a person even get in one lifetime?
So I said my initial assertion was wrong—but within that shorter time window, it wasn’t wrong. That’s the first one.
Zhang Xiaojun: What about the other two?
Liao Heng: The other two—I think knowing what won’t work, what will work. When we wanted to build the Ascend chip, my leadership said we absolutely need to build a flagship product that can do training, that can be built into large clusters. At the time I tried to persuade leadership: absolutely don’t do this—it won’t sell—because I hadn’t yet seen that day where China’s talent would shine so brightly. Back in 2016, honestly, all the important algorithms and breakthrough inventions really were happening in the US. Chinese people hadn’t yet made their mark in that way—today, in the US, it’s almost “Chinese in America competing with Chinese in China”—that dynamic wasn’t so obvious yet back then.
So I felt building a training-focused chip was a mistake—plus, Chinese internet companies hadn’t yet started training their own models. So I thought this product would definitely fail to sell—we should instead do small-scale production. When we first launched the Ascend architecture, we wanted it to cover a range from one-dollar to hundred-million-dollar products—six or eight orders of magnitude of broad coverage. And we did achieve that—maybe even more than eight orders of magnitude, whether in unit count or price. The smallest product is in our earbuds—the Clip earbuds. The biggest product today is a training cluster with hundreds of thousands of cards.
But at the time, I was very bearish on data-center products and very bullish on edge products. Events proved that today, data center—in terms of revenue, profit, or scale—is far larger than edge-side products. Even though our earbuds sell reasonably well too, so I kept telling people, let’s be realistic, let’s build things that can actually sell. But events proved I was wrong again, right? So even with a lot of prior experience-based judgment, it’s still not necessarily reliable—but it still has some value.
Zhang Xiaojun: Looking at the mistakes you just described, they all seem based on pessimistic expectations shaping your decisions.
Liao Heng: Having lived through the decline of the American chip industry—maybe there is a connection there—maybe it trained a kind of pessimism in me, a tendency to first see the downside of things, to try to escape that destiny, that “doomed to fail” expectation.
Zhang Xiaojun: And all the examples you just gave happened before the tech war.
Liao Heng: That’s right. Including our autonomous driving—I already mentioned it took seven years. We started building the self-driving chip in 2019.
Zhang Xiaojun: Before the tech war, when the external environment was still relatively favorable, the choices you all made were actually small, cautious choices—you didn’t dare to bet big.
Liao Heng: At the time, I have to admit, some of Huawei’s key leaders—including HiSilicon’s leadership, or certain group-level leaders—had a much bigger vision than I did. Because I was like a loach in a small pond, suddenly dropped into this vast ocean—I hadn’t yet learned to see things from that grander perspective. It was only after 2019, once we were put into that kind of difficult situation, that we were forced to learn to look at problems from a much bigger angle.
Like the examples I mentioned—at the micro level, they didn’t look wrong at all. But at the macro level—I said my judgment was wrong because I was wrong at the macro level—I didn’t see the bigger picture, didn’t foresee what the future would become. This is really about shallower experience, or having sat in a small pond, staring at the sky through a well, for too long, which naturally creates that kind of limitation. Or put another way, the younger you are, the more prone you are to this kind of mistake. I’m in my fifties now—back then I was only in my forties.
Zhang Xiaojun: I have a technical question—when did people start to feel Moore’s Law was slowing down? What’s the underlying reason?
Liao Heng: First, Moore’s Law has been slowing down regardless of whether there was a tech war—that’s an objective fact. Moore’s Law has three dimensions: economics, performance, and energy efficiency—what we call PPA. The area keeps shrinking, so the average cost per transistor keeps falling. That’s what made electronics the one product category in the world that’s deflationary rather than inflationary. If you buy anything today versus ten years ago, it’s almost always more expensive now—except electronics. That’s the biggest contribution of Moore’s Law: it’s anti-inflationary. Electronics bought ten years ago have almost no value today, apart from maybe collector’s value—that’s an economic benefit from Moore’s Law.
Second is performance—the assumption used to be, the smaller the transistor, the faster it runs, so performance improved. Third is energy efficiency—energy cost—like tipping over a smaller bucket of water uses less energy than tipping over a bigger one.
But around 7nm, or even 16nm, Moore’s Law’s economic benefit basically stalled—the cost per transistor stopped falling and actually started rising from that point on. So the economic benefit disappeared. The performance benefit has also become very small—it’s still improving slowly, but not dramatically. What’s really still delivering benefit from Moore’s Law today is energy cost—because that “bucket” got smaller, so each time you “tip it over,” the energy consumed per switching event is smaller—the picojoule or femtojoule energy cost per switch keeps shrinking.
This actually gets to something quite fundamental. I can tell you this plainly: the gap between Chinese semiconductors and TSMC’s semiconductors is mainly in this energy cost. For example, if you ask me to provide 100,000 GPUs’ worth of compute, I can easily provide that—it’s just that compared to the world’s number-one product, my energy efficiency is a bit worse. But we have abundant energy. So for us, for China, that’s not necessarily a huge problem.
There are actually several dimensions to this question. First, people shouldn’t worry that if we were completely cut off, China would have no compute at all—clearly, the answer is no, we absolutely have compute. Second—is it more expensive? Not necessarily. Third—does it use more power? Yes, it does—no way around that. But you mentioned China has abundant energy, right—our power supply may be triple America’s, and we have a lot of spare energy. China’s data center electricity costs might be a quarter or a fifth of what they are in the US, Singapore, or elsewhere around the world. On average, our energy costs might be about a quarter or less of theirs. So first, we’re not trying to use this argument to convince people to accept something that uses more power—I can only say that for a baseline guarantee, there’s really no need to worry too much, and China’s energy supply is more than sufficient.
Moore’s Law continues, of course—it keeps developing—but I think a more precise way to describe it should no longer use nanometers as the unit of scale. I think we should describe it using atom count instead—because an angstrom is a tenth of a nanometer. If you’re describing things at the angstrom scale, you’re already at the atomic scale. So we know Moore’s Law will eventually reach an end—and now, because the smallest feature size of a transistor has already reached, say, the angstrom level—this is a scale measured by number of atoms. From now on we should really measure it by number of atoms, because ultimately, you can’t shrink below a single atom—it’s either there or it isn’t; that’s its ultimate endpoint.
But I think we should still expect Moore’s Law to keep going for now—its economics will just keep getting worse, keep getting more expensive. And here’s a second point I mentioned: if people are used to measuring how advanced a transistor or a piece of silicon is by a single dimension—a spatial one—maybe we should switch dimensions and ask instead: can the operating speed get faster? Switch from a spatial scale to a time scale—because what we actually need is more computation done per second.
And can we switch to an energy scale—how much energy does it take to flip a transistor once? Once you switch scales—measuring in picojoules or femtojoules—you find there’s a fundamentally different picture. It’s like measuring a person—is he good-looking, or is he smart, or is he good at math versus good at language? Change the subject and the same person might score differently, right? So which “subject” you’re testing matters a lot. What we call “Yao’s Law” measures the time dimension—I want faster circuits. We could also switch to a “Joule’s Law”—measuring how much energy it takes to do the same job with less consumption.
Change the question, and the answer changes too, doesn’t it? There’s actually a lot of interdependence between the answers, even though they’re different. Because anyone with an engineering background knows the basics of circuit theory: what affects a circuit’s switching speed is capacitance plus resistance. Moore’s Law faces two factors that are certain to get worse: if a wire gets thinner, its resistance definitely goes up, which slows things down. Second, capacitance: the closer two plates are to each other, the larger the capacitance—so from that angle it’s getting worse too. But as the plate itself shrinks, capacitance also shrinks. I can tell you why Moore’s Law stopped delivering speed gains after 16nm or 7nm: fundamentally, resistance got worse, and capacitance didn’t meaningfully improve—because one factor (smaller plate area) should reduce capacitance, but the plates being closer together increases it—the two effects cancel each other out.
Actually, all semiconductor progress has really come from shrinking capacitance, while resistance keeps getting worse—that’s the essence of Moore’s Law. So when we see that shrinking the die doesn’t make it faster, the main reason is that capacitance only shrank slightly, because two opposing factors partially cancel each other, so there’s no dramatic speedup. But because the overall thing shrank, the number of electrons it needs to release also decreased, so there’s still an energy-efficiency gain, but not a speed gain. Manufacturing something this tiny definitely costs more—that’s the difficulty Moore’s Law faces today. But this difficulty doesn’t really matter, because the whole world still expects this industry’s scale to be enormous, so people should keep pushing forward.
But like I said—if you change the question being asked, doesn’t the answer change too? Let me give you a small example: if we want to reduce resistance, we should make the wire thicker, right? But if we want to reduce capacitance, we want the overlapping area between the two facing electrodes to be as small as possible. Think about it—if two wires run parallel to each other, their overlap area is large. If we turn them perpendicular, the overlap area is reduced to just the intersection point. So a structural change can bring miniaturization. If I don’t want to change the structure and just follow the existing layout, there’s one kind of answer. But if I say I can’t shrink any further and I switch the two parallel lines into perpendicular ones, wouldn’t the speed increase significantly? The answer is yes, definitely.
This reminds me—you might have asked me this before, a kind of “cold” question, where the answer isn’t obvious. Let me throw you a cold question—it’s really about illustrating the difference between Moore’s Law and “Yao’s Law.” Take a wild guess: how many patents exist in human history for mousetraps? Just guess. 300? I don’t actually have the exact answer either—you all can check with Doubao, or search a patent database. I recall that around 1900, the head of the US Patent Office actually wrote a letter to Congress proposing that the Patent Office be abolished—the organization no longer needed to exist, because humanity was so clever that it had already invented everything worth inventing. One example given was mousetraps—apparently there might have been a thousand patents by 1900, for various ways of killing pests, killing mice, because humans generally hate mice.
Why is this an interesting question? Because back in 2020, I put a lot of effort into learning how lithography machines are made. I found a mechatronics textbook—mechanical-electrical systems design—and it left a deep impression on me, because the very first page had an image showing ten different ways to kill a mouse. I can’t find the book right now, but let me try to describe it. If you’re a chemical engineer, you’d think: to kill a mouse, I need to make poison—put something the mouse loves to eat, whether cheese or a piece of meat, and lace it with poison, and the mouse eats it and dies. If you’re a mechanical engineer, you’d build a trap—or a hundred different kinds of traps—the mouse comes to eat the meat, triggers the mechanism, and either it gets crushed, or it gets trapped in a cage it can’t escape. All kinds of approaches, right? If you’re an electrical engineer, you’d rig up a high-voltage setup—the moment the mouse crosses a certain point, it triggers 1,000 volts and instantly kills it.
What I’m trying to say is: Moore’s Law, or “Yao’s Law,” is really just engineers approaching problems the way Edison would. First—what is the problem? Second—have I found an effective way to address it, to confront the problem and find a good answer? Of course, if you want to build a product, you also need economic viability—the solution needs to be affordable enough. Fourth, the solution needs to be reproducible—you can’t say the quality is inconsistent, that this batch works but the next batch is unreliable. All these factors stack up—this is what I meant about “preconditions.” Because if we sit here and think about how to replicate TSMC’s or Nvidia’s capability—first, we don’t need to. Second, no matter how hard we think about it, we don’t have their preconditions—you don’t have that context.
I’ve said a lot here, and it’s already drifted into a somewhat metaphysical, or even philosophical, perspective. That was us reviewing the chip history you’ve lived through.
Zhang Xiaojun: For the second part, I’d like you to act as a tour guide—walking us horizontally through this industry, this whole supply chain. Because you told me it’s an 18-story pagoda—you don’t fully agree with Jensen Huang’s “five-layer cake”—you think it’s 18 layers? Take us on a tour, give us a horizontal overview.
Liao Heng: We mentioned earlier—of course, maybe it’s a cultural-background thing—Jensen probably eats a lot of French pastries, so he uses a cake as his metaphor. As a Chinese person, when I picture a layered structure, my first association is something like Yingxian Wooden Pagoda—a pagoda. But really it’s describing the same thing, isn’t it?
If we set aside the pagoda metaphor itself, you’ve probably all heard of “co-design”—the idea of collaborative optimization: software-hardware co-optimization, or algorithm-and-chip-infrastructure co-optimization. The moment you’re discussing co-optimization, you’re already touching two layers, right—the 7th-floor colleague and the 6th-floor colleague know each other, understand each other’s difficulties, and have to work together, each making concessions and accommodations. “You don’t have enough memory bandwidth here, so I’ll try to save bandwidth in the algorithm.” “You have surplus compute over there, so I’ll spend extra compute to save bandwidth.”
Let me point out something very intricate, very micro-level. If you look at Blackwell versus what we’re producing, you’ll find a difference that maybe no one has ever described at this granular a level: vector compute versus cube compute. Because “cube,” or what they call the Tensor Core—we call it “cube” too, it describes the same 3D matrix-multiplication unit—say their ratio is 32:1 and ours is 8:1—what does that mean, what’s the significant implication?
I think this gap is a bit like—imagine a family of four that can only afford a 100-square-meter apartment when they’re young. Later, your career goes well, you get a bit older, you’ve accumulated some wealth, maybe you buy a 200-square-meter apartment. Maybe five years later you’ve done even better and buy a 400-square-meter villa—but you’re still just a family of four. You’ll find that once you live in a 400-square-meter villa, you might accumulate a lot of clutter—like delivery boxes that never get opened, just piled up in some corner nobody visits—lots of redundancy. What I want to say is, once you go back to living in a 100-square-meter apartment, you can still get by—you just have to be much more careful.
An 8:1 ratio versus a 32:1 ratio means: when you’re utilizing that compute, your model design absolutely has to take that into account. If your space is abundant, you can afford to waste it however you like. If your space is scarce, you have to think really hard—you have to fully make use of every bit of that space—you either build something economical and efficient, or you build shelving, storage racks—clear out everything unnecessary.
You’ll find, if you look at the DeepSeek model, I think they made some very deliberate, very conscious choices. I think their sense of direction is extremely clear. For example, before 2025, they already recognized: we must save compute. Because everyone believes in Scaling Laws—more parameters bring stronger capability—but does every parameter really need to be brutally computed every single time to gain that capability? Could I compute just 1/32nd of them instead? I think DeepSeek’s V2, or V3/R1, thoroughly proved this point—you don’t need to. This is what’s called “sparse activation” of parameters—I only need to intelligently select a fraction, say 1/32nd, of the parameters to participate in computation, and I’ve saved 32x the compute—and the cache is saved 32x too. Once someone consciously confronts this problem, they gain that 32x acceleration.
Then, from 2025 into 2026, they made another very direction-focused move: if the sequence is very long—say a 4K sequence versus a one-million-token sequence—those might differ by 256x. Do you really need to pay that 256x compute cost to process a long sequence? Actually you don’t. They did compression, sparsification—out of 256K tokens, I don’t need to compute every one; I just need to select something like 1K or 512 to compute, and I save that much compute. But that selection process comes at a cost of added complexity.
So you can see this model was very deliberately, very consciously designed through co-design. What’s interesting is: it’s like we only have a 100-square-meter apartment, and we still have to raise a family of four in it, and we still want our kids to be exceptional—not worse than Musk’s kids—but it requires much more effort, more blood, sweat, and greater complexity. And that complexity actually comes from what I mentioned earlier—vector computation. So you could say, from this angle, DeepSeek is very well-matched to an 8:1 compute ratio, because to select what actually gets sent to brute-force computation, they had to pay a higher cost in vector computation to pick it out. I believe Anthropic’s models, or OpenAI’s models, don’t need to do this—because the machines they buy already have 32x CUDA compute.
So put another way: our chip happens to have a quarter of Blackwell’s specs—but when running DeepSeek, it seems to work just fine.
This is co-design—this is co-design across two layers—the algorithm colleague realized where the model’s advantages should be built, right? That’s designing at this co-design level, and it’s precisely this kind of thing that means, when I run inference on something like a DeepSeek v4 Pro, it doesn’t feel to me any slower in token generation than what I use every day—Claude 4.8, or something else. I don’t feel any lag, right? Of course we still need to respect what others do well, catch up where we’re behind, and surpass where we can, right? But I can say—whether we’re doing chips or models—if you have this level of consciousness, this understanding of the preconditions, your direction of effort will end up dramatically different.
I have to give President Liang—Liang Wenfeng—a big round of applause, because he consciously, actively chose a higher level of complexity, a harder-to-converge algorithm. Algorithms this complex can be difficult to converge, generating all kinds of anomalies—but he worked hard to overcome those problems, because he had an early sense of the coming constraint. He didn’t wait until five years later when he was suddenly hitting a dead end with zero compute—he foresaw the problem and proactively went about solving it. This is the value of a trailblazer—it leads others to keep working hard along that same path.
Before 2019, we were also living in a kind of “happy” world, because back then HiSilicon, though very low-profile, very quiet, was already at massive scale—probably one of the largest in the world, in terms of wafer procurement volume and product variety, probably top 3 globally. It’s just that the scale of the semiconductor business only served Huawei’s own products. But what I mean by “happy times” is: everything we purchased was the best in the world, because we could afford it. For example, our wafer usage—we used TSMC’s most advanced process—even more advanced than what other companies, including Nvidia, were using at the time; we weren’t the first to use it, but the second. Because our products’ positioning could support that kind of spending, right? But once that was cut off, we suddenly had to face much harder challenges.
The second lesson I learned over these years is: it’s also something we could do. When you’re forced into a corner, you have no choice but to learn how “downstairs” really works—what exactly it takes to master a 7nm, 5nm process. I think maybe the most valuable insight is: what looks like an impossible problem, once you break it down, becomes ten, or a hundred, more concrete problems. Then you break those hundred problems down another layer, and they might become a thousand physics, chemistry, and math problems.
That breaking-down process turns a macro-level problem into something manageable. For example, today you might say, “can you match TSMC?” The simple answer is no—but that answer doesn’t actually help you. So what do you do? You start breaking it down a layer—you find that if some parts are impossible, what parts are possible? So an unsolvable problem becomes concrete, tangible. Once it’s made tangible, you might be able to solve a hundred of these sub-problems—maybe solve eighty of them—and you find you’re not actually in that bad a shape. You keep improving on those eighty, and you find that for ten of them, you’re actually doing better than others. Maybe you can play to your strengths to compensate for weaknesses.
So—what are the 18 layers? The very top layer is applications—like Doubao, or WeChat, or Alipay from Alibaba—these are application layers, the tip of the pyramid, because they can be monetized directly. They have their own technology, of course, but it’s not necessarily some Einstein-level technical breakthrough—it’s more about operations or business design. The very bottom layer is very physical, tangible things—like mining ore, going to dig a mine in Nigeria, or whatever—because at the very bottom it’s really physics, chemistry, mining. Because at bottom, a semiconductor is essentially a pile of sand plus certain rare metals.
Going up a bit, there’s device design—for example, FinFET transistors, or GAA transistors—or, like I mentioned, if I can change a transistor’s design from two parallel lines to two perpendicular ones, doesn’t that potentially give a 10x speed boost, because capacitance shrinks significantly?
Going up further: how do you manufacture trillions of transistors on a single wafer, reliably, so they’ll keep working for five, ten years without failure—that’s extremely difficult. Fortunately, the Chinese nation has many extremely capable people—some Taiwanese, some mainlanders—who collectively hold this capability. When they hit obstacles—being limited by equipment, not having the most advanced lithography machine, not having hundreds or even thousands of types of the most advanced equipment—countless people are working hard to build these tools, one customer at a time. What I’m describing might already be around the sixth layer, right—this process-node layer—can you make a reliable 5nm process? Can’t do 3nm—what do you do? Can’t do a single wafer layer—can you stack layers? Stacking requires what kind of equipment? How do you reliably stack these layers without causing thermal problems, power delivery problems—that’s the fifth, sixth floor problem.
Back up to the seventh layer—that’s roughly where we say chips live—how do you best make use of what’s below you, while confronting your own physical constraints? Because—as I mentioned—if our preconditions differ from others, if our process node differs, where others don’t need to stack, we might need to stack. In the face of these physical constraints, you also need to design, four years ahead of time, the next chip’s architecture, in a way that makes it easy to program.
So above the chip is the compiler, the programming language, the best way to parallelize and partition work, reinforcement learning, how to build KV cache, how to compress models, how to sparsify, how to invent new data formats that let algorithm engineers use lower precision without breaking model convergence. From the algorithm layer up, most people are probably pretty familiar—right, you have a good slow-thinking model or an agentic model, and then how do you build a KV cache system, how do you build an agent framework, and finally build a valuable application. Most people can see those upper layers—so I’d say more than 70% of what people see is up there. Everything below the seventh layer is basically the basement—the people who can see it are relatively few.
Zhang Xiaojun: Which of these layers are you most familiar with?
Liao Heng: As I mentioned, none of these layers do I know deeply enough—so I basically just muddle through, drifting between them, trying to help each layer’s specialists bridge the gap between one another, because it’s very hard to be both deep and broad—it’s like digging a hole, right? You can be a needle—a very sharp, narrow point that can pierce straight through a watermelon—or you can be a knife, and cut the watermelon in half. You can’t really be both a needle and a knife at once—that’s the impossibility of being both deep and broad. But I think what’s actually scarce isn’t specialists at each layer—we’re not particularly short on those, not just at Huawei but across the whole world, because that’s its own specialty, and if you work a certain job for long enough, naturally your understanding grows, your experience accumulates. What’s scarce is the ability to work across layers.
For example, as I mentioned—maybe there are 100,000 AI algorithm researchers in the world, but maybe only Liang Wenfeng deeply recognized that he needed to break through this sparsification problem—trading greater complexity for less compute. That’s a directional choice—most people wouldn’t choose that. So this cross-layer capability is the core of co-design. When a person can cross these floors, they act like a thread, stringing many pearls together into a necklace.
Zhang Xiaojun: Laying out these 18 layers, where does China’s advantage lie? Where are the weak points?
Liao Heng: I think the advantages are quite clear. First, the application layer is very strong—things like Alipay, WeChat Pay, which aren’t nearly as convenient elsewhere. Now I basically haven’t seen paper cash in years—I haven’t touched money in years. But if you travel anywhere else, you still need a credit card, otherwise your hotel might not even let you check in. Or you still need to carry some cash, in case you can’t buy a train ticket in Japan, for example. So the application layer is a strong advantage for us—and because Chinese people have already become so digitized, especially in the consumer population.
Right up through everything before chips, I’d say China’s competitiveness is at least not inferior to anywhere else—and algorithms are even more of a galaxy of stars. I think globally, maybe 70% of top algorithm talent is Chinese. Why? I recently read an article by a Berkeley professor, saying that even in their engineering courses now they have to teach students the distributive law—A times (B plus C) equals A times B plus A times C. So the gap [in basic education emphasis] is enormous. Every family in China places a huge emphasis on their children’s education, and we have this imperial-examination tradition—the idea that scholarship leads to advancement—and the upward mobility path for ordinary people still exists—through your own effort, you can still test into a good university, learn well, still have a shot at success. So China’s talent supply is not lacking—it’s overwhelmingly, crushingly abundant. And of course our ability to monetize that talent isn’t lacking either—because Chinese people all pursue a better life, everyone wants to live better, earn more, achieve greater success, whatever “success” means to them.
Zhang Xiaojun: Chips happen to sit right in the middle of this 18-story pagoda. Below is the basement, above is the building floors—it’s the connecting point. You returned to China in ‘16—did your feelings about the chip industry change a lot between then and ‘19?
Liao Heng: Even though it was only three years, I felt in ‘16 I didn’t believe it—didn’t believe what? Didn’t believe China had this capability. As to why you came back—I came back because of my wife—she was very determined that our kid absolutely must not grow up as an ABC [American-Born Chinese] without their own identity—caught in that worldly contradiction searching for a sense of self.
Zhang Xiaojun: How old was your child at the time?
Liao Heng: I came back in 2008—my child was three at the time. So you came back in ‘08—although you were still with the same company, you were physically in China. Because I think my wife’s understanding came earlier than mine—she was more resolute—she said don’t stay abroad, we absolutely need to start fresh in China. Especially considering the next generation, she was very determined about returning to China. So my return, at the time, was a somewhat passive choice. But by ‘16, when I joined HiSilicon —
I think, in terms of what you’re calling a professional perspective—the bigger shift in understanding happened around that transition point. I think before 2016, I was completely unbelieving—note, foreigners tend to have a kind of arrogance, a blind confidence—the first layer of that blind confidence is believing they are the best in the world. If you work in that environment long enough, you can absorb that same mistaken belief—that certain things can only be done by “them.” Even though you’re Chinese, once you join that environment, you become part of that group’s identity—you’re one of them. So you start believing that difficult things: only “we” can do it, nobody else can. Because as I said earlier, for something very difficult, once you break it into a hundred sub-problems, each one becomes more concrete and solvable. If someone else is willing to solve those hundred problems, maybe they’ll solve them even better than you.
That’s exactly the shift I went through starting in 2016. My first shift was: I originally believed China’s industry at the time was completely incapable of it. But in fact, our results far exceeded my expectations—I find that really interesting. I believe the vast majority of Americans, American practitioners, still hold the mindset I held back in ‘16—and that’s actually given us a huge opportunity. Because if you view them as a competitor, and your opponent seriously underestimates your ability, you have a huge advantage—because they inherently believe they can compete with you while comfortably knocking off at 5pm every day, grabbing coffee, going surfing. Actually, if you just push a little harder, you can easily surpass them. That was my biggest realization: whether in 2016 or years before, our Chinese peers, in terms of effort, in terms of understanding of problems, as long as their identification of the problem is at the same level as everyone else’s, their ability to solve it far exceeds people in that field elsewhere—because they simply work harder. Put simply: one unit of effort, one unit of reward. But the precondition is you need the right people to lead them toward the right problem definition—because if the problem is defined wrongly, all that effort is wasted.
I mentioned Liang Wenfeng’s example earlier—he defined two problems: first, model parameters need to be sparsified—he solved that. Second, attention over long sequences needs to be sparsified—he solved that well too. You see, defining the problem correctly, if we’re distributing credit, probably accounts for 80% of the credit. Actually finding a specific method to solve the defined problem might account for only 20%—and that 20% needs intelligence. Maybe that intelligence comes from an intern, but the problem definition comes from a team leader—because the leader is the one who guides others on what problem they should be solving.
The History of Ascend
Zhang Xiaojun: Now let’s get into the heart of this interview—I really want to hear about Ascend. Because your team has stayed very low-profile these past few years. If you had to describe the ten years from 2016 to 2026 in three words, what would come to mind?
Liao Heng: I think it’s been an ordeal—an ordeal. Or maybe, an ordeal held together by hope.
Zhang Xiaojun: The 910 and the 950—the environments you faced were very different. Do you think the problems you were defining when designing the 910 versus the 950 were fundamentally different? What were they respectively?
Liao Heng: At the time of the 910, we were working with the world’s most advanced logic process—but we were fairly lost about what direction the AI ecosystem and models would actually develop toward, and what applications would emerge. Actually the whole world was somewhat lost at the time too. Although I believe Nvidia was probably the entity that sensed this explosion coming the earliest—but I think globally, at that time, because models weren’t yet capable enough, hadn’t yet produced huge scale—so the confusion back then was mostly about guessing which direction things would go. At the time, we already anticipated the need to do model training, but the inference market simply didn’t exist yet. So we were trying to find monetization paths on the edge-device side—earbuds, phones, things like that. Data center, we thought, would be for training—train the model, then deploy it to all sorts of small edge devices.
At the time, we hadn’t foreseen large language model inference at all—not even things like AI coding—none of that was in our imagination space. So back then, we just had very good conditions but no clear sense of where the biggest monetization opportunity would be.
Then by the time of the 950, the situation completely flipped—right, conditions became extremely harsh. We’re now facing an opponent that’s already the world’s largest company by market cap, the strongest in Silicon Valley. So at this point, the question becomes: how do we survive? How do we serve customers? Maybe we don’t have lofty ambitions to compete for the world’s number one position—the first step is just to survive, to serve customers who genuinely need us, to satisfy that basic principle of being customer-centered.
Zhang Xiaojun: What would happen if you didn’t do this at all?
Liao Heng: We could not do it—what I mean is, honestly, the world would keep spinning without any of us. It’s just that, on one hand, we still hold onto expectations—we don’t want to just lie flat, so we still want to make the effort. And on the other hand, maybe many of our peers have expectations of us too, so we don’t want to give up so easily.
Zhang Xiaojun: What changed in your mindset—designing the 910 versus the 950, before and after the export ban?
Liao Heng: The change in mindset—I think, saying it now, it’s not that meaningful anymore. Try to imagine—if you’re Dong Cunrui about to throw himself onto the explosives—normal people can’t really imagine that state of mind. Or the soldiers at Shangganling—what was their mindset? Or I once saw an interview of a Chinese soldier from World War II, also on a Huawei propaganda poster—a reporter asked him what he wanted after the war ended, and he said he didn’t want anything, because his parents were already dead—no need to think about that anymore. So I’d say this change in mindset doesn’t hold that much meaning anymore. It’s just: when you face a difficulty, you want to solve it.
Zhang Xiaojun: You said when you first joined Huawei in ‘16, you didn’t believe [China could do it]. When did that change into belief?
Liao Heng: That shift happened fairly quickly, because that disbelief was based on a foreigner’s blind arrogance—the belief that the capabilities they hold, others simply cannot possess. Reality quickly educated us on that. Because after joining HiSilicon, I very quickly came to understand that my colleagues—though different from me—for instance, HiSilicon might tape out maybe 100 chips a year. Actually, when our president [He, Ren Zhengfei—actually referring to a Huawei leader] gave the speech at the “Yao’s Law” launch, we’d done 381 chips over five years—that’s an average of about 70-plus new tape-outs per year. But I never once heard of a chip coming back “smoking”—failing testing, unable to be mass-produced. What does that tell you? That kind of hit rate means every team was extremely qualified. Regardless of their background, their resumes, they were all deeply responsible in their respective domains. As I said, being an engineer isn’t something you fully grasp as a student—you only realize it after working for a year: being an engineer first and foremost means being reliable, being responsible—that’s the foundation of everything. So that kind of hit rate, globally, should be considered first-class. That result gave me a rapid education and shift in belief.
Zhang Xiaojun: I believe, regardless of what era you look at Huawei’s history from, the export ban has to be one of the most significant events. When it happened, what was it like internally? What was the initial reaction?
Liao Heng: I think that letter from [Ren Zhengfei] was written before the ban actually took effect—but once the ban actually hit, there was never really a chance to write that kind of letter again. But I already said I can’t imagine Dong Cunrui’s state of mind—I can only say each person, depending on their role, probably felt it differently. I can’t fully describe it for others—I can only speak to my own—my memory of it is a bit hazy now, but I think there were two sides.
One side, we probably needed to comfort the team below us, tell them not to be afraid—because everyone was feeling this enormous fear inside. On my own personal side, maybe I can remember—I think it ignited a huge competitive drive in me—an enormous passion—a determination that this must be solved. First, I need to know what problem we’re actually facing, and I need to break these problems down, one by one, and solve them one by one. For an engineer, that’s a golden, once-in-a-lifetime opportunity. Because normally, you’d never need to learn how a wafer stage achieves nanometer-level precision, how to measure something’s position, how to control a stage moving at nanometer precision while it’s in high-speed motion, and how to locate it—do you measure with laser, or something else—this whole chain of interconnected, fascinating engineering problems that engineers get to solve.
Zhang Xiaojun: Did the team atmosphere change before and after?
Liao Heng: I think the team atmosphere did change quite a lot, but overall it held up well. Because right in the most difficult period, I think maybe only 5–10% of people quickly went off to find new opportunities elsewhere—but maybe 80–90%, especially the strongest core backbone colleagues, I believe to some degree felt similar to what I felt—each of them ignited their own drive to solve problems. So most people chose not to dwell on things too much and instead threw themselves into breaking down the problems and solving them—really interesting problems.
That period was just—every person has both good and bad within them—a desire for survival, for more benefit, but also a vision for something bigger, a greater sense of mission—it’s just about which side gets a bigger share of the weighting. It’s like Gollum in Lord of the Rings—that kind of creature—on one hand wanting the ring for himself, for more power, but on the other, sensing something isn’t right about it.
Zhang Xiaojun: When was the most difficult period, roughly which years?
Liao Heng: I think every phase had its own difficulty—just different kinds of difficulty in each stage. The absolute rock-bottom period was when people genuinely thought that within a few months we might completely lose the ability to manufacture chips at all—the entire business cycle would break. But fortunately, these difficulties were mostly shouldered by a few particularly resilient leaders, who kept it tightly contained within a very small circle—so those individuals bore the overwhelming majority of that pressure, while everyone else was somewhat unaware of the full picture.
Of course there are also some landmark events—though maybe not a single one. For example, China’s 5G network today is everywhere—throughout that whole process, we never once had a supply disruption, never caused the buildout of the communications infrastructure to stall or fall significantly behind schedule. Behind that were countless people, working relentlessly day and night, to make that happen. I can say maybe 5,000, even up to 10,000 circuit boards had to be completely redone from scratch—every single component swapped out—quickly restored to production readiness. That’s the result of an enormous collective effort by an enormous number of people. It’s just that there’s no single defining moment for something like this—when something doesn’t fail to happen, what is that? It’s like—no torrential rain today, no earthquake today—that’s just how things are supposed to be, in daily life. What I mean is: behind that result lie countless people’s day-and-night efforts that made it happen.
Zhang Xiaojun: Where does the name “Ascend” [Shengteng] come from? When was it decided?
Liao Heng: I don’t remember this too clearly anymore, but I believe it drew from the Chinese phrase “the sun rises, the moon is steady” (日昇月恒)—expressing a beautiful hope for the future.
Zhang Xiaojun: You just talked about the overall architectural design philosophy. From generation to generation—the 910A, B, C, and now the 950—what were the architectural changes and evolution like?
Liao Heng: I think architectural changes came with some fairly clear, important realizations. First, during the four or five years we were stalled, our algorithm peers were tirelessly working, and large models exploded, achieving substantial results. Those results all gave us a lot of design reference. Second, a very important event: some models, though starting from LLaMA’s open-source path—LLaMA’s first generation essentially “opened up the body” for everyone to analyze—really let people perceive what a large model actually looks like inside. Every micro-level detail gave us really good design references.
Of course, after that, a series of Chinese models—whether DeepSeek, or Qwen—kept releasing open-source models, and in terms of capability were near the first tier as well, giving us good reference points. Beyond that, there was also a lot of open-source work in non-language domains—for example, Chinese video generation, multimodal work—including Alibaba’s Wan series, and other internet companies open-sourcing a lot of valuable work in those domains too.
This work, plus the real-world performance feedback we got from things like our own autonomous driving system and phone business closed-loop, continuously improved the models themselves. So simply put, by the time we got to the 950, we had much richer design reference and evaluation criteria—right, a lot more of them.
Second: after DeepSeek’s R1 release last Chinese New Year, it directly triggered an explosion in model inference demand that we hadn’t anticipated—at a scale that turned out to be quite substantial. So we developed a much more realistic understanding of the capability required for inference—because inference doesn’t just need functionality, you also need fast inference, low latency, and under low latency, you also need each card to produce high throughput—a high token count per card. Balancing these factors to an extreme degree gave us a real sense of reality—because before we deployed large-scale inference systems at scale, we didn’t know these requirements existed. Once we quickly started serving inference customers, actually deploying inference systems, we ran into all sorts of engineering problems. One of the most obvious ones we discovered was—within a month or two of moving into inference deployment—we found that what we thought was important, compute, actually mattered less than we expected in inference—especially for the decode portion—while memory bandwidth turned out to be much more critical.
If you look at the ratio between compute and memory bandwidth—using the kind of proportional-scale analogy I mentioned earlier—you’ll find there’s a significant difference between the ratio needed for inference decode versus the ratio needed for training/prefill. This process also validated one thing—our belief that a “super node,” tightly coupling a large number of chips into extremely tightly synchronized collaborative work, was a key reason our 910C, or 950, was able to survive. Because we had a somewhat fuzzy vision earlier about wanting to combine more chips together, but it was through actual business deployment that this got validated—and of course, we also discovered where we hadn’t done well.
That is: when doing communication, say moving a chunk of data of 1MB versus moving 7K/8K worth of data—the granularity is different, and that’s actually a very fundamental issue. We found our previous granularity had been too coarse. What do I mean by granularity? Imagine filling a bucket, or a wooden crate, with pebbles versus filling it with sand—there’s a lot of empty space between pebbles because the granularity is coarse. If you fill the same box with sand, it packs much more tightly because each grain is much smaller. And if you then pour water into the same container, as long as the container doesn’t leak, it fills up even more completely, because water is a single molecule—even finer than sand.
So—we originally assumed this tightly-coupled network was meant to move large chunks of data. Later we discovered it’s actually used to move very small chunks of data—we thought we’d be moving 1MB at a time; it turned out to be 7K/8K. That’s a big difference, right? If designed poorly, you might have high efficiency moving 1MB of data but low efficiency moving 7K/8K—so all of this needs continuous refinement through real-world practice.
So that’s why I can say the 950 will be significantly better than the 910C—one reason is that the 910C is simply old now, and second, it was designed largely on blind guesses. By the time we designed the 950, we had a lot of real-world evidence to validate our guesses—where we guessed wrong, we could quickly refine and adjust. And products also started to diverge—as mentioned, the ratios needed for inference versus training differ, so different memory types may be required.
There’s also a more profound lesson from the past six months. This lesson: we actually recognized seven years ago that memory would be extremely important—that instead of selling another AI compute chip, we might as well be selling a high-bandwidth memory chip. That realization directly drove us to establish a dedicated department seven years ago to build high-bandwidth memory design capability. Even having spent seven years preparing, once the memory storm actually hit, we found our preparations still weren’t sufficient. This memory price surge—everyone’s seen prices go up more than tenfold—has hit the entire electronics industry with a huge shock. What I want to say is: even with this level of forewarning and preparation, when that memory tsunami actually landed, there were still a lot of regrets—why didn’t we do more back then? Why didn’t we push harder, prepare more?
Zhang Xiaojun: People probably always have some sense of complacency until they’re actually facing the tsunami head-on. For training chips versus inference chips, which do you think will have greater demand long term? What ratio?
Liao Heng: I think this has probably already happened—inference will definitely be greater. If inference isn’t greater, that would mean the economics aren’t working. Because inference is the process that generates revenue, while training is the investment that builds that capability—one is purely cost, a “cost center,” the other is a “profit center”—so the profit center inevitably needs to be bigger than the cost center.
Zhang Xiaojun: How does your thinking about inference chips show up in these generations of chips?
Liao Heng: I think this generation shows it quite clearly. For example, we’ve split into two tiers—one with very high bandwidth, and one with relatively lower but more economical bandwidth. For training, for instance, you could use the more economical version.
Zhang Xiaojun: Back in 2021, at the STW system technology symposium, you proposed that computing architecture has a new trend—moving from a heterogeneous architecture where CPU, GPU, and NPU are separate, toward a new kind of homogeneous computing architecture. Could you explain what heterogeneous means, what homogeneous means, and the difference between this “new homogeneity” and traditional homogeneity?
Liao Heng: I think, to some extent, that could be described as a bit of a marketing hook—used to counter some skeptical voices.
Zhang Xiaojun: Was that skepticism internal or external?
Liao Heng: Both internal and external. First, one of the most common debates in this industry is the SIMD-versus-SIMT argument. Some say GPUs are SIMT—”single instruction, multiple threading”—that’s the GPU architecture. Ascend, in the 910B/910C, is currently a SIMD architecture. What does that mean? When you’re processing a computation, do you express it as operating on a single element, or on a large batch of data all together? In other words, it’s the difference between a truck and a private car, right? If you drive to work by yourself, everyone’s driving their own car. If you take a bus, one bus carries a whole busload of people—dozens of people sharing one vehicle, right?
Why do I say this is somewhat of a “marketing hook”? Because actually, today’s SIMT processors are also internally SIMD. And today’s so-called SIMD processors also have multiple cores running essentially the same code—an SPMD-style pattern. But some people really like to argue this point, or really enjoy leaning on the CUDA ecosystem’s advantages, feeling SIMT is superior, because a large amount of original research work has been developed on top of the CUDA ecosystem. So if your processor looks exactly like a GPU, you can just directly download the code and run it without modification—that’s the first-mover advantage of the ecosystem.
But as I mentioned, SIMD itself—SIMT internally also has SIMD elements. For example, if you’re processing 8-bit floating-point numbers, but your register might be 32 bits—you have to pack four floating-point numbers together to compute them. That’s a small trick too—you’re never really operating in a “pure” fundamentalist way where you compute one number at a time—if you did, you’d waste the compute unit’s capability. So even in SIMT-style code, people pack four 8-bit numbers together to compute. It’s really just a question of whether the “bus” is wide or narrow. What can’t be denied is that a smaller bus means smaller granularity—for packing purposes, like the pebble-versus-sand analogy I gave earlier: sand can fill a box more completely than pebbles, right? So we’re just really talking about how large or small a data block’s scale should be. The difference is between 128-bit and 512-bit—actually, the GPU’s natural width is 128 bits. Our Ascend, especially early products, had a width of 512 bits—a bit too wide, honestly. We have to be truthful here—our pebbles were a bit too big.
Although, in the vast majority of, especially transformer-style, networks, the data is naturally large-block anyway, so having larger pebbles doesn’t hurt much. But in some networks, like recommendation networks, or certain early-stage, particularly edge-side vision networks, finer granularity may be needed. With the 950 generation, what we call “new homogeneity” means the processor supports both SIMD and SIMT, so this problem gets significantly alleviated. Because rather than getting caught up in this kind of “religious” debate, you might as well use different modes for different situations, since the essential issue—as I said, that granularity, how large the “bus” is—is now something you can dial in either direction. So by today, I’d say 90% of this debate can be set aside—technology has already evolved to a point where arguing about this is no longer the core, critical issue.
Instead, the real questions now are: how much memory bandwidth should be allocated, or how to build a super-node, how to achieve extremely tight fused computation (”mega co”) when doing computation, how to reduce synchronization overhead among multiple cores, even multiple chips. Those are the problems we consider basically solved now, but they’re the more important problems at this current moment.
Zhang Xiaojun: Do you have answers to these currently more important problems?
Liao Heng: We have some staged thinking on this—of course we have our own answers, or otherwise the product line wouldn’t be moving forward. We have the 960, 970, 980 all in our development pipeline. We keep developing more answers.
But let me tell you a more serious underlying reality: in the past, AI algorithm developers were used to using PyTorch—describing an algorithm using a bunch of “operators”—each operator is a function, a basic computation building block. Then you string these operators together to represent the whole model’s computation process. But the more important issue now is: if you want to develop a modern model this way, and achieve something like one-millisecond inference latency—meaning, when you’re conversing with a model, whether human or agent, submitting a prompt, and getting output in one millisecond—that means outputting 1,000 tokens per second. Today’s typical mainstream chatbots output maybe 20 or 30 tokens per second. Going from 20 to 1,000—that’s a 50x time compression. That’s a huge time compression—going from something like 20 milliseconds down to 1 millisecond—maybe a 20x compression. Right—that’s 20–30 tokens, or 40–50 milliseconds down to 1 millisecond, tens of times faster.
When you compress time like that, the more important problem is: for almost every use case that needs performance and low latency, you can no longer call individual operators one at a time. Because every call has launch overhead, latency, and data-transfer overhead. So the more important problem becomes: when a processor writes the entire model as one giant “operator,” what should that even look like? This is what we call a “mega kernel,” or kernel fusion—fusing dozens or even a hundred operations into one larger operation, to reduce launch and data-transfer overhead. This shift has been happening rapidly over the past year or two. The most famous example of this work is FlashAttention. And after FlashAttention, more important work followed—like DeepSeek’s various open-sourced kernels—all of them are exemplary models of fusion.
This fusion process actually has a huge impact on processor design. In the past, people designing a chip only cared about raw compute speed. Now, if you want to express extremely complex computational processes, you need programming to be relatively easy, and you need to be able to quickly modify programs on your processor to squeeze out maximum performance. This has given rise to some modern programming paradigms. As algorithms evolve rapidly, people can no longer afford to express such complex computation processes using low-level programming.
So naturally, new programming languages and compiler technologies have emerged—including Triton, an OpenAI project. Later, Peking University’s Professor Yang Zhi and his student Wang Lei built TritonLang, which is now also DeepSeek’s primary development approach—including our own effort, PiTorch, which is also trying to express extremely complex computation at a higher level—writing an entire model as a single operator. This is when compiler design and processor design both start to have a much deeper impact from the software side.
Because if you look at Hennessy and Patterson’s classic textbook, “Computer Architecture”—the first time I read this sentence, it deeply struck me—they wrote: “Computer architecture is the interface between software and hardware.” It’s an interface—a communication boundary between the two sides. Most people who did core processor design in the past probably didn’t have particularly deep awareness of this—because they only saw their own side, designing a compute unit, worrying about efficiency, about area. But once you elevate your understanding to this “interface” level, you realize that changes in programming languages are inseparable from this—this interface has two sides, like yin and yang—and where exactly the dividing line sits is determined not just by hardware, but also by programming methods and compilation methods.
So now, the design of AI processors, or Ascend processors—I think everyone designing processors is facing this reality: this interface is being reopened. Or put another way, CUDA’s ecosystem moat is rapidly dissolving, because it’s no longer the only main interface. If all of DeepSeek’s work is built on top of it, then it becomes an extremely important interface [in its own right].
Zhang Xiaojun: So this is also the significance of what you’re building now.
Liao Heng: Right—we ourselves recognized this quite a few years back, even though we faced a lot of skepticism at the time—we’ve also been doing open-source work on this PiTorch series of compiler projects. This work is purely built on first-principles thinking, not with the intent that our specific method or design philosophy should only apply to us—we believe it should be universally applicable—meaning, even though we’ve open-sourced this work, I believe all AI processors could benefit from the same compilation stack—that is, cross-platform.
Zhang Xiaojun: What’s your expectation for CANN?
Liao Heng: CANN is our software stack—it’s basically our whole runtime environment, compiler libraries, and models that we’ve already ported over.
Zhang Xiaojun: Do you expect it to become the next CUDA?
Liao Heng: Right now it’s—first, it’s fully working to protect the “legacy”—meaning it needs to have every capability CUDA has, layer by layer, one-to-one mapping. You can run PyTorch, so can we; you can run vLLM, so can we—though honestly, there’s still a gap there—that’s the legacy part. But the part that’s advancing rapidly, day by day—legacy mainly serves the general public. What does “general public” mean? If you’re a college student just starting to learn AI, deep neural networks, you’ll most likely use PyTorch, use torch operators—if the professor teaching your course uses Ascend hardware for coursework environments, you’ll need to use CANN. That’s the legacy—it’s long-tail, it’ll persist for a long time.
But serving the smaller group—that’s when I need to deploy something like GLM-5.2, and I want to maximize economic efficiency—highest token throughput per card per second, and lowest possible latency. That’s where the “cutting edge” part comes in—we need to use the most advanced methods, in the shortest time, to help people quickly achieve high-throughput, low-latency tuning. This tuning is extreme optimization—every additional 10% throughput is an additional 10% revenue, right? So this is real money. And this “cutting edge” part isn’t necessarily something ten thousand PhD students are using—maybe only 100 people, working around the clock, pushing hard to maximize the throughput of a production business system, because it is a production system.
We’re constantly working on both fronts. Ascend can only sell if progress happens on both. If the production system’s performance isn’t good enough, nobody’s going to buy it, right? Because if you spent that much CapEx and got out less token throughput, or your training doesn’t converge, or takes five times longer than someone else—nobody can accept that.
And what I mentioned about the ecosystem—that serves the broader community. This community, we’ve also been building through full open-sourcing, working with countless university professors, hoping they’ll co-build this ecosystem with us. Now, the CANN community has become one of the most active communities in Chinese open source—because so many people care about AI, and there’s a lot of enthusiasm here. Given this “many hands make the fire burn higher” dynamic, we’ve made substantial progress over the past year.
Zhang Xiaojun: Do you think achieving nanometer-level extreme physical pursuit is harder, or building an ecosystem that surpasses CUDA is harder? Which is more difficult?
Liao Heng: I think surpassing CUDA is harder. Even though pursuing the physical limits at the nanometer level, achieving physical breakthroughs, is difficult—why is CUDA harder? Because one relies primarily on your own effort, while the other requires changing the collective habits of a massive group of people. For example, imagine inventing a new instant-messaging app and trying to convince everyone to abandon WeChat and use your new app instead—that’s incredibly hard, right? Because this is a collective, group habit, and this collective habit needs to constantly receive positive reasons in order to make that kind of transition.
Developing the next more advanced chip—as I mentioned—is mostly about our own effort. As long as we work hard, maybe 70% of the factors are in our own hands. Maybe 30% are objective physical constraints, and we just have to find ways to break through those constraints. But these two problems fundamentally require different specialists to solve them.
Zhang Xiaojun: Was open-sourcing CANN a hard decision for you all?
Liao Heng: It wasn’t a hard decision, and here’s why—this is just my own personal view, everyone has their own perspective. First, I think human-to-human communication—even though the human brain is very sophisticated, the bandwidth between people connecting with each other is actually very poor. For example, right now, doing this interview with you, my rate of speaking out words is maybe three to five characters per second at most, if I’m going fast, right? So my communication bandwidth with you might be about 1 kbps. But if you’re an Ascend chip, and I’m an Ascend chip too, our communication bandwidth is 7.2 Tbps—terabytes per second—an insane number, right, some huge power of two. So the gap between these two numbers is enormous.
Why use this analogy to explain the importance of open source? Because when you’re trying to deliver an extremely complex technical system, you actually need the person using it to be able to quickly understand the system—what its interfaces are, how it’s designed, how it should be used. This communication bandwidth is actually an enormous bottleneck when you’re trying to convey this. And if on top of that you’re also limiting that bandwidth—controlling when a certain document can be shared with whom, requiring an NDA just to see CANN—that’s naturally adding unnecessary security gates on top of an already very constrained communication channel, right?
What we mean by open source is really extreme openness—meaning, once I show you all the code, I don’t actually need to communicate with you anymore, because you can just look through the countless files in the countless directories yourself, right? No need to waste more breath explaining. So it solves the single biggest bottleneck—sharing ideas, aligning understanding.
So after we open-sourced it, both on our development side and on the customer side, there was an immediate feeling of relief—because previously there was this really difficult problem, having to make a phone call every time to ask a question—now that artificial bottleneck has disappeared, right? The speed at which problems get solved improved dramatically.
Of course, on top of that, more people are now contributing their own work—and AI-assisted coding is also becoming an important, even dominant, way of working. Once open-sourced, those models—agents—can more quickly gain access to abundant training data to train on CANN programming. So those models also play a good role in CANN development.
Zhang Xiaojun: I think Ascend’s early story has a bit of a “forced into rising up” feel to it. Now that you’ve come this far, do you think Ascend is becoming more like Nvidia, or less like it? Have you two walked the same road, or different roads?
Liao Heng: I think in the early days there were more similarities—and to some extent, they’ve become more like us, right—like that “cube” I mentioned earlier—their Tensor Core has gotten bigger and bigger, more like our cube—becoming larger to achieve higher compute efficiency. But apart from these details, at the macro level, back then things were relatively similar—because everyone was primarily thinking single-chip, similar deployment approaches, and so on.
But at a bigger level, they’re increasingly not alike, I think—increasingly less like it. The most fundamental reason for the “not alike” is exactly what I mentioned before—the difference between small and large—granularity—the “small defeating large” approach, and super-node level system differences—increasingly not alike at the system level. At the micro level, there’s still a lot we could learn from them—like realizing our own thing had become too large and should shrink to be closer to their size—that’s better for algorithm-adaptation efficiency. But at the system level, it’s getting increasingly different.
Let me give you a small example to try to illustrate this. We can boldly guess, or infer, from the roadmap Nvidia has revealed at GTC, that Nvidia has been relentlessly pursuing extreme density at the rack level—and the rate of improvement is stunning—almost every generation doubling the compute crammed into the same rack footprint. Personally, I think there has to be a limit—you shouldn’t push rack density infinitely. Why? There are several dimensions here.
First, I don’t think a single chip’s spec should be pushed to the absolute extreme—or maybe it can’t be pushed any further anyway. Single chip, single package—why shouldn’t a single rack be pushed to its absolute physical limit? It needs to keep improving, but it shouldn’t exceed physical limits—because once you exceed physical limits, you run into catastrophic consequences.
Let me start with a single chip. Everyone’s seen the biggest crisis facing the AI industry right now—the HBM shortage—because demand has exploded, and everyone’s using larger and larger, higher and higher bandwidth HBM. Let me use a somewhat visual example to explain why 2.5D scaling has an upper limit. If we treat the compute die—although now it’s multiple dies, let’s simplify—as a square with edge length N, its area is N-squared, right? A square with edge length N has area N-squared—so its compute power scales with N-squared, because each compute unit needs a certain number of gates, and if you have N-squared gates, your compute scales proportionally with N-squared.
But when you have N-squared worth of compute, the memory bandwidth and interconnect bandwidth you need scale proportionally with it too. But if I put one giant compute die in the middle, and keep growing N larger and larger, while all my HBM, my I/O, my power delivery all sit around its perimeter—doesn’t every memory access bit have to cross one of the four edges of that square? And the perimeter of four edges is 4N. So you’ll find that as N grows larger, your compute, your bandwidth, your interconnect, your power all scale with N-squared, but your four edges only scale with 4N. So doesn’t the gap between a quadratic curve and a linear curve keep growing bigger and bigger? That contradiction just keeps exploding, right?
So this tells us that relying on ever-larger compute dies plus ever more HBM has an upper limit. Of course, before hitting that physical boundary, we’ll keep pushing toward it. But once you go past that boundary, you get catastrophic consequences—it’ll be an avalanche.
Zhang Xiaojun: Where might that boundary be?
Liao Heng: That boundary might be four reticles—each reticle is about 800 square millimeters—maybe six reticles, maybe eight reticles. I think that ultimately comes down to our entire upstream chain—what I mentioned earlier about the manufacturing layers—floors five, four, three—the basement of the basement—how far can they push that boundary? But I already know clearly that this N-squared versus N physical gap is irreconcilable. So this kind of architecture might work for one or two more generations, maybe three, but probably by the fourth generation it won’t hold up anymore. Nvidia is still on this architectural path—that’s the current target they’ve announced, at least.
So, I don’t think a single chip can keep doubling every single year. If it keeps doubling every year, this gap just keeps growing more and more irreconcilable, right? The second question is: why shouldn’t a single rack scale indefinitely, and why is it unnecessary?
Because if a rack’s power today is 100kW, next year it’s 200kW, the year after 400kW, and the year after that 1MW—reaching 1MW might only take four years—that’s already several times over—2 to the 4th power. Why shouldn’t we push it this high? Because as you pack more and more chips together, they need countless interconnects between chip and chip—and those interconnects also need bandwidth. The more you pack in, the more the interconnect load doubles too—not just compute doubling, but interconnect doubling, power doubling, heat dissipation doubling.
Let’s try to estimate: if it’s a 1MW rack using natural air cooling, how much heat-dissipation tower space would you need? For example, if you need 500 square meters to dissipate 1MW, then a gigawatt data center might have only 1,000 racks—but it might need one square kilometer of land just for cooling towers. Between these two—one rack’s footprint might be less than 2 square meters, but if it needs 500 square meters of cooling space, that’s a 250x expansion in space. When that ratio gets this badly out of proportion, you might need a water pipe stretching a kilometer just to reach the cooling tower, because that cooling tower footprint is far larger than the machine’s own footprint.
What I described just now goes way beyond processor architecture design itself—but it’s part of a different layer, isn’t it? If you’re building this “18-story pagoda”—if you’re building an AI gigawatt data center, someone has to build the building, lay the water pipes, design the cooling system—this ratio can’t be allowed to get this badly out of balance. We can’t really comment on how others are doing it. We can only say—in our own system design, maybe three years ago, we already recognized we didn’t want to cram too many chips into a single rack, meaning the physical distance between our chips would need to be considered.
If someone else crams 100 chips into one rack, and I need four racks to hold that same 100 chips, then the physical distance between those 100 chips would grow longer. Of course, a longer physical distance brings extra latency and overhead, but if you recognize this is fundamentally an unavoidable problem, and you know you’ll eventually need to spread things out, you need a low-cost way to connect them, so the transmission distance doesn’t become a signal-integrity failure.
So we’ve actively designed our first generation of UB—the “LingQu” (灵衢) bus, as we call it—to be able to run over both copper wire and fiber optic cable. You need to understand—the diameter of a fiber cable versus a copper cable is very different—maybe as much as a tenfold difference in diameter. When you’re connecting one cable, that difference doesn’t matter much—but if a rack needs 5,000 cables, the physical bulk of those 5,000 wires bundled together would be roughly as thick as an elephant’s leg. That’s honestly a bit brutal.
On one hand, they’re pushing signaling speed higher and higher—everyone’s trying to transmit more signal per wire. But on the other hand, they’ve had to use the world’s most premium board materials, which has directly caused a shortage in high-end PCB manufacturing capacity too. They’ve also employed some very advanced physical-layer technologies, right—all fantastic, extremely excellent work on their end. But as a systems designer, I don’t want to challenge the world’s toughest problem at every single layer simultaneously. Why? Because I need to reliably deliver a new, high-quality, high-reliability product to customers, on schedule, every generation. If I try to break through twenty physical limits at once, and even one of them fails, my product falls apart. I’d rather focus my energy on breaking through, say, five particularly hard, high-value problems, and leave some other hard problems for others to solve. That way, the probability of catastrophic failure drops dramatically, right? Because even if each individual success probability is 99%, 99% to the 20th power is close to zero. But 99% to the 5th power is something you can live with. So that’s a kind of design philosophy—play to your strengths, avoid your weaknesses—don’t try to compete on every single front with everyone else.
Zhang Xiaojun: So on the question of whether a single chip keeps getting bigger and bigger, your answer is no?
Liao Heng: We are growing it—but we want to grow it within our own “comfort zone,” not infinitely. We’re not going to push it to the point of self-destruction. Whether a single rack’s density should push to 1MW—we don’t need that, the answer is no, we explicitly know we don’t need it. Because our physical space is dirt cheap—every square meter of our data center might cost just 2,000 RMB. So compared to a 10-million-yuan rack, that 2,000 RMB is basically negligible. So we have plenty of space—through optical interconnect, we let these chips be comfortably spread out over a larger physical space, to build data centers with 100,000, even 500,000 cards—that’s all achievable for us. A single data hall today can already reach 100,000 chips without any problem.
Zhang Xiaojun: People in the large-model industry often say—”roughly a gap of a few months.” What do you think the real generational gap is in the chip industry?
Liao Heng: I think it’s natural for people to have this concern, this instinctive question. Earlier when I mentioned “Yao’s Law” being announced, part of the point was to let more people know that shrinking transistors isn’t the only path—when you can’t shrink further, you can stack. And this stacking can happen at the chip level, and also at the system level. I think by fall this year, when you see the newest phone releases—because there will be a lot of teardown analysis, reverse-engineering reports—you’ll be able to see whether this “Yao’s Law” stacking, folding logic can actually close, or shrink, this gap. My answer is: to a large extent, yes.
And I think people often overlook the most critical point: when you’re buying a GPU or NPU today, you’re actually buying memory. Because at today’s cost structure, 80% of an AI chip’s cost is HBM—only about 10% goes to the advanced logic die. So the primary competitiveness of an AI chip actually comes from how much bandwidth it has, not from how large its compute spec is—or rather, as compute specs grow, bandwidth must scale proportionally with it. After all, HBM can’t be sold independently—it has to be integrated inside some GPU or NPU, so its monetization depends on that final integrator.
Why am I going into this? Because the biggest core competitiveness doesn’t necessarily come from how advanced that single logic die is—more, it comes from how advanced the logic die is together with its matched memory.
Zhang Xiaojun: You just gave a lot of conclusions—when did you actually arrive at this understanding, through this whole exploration process?
Liao Heng: I’d say we’re actually somewhat lacking in this area. First is instinct—if you get lost in a deep forest, a good hunter has a kind of innate instinct about which direction might lead to water, which direction leads home. Maybe they read the stars in the sky and determine a direction, and choose to go that way. That’s what “instinct” means—making a choice without full evidence—instinct plus your own judgment. This judgment may not be a rigorous, mathematically provable conclusion.
Second, you need to test this direction, run experiments. And I’d say we’re still somewhat lacking here—we need to find effective dialogue with people working at the frontier of modeling, or of quantitative science, to see whether our instincts align with theirs. This is part of that “four-year head start” I mentioned earlier—you need enough credibility yourself so that others are willing to engage with you, to form effective communication. A lot of this needs an exchange of ideas at a thinking level, ongoing alignment. Much of that capability comes from the pioneers building the models.
Zhang Xiaojun: Can you talk about some stories of Ascend co-optimizing and adapting with Chinese large models over these past two years—like DeepSeek?
Liao Heng: I can’t get into specifics on that one, but let me say this: on the so-called AI ecosystem, there are roughly three tiers, each with increasing difficulty and demands. The first tier is: a model arrives, and I can do inference—that’s relatively simple, because once a model is adapted, it typically has a usage lifespan of maybe half a year to a year—once you’ve adapted it and tuned its performance well, people just keep redeploying it, replicating that work. This is a one-time task that can then be replicated.
Second is training—taking a model architecture a research team has already finalized and putting it into initial production-grade training—that workload might be three times that of adapting an inference system, because all the operators already validated for inference get reused for training too, but there are additional backward-pass operators, which tend to be more complex—though the extreme performance-efficiency demands are relatively lower than for a production inference system. Regardless, large-scale training is also somewhat one-off—you might do it once every three months, then move to the next version. This should still be a solvable challenge through sheer manpower.
The third tier is research—because a leading research team might have dozens of excellent algorithm researchers, each running different experiments every single day, constantly modifying their code. To really break through at that level, you need a much more advanced compiler stack—because researchers today have already shifted to using higher-level languages, since it’s easier to iterate, and it doesn’t matter whose processor they’re using—everyone faces the same problem there. If you tweak an algorithm and have to rewrite CUDA kernels each time, nobody can tolerate that. So we’re also working very hard, whichever company leads the compiler stack, to make sure our side handles this well too. Once that’s solid, we’ll have the conditions to expect that, within the next year, we could reach that “research-tier” level I described. But it does take time.
Zhang Xiaojun: As of today, is there original, genuine innovation in your work? Where does it show up?
Liao Heng: Maybe I can speak to a few dimensions. One is the chip itself—I’m personally a firm believer in not copying homework. This “not copying” naturally creates a lot of confusion, because not copying means you can’t borrow someone else’s ecosystem—you have to pay that price. But I have a deeply rooted belief that no tech company in the world has ever achieved great success by copying homework—because technology itself is about innovation, your main value-creation lever is building a differentiated advantage. Even if you have 100 flaws, you must have at least one strength—that’s my understanding, built over this long career. I’m particularly afraid of building products that look exactly like someone else’s.
Because if we’re not—like—Haining Leather City, or the Wenzhou small-commodities market, where everyone’s Christmas trees look exactly the same, and it just becomes a race to see who can be cheapest—that model might work in low-end industries, but it’s absolutely not viable in an industry that’s meant to lead the world. You have to have your own strengths, because only those strengths translate into product value.
Zhang Xiaojun: When did you decide not to copy homework?
Liao Heng: From day one. From day one, and every product I’ve ever worked on has never been a copy of someone else’s—the moment I hear “let’s make it like theirs,” I immediately feel that this won’t work.
Zhang Xiaojun: What’s the cost of not copying homework?
Liao Heng: The cost is—for example, the ecosystem gap starts off enormous. Everyone’s used to using Windows, and suddenly you want to make a non-Windows PC—you have to change people’s habits. Or, right now everyone uses iOS or Android, and suddenly you want to build HarmonyOS—you have to put in enormous effort to build a distinctive advantage, or give people a reason to switch. That’s the cost. And I think we’re now past that painful plateau, and hopefully entering a better phase.
Zhang Xiaojun: Isn’t copying easier than not copying, for you personally?
Liao Heng: For someone in my role, our mission isn’t about getting fewer complaints or a bit of praise from a customer meeting this year. Our mission is building a sustainable, ever-evolving product system. In that system—as I mentioned, quoting Hennessy and Patterson’s textbook—architecture is the interface between software and hardware—two sides of the same coin. Copying homework means: I want to make my hardware look as close as possible to someone else’s, and then just leverage their existing software—that way I don’t have to put in the effort to build my own ecosystem.
But I deeply believe this approach has several fatal problems. First, you can never assume you can transform yourself into something exactly identical to someone else—I could never truly become exactly like you—and there will always be consequences from not being exactly the same. Second, this kind of technical system becomes completely non-evolvable, because you’re following one step at a time, forever a step behind. You completely lose the possibility of independent design and innovation, because the other side’s software already exists, and you’re forced to shrink your own foot to fit into someone else’s crystal slipper. Once you’ve cut once, you’ll have to keep cutting forever. So I think it’s a “follower forever remains a follower” trap—the very act of copying already determines that you’ll always be behind.
And we don’t want to be behind—we want independence, self-reliance. And we have our own constraints—for example, even if I copy them exactly, their interconnect capability isn’t the same as mine—what do I do then? I’ll definitely fall short on the communication side. And if the communication falls short, then the fused compute-and-communication operators become unusable. So this problem isn’t something you can solve by copying just part of it—you’d have to copy everything.
Zhang Xiaojun: We just talked about innovation—you mentioned chip architecture—what about other aspects?
Liao Heng: The important one I already mentioned is the system level. We took a completely different design path—he wants dense, I want sparse; he wants tight, I want loose—I want everyone to have their own bedroom to sleep comfortably in, while he wants 20 people sleeping on one heated brick bed.
Zhang Xiaojun: For the generations that haven’t officially launched yet—including 950DT, 960, 970, 980, 990—what’s worth looking forward to?
Liao Heng: I think the next three generations, as far as we can see, will probably look fairly similar to our competitors—roughly doubling each year, that kind of cadence. At the system level, I think the most anticipated question is one we’ve been constantly thinking about: should we build a highly integrated “brick”—a monolithic building block—or should we build smaller “bricks” that different users, in different scenarios, can more flexibly combine through interconnects? This is a fairly interesting question at the system level.
Look at the equipment historically used in data centers—first it was servers, each in a 2U enclosure. Then AI servers grew to 6U, 8U—bigger boxes. Then “pods” emerged—basically an all-in-one form factor, where the whole rack is a single unit, with compute boards, switch boards, and a backplane on the front and back sides. Both forms have their own advantages. The all-in-one design has high integration density—you can achieve extreme density—for example, Nvidia’s NVL72 is exactly this kind of all-in-one big machine, right, and their roadmap has even bigger machines coming. This is an important form factor—we have similar designs too. But I think the advantage of this form is high density, and being able to pre-integrate everything—but the challenge it brings is: if the ratio of components in that machine needs adjusting later, you simply can’t—because the machine has already fixed every insertable component and their relative proportions, and their interconnections.
What we’re seeing instead is: because the business itself, the model itself, went from six-hundred-something billion parameters to 1.6 trillion, and might reach 5–10 trillion next year—the resources it needs, its KV cache which used to sit in DRAM/CPU memory, is now being moved to SSDs—and model inference speed, or latency, is going from 20 milliseconds down to 1 millisecond—these are all order-of-magnitude, drastic changes, whether in size, speed, or latency requirements—roughly half an order of magnitude to a full order of magnitude of dramatic shift.
These changes are very likely to be unpredictable at the time you’re designing the “pod”—two years later when the machine ships, that original ratio may have shifted dramatically—what do you do then? So we’re going to think more actively about needing a more flexible, reconfigurable “building block” approach. These blocks need to be combinable in the simplest possible way. This actually ties back to what I mentioned about the relatively “loose” design philosophy, and using all-optical interconnects—these ideas are mutually reinforcing.
In other words, once you go all-optical for interconnect, you can build smaller boxes—and these boxes can be freely recombined into optimal configurations just by connecting a few fiber cables. This is a kind of system-level innovation, or something worth looking forward to. And our super-node domain will also scale up significantly—because we found that for something like a 20-trillion-parameter model needing ultra-low latency, the communication domain needs to get much larger—hundreds, even thousands, of nodes.
Zhang Xiaojun: In the chip domain, has China charted its own path?
Liao Heng: As of right now, honestly, I think that phrase doesn’t really matter that much. I think you can look at these companies’ financial reports—those numbers actually reflect a kind of aggregate picture, because they’re mostly fabless companies—everyone sends their chip designs out to various fabs for manufacturing—so it roughly represents the industry’s total sum. This kind of data speaks volumes, because it represents, first, economic scale, and second, whether the business can actually turn a positive profit. Because whether we’ve “charted our own path” isn’t just about technical breakthroughs—it also needs to be economically closed-loop. You can’t rely on subsidies forever, or run losses forever—a long-term drain isn’t sustainable, right? So I think if you look at those numbers, you’ll get a pretty good answer. And this isn’t just about the logic-chip makers—you also need to look at memory makers, at panel/LCD manufacturers, at packaging companies.
My direct, or indirect, sense is that—especially over a five-year time horizon—they’ve all shown roughly order-of-magnitude growth. The next question is whether this curve is sustainable, whether it might fall back down. Personally, I don’t think it will—because, first, in many domains we’re actually not behind at all—advanced packaging, for instance, that’s not a weak area for us; DRAM’s gap isn’t that large either, HBM’s gap isn’t that large either—there’s some gap, but as long as logic can achieve a positive closed loop, that’s already quite good.
And of course, there’s also the macro question of whether the US and China will return to that old “strategic partner” style of mutual trust and mutual willingness to depend on each other. I think it’s pretty clear that more of the friction is coming from the other side, right?
Zhang Xiaojun: They’re unwilling—everyone says there’s a compute shortage—every model company is desperately short on compute. So what’s the real state of China’s compute today?
Liao Heng: Honestly, I don’t know the exact state, but I can make an inference from some indirect data points—and it’s probably a pretty inaccurate guess. Supposedly, last year China deployed something like 1-point-something gigawatts of data centers, while the US deployed around 7 to 8 gigawatts. So we’re at roughly a fifth to a seventh, an eighth, of that scale—a relative ratio. And of course, in a lot of other places, due to those special factors, they’ve also deployed quite a lot.
Zhang Xiaojun: Like Southeast Asia—can your own supply capacity be scaled up?
Liao Heng: We’re working hard on that—maybe by, say, this time next year—how much is it now?—this isn’t something I can say precisely, but I can say we should be able to resolve this fairly quickly.
Zhang Xiaojun: When did you personally realize you’d made it past the hardest days? Which year?
Liao Heng: I think when I started receiving more and more criticism—especially internal criticism—that’s when I felt we’d made it through the toughest days. Why? Because when things are truly at their worst, no one bothers to criticize you.
Zhang Xiaojun: Was that last year, the year before?
Liao Heng: Roughly last year or the year before, that’s when we started getting challenged internally more. Here’s why I use this as a marker: if you’re walking through a desert and haven’t had water in seven days, and you see any water, no matter how dirty, you’ll feel it’s saving your life, and you won’t be picky about how it tastes. After that first bottle of water, everyone recovers. But by the second bottle, people start noticing the taste isn’t great. That’s what I mean by “criticism”—so I think the hardest time is actually when nobody criticizes you at all.
Zhang Xiaojun: Does being criticized feel unfair?
Liao Heng: For me, personally—maybe my character isn’t cultivated enough, or it’s a flaw—I still feel some resistance in the moment. But afterward I usually come around to thinking, this is just an inevitable part of life, so there’s nothing really to feel wronged about.
Zhang Xiaojun: You described that decade-plus at the American semiconductor company as a long period of decline—a shrinking, sunset period. What did that decade feel like?
Liao Heng: For China, or for the environment I’m now in, this decade—despite going through this near-death ordeal—has absolutely been a period of explosive capability growth.
Zhang Xiaojun: What did it feel like?
Liao Heng: I’d say, apart from maybe being a tiny bit behind in process node, a tiny bit behind in specs, the vast majority of chips America can make, China can now also make itself—design through manufacturing. And I don’t just mean our own team—I mean the entire ecosystem outside Huawei too. You’ll find countless companies building autonomous driving chips, or embodied-AI chips, or WiFi chips—this capability has become widespread. The ability to design fairly complex, non-trivial SoCs has become something that’s blossoming everywhere—meaning this capability is no longer particularly scarce.
Second, an interesting point: the average age of people in this industry [in China] is about 20 years younger than in Silicon Valley. Twenty years younger than Silicon Valley’s semiconductor workforce—yes. I was considered relatively young—not old—in Silicon Valley’s semiconductor scene. But in China’s industry, I’m at the age where I might soon be considered “due for retirement.” So this represents hope—because young people have more explosive energy, and they have more runway ahead in their careers.
Zhang Xiaojun: Here’s a question—you see so many people now entering the chip industry, including large-model companies moving into it—how do you view this domestic competition?
Liao Heng: Honestly, I’m not particularly worried—I think competition is quite healthy. I hope they all end up strong, but I’m not afraid of whether they’ll hit near-death moments in that process.
Zhang Xiaojun: What was the most desperate moment?
Liao Heng: I didn’t personally experience one, so I don’t want to speak for others, project myself into different roles to describe that moment. Though maybe the closest thing I was indirectly involved in was the Kirin chip—the phone SoC. But fortunately, there were some very determined people who absolutely refused to give up, and they brought it back.
Zhang Xiaojun: Most of my interviews are with companies more on the software side. In a hardware-heavy field like yours, is the organizational and cultural design very different from those more software-oriented companies? Do you have any unique culture?
Liao Heng: This topic runs a bit deep. I’m not really someone responsible for managing an enormous organization or HR, so anything I say about organizational-level talent management might not be entirely representative. But I can speak from a personal perspective about what makes a person particularly valuable on a chip team, how they should grow—maybe that’s more meaningful.
First: to be an engineer or architect building chips, you must have actually built chips before. A brilliant PhD graduate can’t do it alone—you have to go through module-level design, then subsystem design—a long process, usually at least 18 to 20 months—and you must go through it more than once, several times. Why? Because chip design is very different from writing software. When writing software, if you write ten minutes’ worth of code, you’ll habitually compile it and run it right away. With a chip, you might only get to “press the button” once every 18 months—so you basically have zero room for trial and error. Press that button, and you’ve spent at least two or three hundred million RMB.
Second, chips involve a huge amount of physics. A smart person might be strong in the digital, logical world—but you have to go through the physical gap too. If a single nanosecond, a single picosecond of timing isn’t met, the chip immediately fails—dead on arrival, right? So it has an extremely rigorous, unforgiving engineering side, and school education simply cannot give you that. You have to go through this process personally to deeply understand why this level of exactness matters—not a single picosecond can be off, not a single wire can be miswired, not a single logic gate can be wrong—no critical bugs allowed. Whereas with software, you can fix it a minute later, right? You find a bug, it fails, you fix it.
So I’m actually not that worried about teams that look “glamorous” on paper, because maybe their leaders or organizers haven’t realized this—they overvalue how smart someone is, how good their resume looks, whether they went to a prestigious school. That has nothing to do with it—you have to go through the process. And this process comes at a cost—you have to go through it, tape out roughly 380-some chips over a year, or over five years, and that process forges a large number of people. Anyone unwilling to go through this process will never grow into the role of architect—that’s one point.
The second point: if fabless chip design sits at, say, the seventh floor, do you have understanding of the sixth floor’s capabilities? The eighth, ninth floor? Can you build shared understanding with someone working on algorithms? I think if you don’t have that basic understanding, people won’t even bother engaging with you—if you can’t even understand what they’re saying, the conversation simply can’t continue. I think this means an engineer needs to keep their field of view, their “curiosity funnel,” wide open—not just caring about their own immediate work, but also caring about what’s happening upstairs, what’s happening downstairs.
Over time, and it’s not just about listening to others—you also need to go and Google things yourself, or now, ask large models a lot of questions. For example, I once asked: how much cooling-tower area is needed to dissipate 1MW of heat? That’s not a question good enough to ask just one model—I’d ask four different models and compare whether their answers roughly agree in order of magnitude, and only then would I trust the answer. You see, these aren’t questions someone focused purely on their own narrow responsibilities would even think to ask—but they’re meaningful questions, aren’t they?
So what I mean by “keeping the funnel wide open” means you have to invest a lot of extra curiosity, and you have to read a lot of papers that seem to have nothing to do with your job. And when you’re able to hold a conversation with the frontier researchers working on the most cutting-edge models, you’ll find yourself entering a whole new comfort zone—not only can you understand them, you might even be able to predict what they’re going to think about next—you might even anticipate problems they haven’t thought of yet—because you have a more foundational, lower-level perspective, right? Over time, this “roaming across floors” experience accumulates, and things start to converge—and that’s how your ability to grasp technical problems keeps getting stronger.
Zhang Xiaojun: So when you’re evaluating someone, if you’re not looking at their resume—say someone who hasn’t actually built a chip yet—what do you look for?
Liao Heng: What I look for first is their values—maybe “values” is too strong a word, but it’s about whether we can get along, whether when discussing technical questions, or expectations for the future, or expectations for their career, our general outlook is roughly aligned. Because even someone who looks great by every other measure—if they only spend half a year, or a year, without the patience to go through those training cycles I mentioned, and they leave—for me, that’s a complete waste of time. So we increasingly look for a general alignment in outlook—otherwise we’d just be wasting each other’s lives.
Zhang Xiaojun: You wouldn’t outbid people with sky-high salaries, would you?
Liao Heng: Ha, outbidding with sky-high salaries—we don’t really have that capability, and honestly, it doesn’t quite fit Huawei’s culture of the “striver” (奋斗者) ethos.
Zhang Xiaojun: Is Huawei’s “striver” culture a kind of military-style management? How does the organization create innovation?
Liao Heng: No, Huawei isn’t—I think innovation here really comes down to, for example, an employee who’s only been here a year being able to walk into my office and argue this point with me—and I genuinely welcome that kind of colleague. I don’t know if I can generalize about the whole organizational culture—I can only say what I see in my own small corner of it—we try hard to make sure the people around us are self-driven, since self-driving is far more effective than being driven by someone else.
Zhang Xiaojun: Some young people today are wrestling with whether to stay abroad or come develop their careers in China—a common concern is that domestic resources are more limited.
Liao Heng: Which resources are you referring to?
Zhang Xiaojun: Compute resources—the gap is supposedly quite large.
Liao Heng: I don’t think this is quite right. First, within academia, domestic compute resources are actually far more abundant than overseas—that’s surprisingly true. Why?
Zhang Xiaojun: Really?
Liao Heng: Absolutely. Look at national labs in Beijing and Shanghai—places with 10,000-plus GPU clusters—that’s a staggering number. Overseas, even the most prestigious universities—a whole department might only have 1,000 GPUs, or a few hundred. So academically, China’s compute resources are absolutely leading—the abundance is almost unimaginable. This is partly thanks to certain senior leaders recognizing much earlier that providing compute is essential to doing good academic research. I think this condition is unmatched anywhere else in the world. So at the academic level, China’s compute is very abundant.
In the corporate world, I don’t think there’s excessive scarcity either—as long as a team is genuinely valuable, they can typically get some compute allocation. Whether compute is so abundant that people can just waste it freely—no, definitely not.
Zhang Xiaojun: Is compute a bottleneck for large models today?
Liao Heng: I don’t think it’s the main bottleneck—not one of the main bottlenecks. I can only say some very excellent teams achieved extremely high breakthroughs using very little compute. Like DeepSeek’s team, for instance—their compute usage was quite small—definitely not on the scale of hundreds of thousands or millions of cards—more like thousands of cards. Maybe Kunpeng too—everyone’s paid so much attention to AI compute, but general-purpose compute is also part of the infrastructure—along with the corresponding network, optical modules, network cards, SSDs—we continue that “eight core components” concept.
First, to build a complete data center, the most fundamental components are CPU and NPU/GPU. Beyond that, there’s memory, then SSDs or related storage systems for storing data, then NICs, switches. And switches also carry a relatively high engineering bar—the more ports a switch can fan out, the stronger its interconnect capability, and this directly affects how many network layers you need. For example, if a switch can connect 512 devices, one layer suffices. But if you need to connect 1,024 devices, one layer won’t cut it—you need two layers—and once you have multiple layers, you get latency, along with additional overhead between the layers, all of which multiplies.
Then there’s the NIC. So when we’re building this whole system, especially network technology, on one hand it needs to be advanced, and on the other, it needs to be interoperable—interoperability means being backward-compatible, because any system has a huge amount of legacy equipment—you can’t just build a whole new world from scratch. So when we’re building this—on one hand, we’re building the LingQu system. And our full Ethernet stack has also been developing quite remarkably over many years—from high-performance NICs, to RoCE, to switching gear—we’re trying to be at least first-tier in every single track, to make sure there’s no generational gap. For example, if someone else is using 51.2T switching capacity and we’re still on 25.6T—that’s a generational gap. That factor really matters, because if we only nail a single component, it’s very hard to build a complete 100,000-card or 200,000-card system—you’d inevitably be constrained by legacy elements dragging down your competitiveness.
This is also part of why we were able to go build our own entirely new interconnect bus—because the moment you’re building interconnect, you inevitably need to connect every necessary component together. On one hand, we have a complete Ethernet stack, ensuring interoperability with all legacy equipment, and that interoperability isn’t weak—it’s world-class. But when we’re building this “super node” technology, again, it’s part of that “eight core components” list—missing even one creates a major weakness—you can’t skip any of them.
This partly explains why so many companies in the communications field emphasize that it’s a standard—you have to go through IEEE/IETF, get certified for market access. Even for something like PCIe, you have to go through compliance labs to test interoperability, because it’s a multi-vendor world—every component might come from a different vendor, and they all need to work together.
I think especially around 2019, when we started designing the LingQu system, on one hand we were forced into it. But even then, we already understood that a technical system has to be self-sufficient—and the precondition for self-sufficiency is that you need to collect the full set of components required to make those connections. If even one component is missing, you don’t have the conditions to build the system, because that connection simply can’t be made.
So with these two factors combined, we happened to already have full self-sufficient capability across every key component, and had already achieved world-class quality—able to stand shoulder-to-shoulder with, or substitute for, the best products in the world. At that point, we had the conditions to define our own private [standard]. Now, we’ve actually made the LingQu protocol license-free and published its spec publicly—we welcome any manufacturer in the world to use it. But from our own perspective, we now have the capability to build a fully complete system—so we took this fairly bold step.
Zhang Xiaojun: What are you still lacking?
Liao Heng: I think our biggest gap is still on the software side. Hardware, sure, has its spec shortcomings—we hope to grow ours a bit bigger, right—every year for the next few years, hopefully doubling with real effort. We hope memory improves too, so it doesn’t hold us back. But the biggest gap is really this so-called software ecosystem. We hope that as our deployment volume keeps growing, more and more people will have both the conditions and the motivation to join in contributing on this new hardware platform, continuing to work on performance tuning—because whenever a manufacturer or customer deploys this system, some team on their end will need to migrate their business onto it. I think there’s a kind of tipping point here—once this community of people doubles from where it is now, that collective, community-driven force will become strong enough for the system to sustain itself in a positive, self-reinforcing cycle. So that’s what we’re most looking forward to.
Zhang Xiaojun: Earlier we talked about the ups and downs of the chip industry—you lived through its boom period, then the sunset period, and now with this AI wave, it’s entered another cycle of prosperity. Right now the application layer built on top of chips is still nascent, monopoly hasn’t formed yet. If it eventually does form a monopoly like those giant American tech companies once did, will the chip industry face another decline? How do you see the future of this industry?
Liao Heng: I think that’s possible—that’s a real possibility. But there are some factors continuously puncturing this assumption—if we think of monopoly as a balloon, some factors keep pricking that balloon. From an investor’s perspective—if I’m an investor and my entire retirement fund is invested in OpenAI stock, or Anthropic stock, I’d naturally want to maximize my return—right, becoming the world’s sole provider of AGI models would be most beneficial for me. From an investor’s perspective, if I don’t think one company is enough, maybe I bet on two—buy both companies’ stock simultaneously.
Why do I think some of these factors have kept monopoly from forming yet? One important reason is that AI technology hasn’t yet reached its saturation point—the plateau of its progress curve. So whoever’s leading today, maybe in six months, three months, someone else surpasses them—that’s entirely possible. This kind of alternating leadership has happened repeatedly over the past few years. We once thought LLaMA was the best model in the world, but clearly that’s temporary—maybe tomorrow it makes a comeback, right? Because—and I have to say this—human intelligence doesn’t have that high a threshold. If you walk around Wudaokou, you’ll find maybe 5,000, even 10,000 students who fully understand the latest model tricks, and they’re running experiments daily at smaller scale, hunting for the next breakthrough. So when something hasn’t reached saturation yet, it’s very hard to monopolize—that’s a purely technical reason.
The second important factor: I think there are also some people with real vision behind the scenes who, even if they had the ability, don’t want to become the world’s “final arbiter.” Instead, they want to make things more universally accessible, so more people can join this industry and make the next invention. Honestly, at least among the one or two leading Chinese teams I’ve interacted with, their entire value vision is like that—they don’t necessarily want to leverage their intellectual lead to maximize short-term profit; instead they hope the smart people around the world can keep racing forward together, and realize the next capability breakthrough. So I think there are two coexisting value systems here—their visions aren’t singular.
And of course, there are other factors—I believe in China’s particular environment, many people, even the country itself, absolutely do not want to see all of AI/AGI monopolized by a single American company—that could even create a crisis for humanity, right? So vision, values, internal drive, plus China’s galaxy-of-talent situation and abundant talent supply—generation after generation of outstanding students keep emerging—sometimes even the most important inventions come from interns—this kind of possibility keeps existing.
So I think, at least in the near term, I remain hopeful—I’m firmly one of the people hoping to puncture that bubble.
There’s also another point worth anticipating: AI right now, in the digital, virtual world, is clearly advancing very rapidly—I now deeply rely on it in almost every aspect of my work, and honestly, for almost anything specific I ask it to do, it does better than I could. That’s basically close to AGI in the digital world—whatever your definition of AGI is, it’s already extremely useful, extremely capable, right? But crossing over into the physical world—there’s still a considerable gap. From our own seven years doing autonomous driving, I already gave you one reason why “legged” robots might struggle to monetize soon—the energy-consumption reason, right? Maybe trailing a power cord solves that problem, but trailing a cord would limit its working range.
But the bigger gap is that physical AI models haven’t yet had their “ChatGPT moment”—this gap will need—maybe it’s coming, maybe within the next two or three years—a major model breakthrough to cross that zero-to-one moment for physical AI. And once that happens, it will spawn an entire industry of considerable, exciting scale, because once something’s physical, it needs a body—it has to be a machine, right? And that machine itself will be diverse—so this naturally leads to the conclusion that once you enter the physical world, monopoly becomes much harder. Look at the automotive industry—it’s been developing for over a hundred years, and there are still so many different brands across different countries and regions, even new brands still being born. And even people’s differing body sizes might require two different car sizes, right? So I think diversity is baked in there. And I think what’s most worth anticipating here is that I’m very bullish on China.
I once heard that Buffett may have said something like “nobody wins betting against America”—meaning if you short the US, you’ll lose. We’re actually not trying to short America—I just want to “long China,” because I think China, in every dimension, holds a lot of promise—it should be able to climb a very big staircase. And including the interview you mentioned earlier—I think a lot of these interviews frame things through a life-or-death lens—as if China winning means America losing, or that China’s superintelligence, or comparable AGI-level capability, would inevitably be used as a weapon to attack American networks. I think that’s a truly ridiculous framing—it’s basically a “devil’s advocate” worldview. Because if you had such a great model, why would you use it to attack you? Why not use it to improve my own life instead? Why not improve our own economy, our own healthcare—improve people’s livelihoods—rather than attacking America’s networks? I think that mindset itself is fundamentally flawed at the root.
Zhang Xiaojun: Is this a difference in values?
Liao Heng: I don’t know if it’s exactly a difference in values, or more a kind of “pirate culture”—they’re accustomed to viewing the world through that lens. China, even at its strongest moments in history, has never really needed to make others weak.
Zhang Xiaojun: Over the next visible five, even ten years, do you think the global chip industry landscape will change? Will there be some kind of reshuffling?
Liao Heng: First question: will things go back to how they were before the US-China tech/trade war? Because this has an extremely real, near-term impact—if this continues for another ten years, I think it will inevitably split into two separate camps—even without more severe conflict, everyone will still need, just to survive, their own complete manufacturing capability. This is because chips have become so essential to modern life—comparable almost to water and air—closing in on that level of necessity. In terms of economic scale, it’s already bigger than the oil industry. Imagine coming home to no rice cooker, no refrigerator—nothing with a power cord—unthinkable, right? So it’s a necessity. And that necessity means—if you ask whether there will be major changes—the first major macro factor is exactly what’s driven by these international dynamics. I’m not an expert in that field, so I can’t predict exactly what the impact will be, but there will definitely be impact.
Second, as you asked—will monopoly form, will we head into that decline period I described? My answer is: not that fast—maybe in ten years, it’s hard to say. But right now, because things change every single day, and the capability landscape hasn’t settled—nobody can say who the leading company is for sure. So maybe you’re leading by three months, six months, but there’s a good chance someone else catches up to a similar level soon. So right now, the conditions for monopoly simply don’t exist yet.
Zhang Xiaojun: Looking upward—at the top of this 18-story pagoda—what would you say to the large-model companies, the AI application companies? What are your expectations for them?
Liao Heng: First, I think the “best-resourced” party doesn’t always win. Like when we all go to the same college—you’ll find that the classmate from the wealthiest family isn’t necessarily the one who ends up far ahead of everyone. They might not do badly either, right, but maybe they end up just average. So I think the organization with the most money, the richest resources, doesn’t necessarily win.
Take America as an example—the top model company in the US isn’t Microsoft, isn’t Google, isn’t Meta, isn’t Amazon. Why? They have plenty of money, plenty of GPUs—countless cards—strong ability to invest, strong research teams, right, well-staffed. Because it’s a new thing, not an old thing. So an established company with scale has advantages built up in its own respective domain—otherwise it wouldn’t have become a giant, right? But precisely because AI is a new thing—its most valuable applications, its most valuable business models, are still being invented, or are still to be discovered.
I gave the AltaVista example earlier—the giant of that era was also a giant, and eventually sold itself off. AltaVista sold for a good price, but no one figured out how to monetize that great technology. When we were grad students, we thought it was fantastic—I told myself, I don’t have to go to the library for papers anymore. From the moment that tool existed, I basically stopped going to the library—maybe only to photocopy a specific paper I’d received, because back then downloadable PDFs weren’t really a thing yet.
So we have to watch what’s happening upstairs—the application layer—where, sure, there are countless giants present. Earlier I expressed my hope—that if giants can leverage their advantages, provide really excellent service, and keep improving—for example, I’d love it if the AI coding service I rely on were provided by a Chinese giant, so I wouldn’t have to go through all this trouble using something overseas, right? But I also deeply understand it might not be them—maybe they don’t want to be in that business. And of course I’d love it if the best coding model were from Huawei—maybe I could just call HuaweiCloud’s service directly. But every enterprise has its own current state, its own culture, its own already-established systems—it’s like an immune system: if a cat cell suddenly appears in your body, your immune system will clear it out. This is exactly why there’s that famous book, “The Innovator’s Dilemma”—describing exactly this issue: a mature organization typically can’t accommodate something genuinely new. But it depends on the organization’s own capacity for self-renewal, or whether its founder has that vision, right?
Second point: this vision really matters. If you say “I want to monetize this year, and rapidly achieve monopoly,” that’s one path. If you have a loftier vision—”I want to make sure that by next year, or three years from now, there are no more incurable diseases, no more unfixable bugs, no more cybersecurity problems”—that’s a completely different vision, and it will drive your team and organization in a completely different direction.
I think—you see, OpenAI created a chatbot, but Anthropic’s main revenue comes from coding, right—via API calls. At that point, these two businesses have basically nothing in common. Don’t you think? They’re selling to different customers, monetizing through completely different mechanisms—nothing similar at the business-model level. So I’d say it’s a new business—and even the question of whether that’s the “best” model—should you control the coding entry point, like controlling something like a VS Code IDE, or do you not need to at all, just sitting behind the scenes providing an API? You see—different people, even extremely accomplished figures, all end up in fierce competition based on some particular fantasy.
Think back to when the internet first took off—some of you might still remember Microsoft’s IE versus Netscape’s browser war—a competition that lasted years, and eventually the US Department of Justice intervened, right, and then Bill Gates retired. That’s a story now, for people today. But looking back, was that competition actually meaningful? Today, Windows doesn’t even ship IE—the default browser on Windows is a Chromium-based Edge—literally Google’s Chrome engine with a different shell wrapped around it. So the battlefield you once thought was the most important, the one you absolutely had to win—once you actually won it, you found it wasn’t worth much—it didn’t deliver the fantastic result you expected. Microsoft never became the internet’s ruler, right? Instead, a whole set of new giants emerged, each finding their own strange track—and nobody expected China to emerge out of nowhere with Alibaba, Taobao, Alipay, WeChat—it was all new. They didn’t compete within the tracks those companies had defined at all, right?
So I think the AI application layer might be exactly the same. And I think, within this application layer, in the digital world, one of the most important thresholds people need to cross is this: AI is already very capable, very smart—but what it’s missing is your personal context. If it can already do 95% of what I do, better than I can, the one thing it hasn’t replaced me on is that it hasn’t been embedded into my life context, my work context. So I become a machine feeding it prompts—even though I might have no problem-solving ability better than it does, I still have one thing it currently lacks: me. It didn’t sit in on your meeting with me, didn’t attend this morning’s meeting, doesn’t know what KPIs my manager set for me, doesn’t know what the most important project I’ve worked on for the past six months, discussed across hundreds of meetings, actually is. Once you cross that threshold, you absolutely need an “entry point.”
I personally expect this entry point to be the next super-app. And of course, this entry point immediately crosses into human boundary issues. For example, if I embed AI directly into all my social apps—like WeChat—letting it see every conversation’s context, it would then know essentially everything about all my non-work interactions—that inevitably crosses personal boundaries, right? But it also means it wouldn’t need me to laboriously describe everything I want—it could just go do it for me. Similarly, if it’s embedded into every tool I use at work, if it can see every email, sit in on every meeting with me, see all my work communications through something like Huawei’s WeLink—it would know my entire work context, an “always present” companion. That entirely could mean I get phased out—because next time, maybe I don’t even need to attend the meeting, since it does everything better than me.
But at that point, we cross into a different kind of boundary—one is that I’d lose a sense of security; the other is that my company would say, you’ve breached information security—how could you let a chatbot or agent know all this confidential meeting content? What if it leaks it somewhere inappropriate? You see, there’s still a huge amount of work to be done at the application layer. I’m just describing that what’s missing in digital AI is that “human-embedded-in-context” piece. Whoever solves that first—finds some novel design—that design might have nothing to do with the underlying model at all—it’s just a better app. That app could become the next super-app. Because if it knows everything you know, and it’s more capable than you, then what do you do?
Zhang Xiaojun: You mentioned the innovator’s dilemma earlier—but what you’re building at Huawei, chips—within Huawei’s own system, this was also something new.
Liao Heng: We’re not exactly newcomers—HiSilicon has existed for a long time, long before I joined. But its dilemma lies in: who gets to define the future? Because that “future” is really a vision—what should I build, how should I build it? In any organization, unless it’s your own—unless you’re the founder, unless you personally control every resource, right—anyone who isn’t in that position, as an employee, or as a contributor, has to convince others: “here’s the future, here’s how you’re going to do it.” But often, at that point—like I mentioned earlier, thinking about living in the year 2030—that kind of forward-looking decision can feel like pure fantasy to others. So this requires building a lot of consensus—this “dilemma,” so to speak, is unavoidable. We don’t over-dramatize it—this kind of collaboration is simply necessary within any human organization, any society.
Zhang Xiaojun: If today were the year 2030, what would you see?
Liao Heng: If I imagine living in 2030—I’d see that we have enough manufacturing capability. I’d see digital AGI has become even better. Maybe I’d see a robot that can help me sweep the floor, getting closer to reality, right—and then I’d start worrying about what I myself should be doing.
Zhang Xiaojun: Have you become more confident over these years?
Liao Heng: Professionally, more confident—because when you’re faced with ten times the difficulty, and eventually you find a way to solve it, you naturally gain more confidence facing the next challenge. And once you cross those thresholds, it no longer feels as hard.
There’s another topic worth adding here—optical communication. As I mentioned, super-nodes need to spread chips further apart, so they need optical interconnects. This is another case where we had to respect physical boundaries. We built something called HiOne—a 7.2T optical module—while the market today mostly sells 800G modules, meaning 0.8T. So immediately you see: 7.2 divided by 0.8—roughly 9x, nine times higher.
And immediately we faced a very interesting problem—once you realize a chip has this much bandwidth needing to go external, imagine a machine the size of this table—your chip sits on a board, and the optical module sits on the chip’s front-panel side, with a connector. So from the chip to that connector is a stretch of copper wire—this cable might be tens of centimeters—the longest maybe 50 cm. Picture it—like an octopus, with countless cables shooting out from the chip, all needing to be routed to a board with optical modules—that’s today’s mainstream industry pattern.
The 7.2T optical module we built is placed directly right next to the chip. So the connection between chip and optical module shrinks to maybe just 3 to 5 centimeters—because that chip’s size is roughly that scale—eliminating that whole stretch of copper cable. This is what the industry calls “NPO”—Near Package Optics.
But when we tried to solve this problem, we very firmly chose NPO—not putting the optics directly inside the package, not inside the chip itself, but right next to the chip. This immediately raises another issue—if you get a chance to interview others in the optical communications field, you’ll find this is a debate that’s gone on for years—should it go inside the package, or outside, or right next to it? But interestingly, for me, I didn’t even need to go through the debate—I directly chose “next to it,” and firmly decided in 2026 that this 7.2T optical module absolutely should be positioned right next to the chip—not outside, and not inside the chip package.
Why? Because there’s a second problem—as soon as you’re dealing with optics, you inevitably need a laser source. What’s that light source? That laser can fail—in relatively hot, dry working environments, it has a certain probability of failure. Once it fails, you need to figure out how to repair it—that’s the second problem. And a lot of people, even if they chose to put the optics inside the package or right next to the chip, firmly chose to put the light source outside—because if it’s outside and fails, you can just unplug it and plug in a new module to swap it, solving the failure problem that way. But we firmly chose the opposite—absolutely do not put it outside, put it inside. And not just inside—we built in two of them, so if one fails, there’s a second one that keeps working, reducing the failure rate.
At this point, you might start questioning—you’re someone with a software and algorithms background who later, sort of pretending, went into chips—what gives you the authority to make such a definitive judgment call, that this is the only right way to do it?
This is actually a fairly involved topic—and it perfectly aligns with the philosophy I described earlier about how a person grows within an industry. First, you have to keep that “curiosity funnel” wide open—look above you, below you, all around, 360 degrees. Around 2008 or 2009, I already realized that chip I/O—because the physical distance a signal can travel over copper is too short—would eventually have to switch to optical. At the time I thought this transition would happen at 56G signaling speed, but in reality it didn’t happen until 224G—four times later than I expected. But even back then, I already understood that eventually optics would have to move right up next to the chip, or inside it.
Carrying that question with me, I started trying to learn how lasers work, how signal modulation works, how signal reception works. And because at the time, optics and chip design were completely separate industries, the first year I did this, I signed up for a summer course—because at the time, the Netherlands, in Europe, had some advantage in optical communications, and a university there offered a summer program teaching optical-chip design. So I went for a week or two during the summer. And I came away with a foundational understanding of how optics work.
That knowledge actually became the basis for a lot of important choices I made later. I immediately understood that every optical component ultimately depends on one basic physical quantity—refractive index. Because for any material light passes through, refractive index represents the speed of light within that material. In a vacuum, light travels at 300,000 km/s, but in glass it’s a different speed, in silicon nitride it’s yet another speed. The second basic fact I learned: refractive index changes with temperature—every material is subject to this—a one-degree temperature increase changes the refractive index. Third, I learned that lasers do fail. Fourth, I learned that all optical connections require “coupling.” Coupling, for electrical wires, is like soldering—you take a wire, use a soldering iron, and that’s a discontinuous point, and the connection is made. But optical coupling—every time you couple, you lose a fraction of a decibel of energy, because passing from one material into another always loses some energy.
You see, all of this looks like fairly basic high-school physics—things a student who studied well would already know. But surprisingly, some people who’ve worked in this industry for decades have forgotten these fundamentals—they’ve learned a lot of other knowledge instead. But this knowledge is exactly what shaped my fundamental view of the chip. First, why not put it directly inside the package right away? Because the inside of the package is extremely hot—there’s a 1000-watt NPU in there. What does 1000 watts mean? Think of an electric heater—a fairly high-power household heater might be around 2000 watts. So it’s extremely hot in there—anything placed right next to that will heat up, and the refractive index will shift dramatically. So placing it inside creates a lot of problems from refractive-index drift, which is why we generally want to avoid that.
Second, higher temperature causes lasers to fail. Third—the other reason I mentioned for not putting the light source outside—if I have 72 channels of light, and I need one light source to feed 72 other sources, doesn’t that require 72x the energy? That 72x energy creates a hotspot—meaning an extremely high-power concentration hitting a single point, and it also has to be split into 72 paths—that single point will inevitably fail. Simple as that. Because that point carries 72 times the power, concentrated at a single location—it will burn out the material, or any glue or dust that gets on it will immediately char, and it stops being an optical communication component—it turns into a cutting laser instead.
So—maybe others made these decisions differently. What I’ve just described are all fairly intuition-based choices—I didn’t run experiments, didn’t validate through a hundred experts weighing in, right? I didn’t need validation—I directly eliminated the wrong option, because I knew it wasn’t good, so I didn’t do it. I picked the option that was easy to get right. And honestly, our 7.2T module today performs very well—it may well become our primary connection method in the next generation. It’s cheap, it has high bandwidth, it’s small, and it eliminates all that troublesome cabling. You see, this kind of philosophy is purely intuition-driven, vision-driven—it doesn’t need endless discussion with countless people—but it may end up determining whether our system can scale across a much larger physical space. These ideas sound very simple when you say them out loud, but someone has to be the one to raise the question in the first place.
Zhang Xiaojun: After all these years, do you still have strong passion for the chip industry?
Liao Heng: My passion comes more from a sense of need. It’s like this—if you’re walking down the street and a stranger suddenly collapses, you’d feel you should go check if something’s wrong with them, right? Right now, what we’re seeing is that this “need” is very easy to perceive. We need compute, we need self-driving cars, and so on—these needs are everywhere. We need our own smartphones—the whole country can’t only have iPhones, right? These needs are what people call “necessity is the mother of invention”—that’s mostly where it comes from. But of course there’s also a negative driving force—as long as a person isn’t doing something particularly good, there’s a good chance they’ll end up doing something bad instead.
Zhang Xiaojun: What’s your sense of mission?
Liao Heng: Right now—whether it’s personal mission, or company mission, or Huawei’s mission—it’s written very clearly: bring digital [connectivity] into every person’s life, every enterprise, every society. My personal mission is just—I feel that within this limited life, I want to leave behind, as much as possible, more good things, more constructive things than destructive ones. Because a person lives—you’re born with nothing, you leave with nothing—but throughout that process, I hope to remain useful, to be of service to others.
Zhang Xiaojun: You also gave me two more questions—where is “the city” and “the valley,” and what’s the difference between AI’s ideals and Wall Street’s ideals?
Liao Heng: I actually already touched on both of these earlier. I think—if you ask, where is AI’s “Silicon Valley”—my guess is it’s probably around Wudaokou, or maybe Shanghai’s Qiantan, or Houhai, or Binjiang—that area, the Xuhui District’s innovation institutes, that sort of place. Why do I say this? Because my own kid is currently at school around Wudaokou, still in third grade, and I often go there for exchanges and discussions. I think if you sit down at any random restaurant or coffee shop there, you’ll find the table next to you having an intense discussion about why yesterday’s training loss spiked, or how to improve reinforcement learning—because that’s the atmosphere, the crowd there. Why not Silicon Valley? Not because Silicon Valley has lost its ability to attract talent—it’s that these companies’ own desire for monopoly limits things. An Anthropic employee probably won’t sit down and casually discuss the latest reinforcement-learning techniques with a Stanford undergrad, because they want to protect that as their competitive edge—because a monopoly is currently forming there.
So I think China has a much more abundant talent supply right now, with far more people actively working on all kinds of related problems—and China is much further from monopoly right now. So in this brief window of time—maybe three years, maybe five—we’re right in the middle of this kind of galaxy-of-stars, hundred-flowers-blooming, extremely vibrant period, with extremely active exchange of ideas.
Zhang Xiaojun: So this open-source culture forming across China’s model companies and chip companies—regardless of which company—might have a far more profound impact than we can currently imagine?
Liao Heng: Yes, that’s how I see it—and I hope this open-source culture lasts a while longer, because there’s no guarantee that companies willing to open-source today will continue to do so next year. I hope it persists—if it’s the founder’s own decision, it’s more likely to last. But if it’s some department within a certain company, that department might not even have the authority to make that choice on their own, right?
Zhang Xiaojun: They might close it up faster then.
Liao Heng: Right—could go either way, may not stay open forever. I don’t know if it’ll close up. Anyway, more challenges may come. Let’s move to a quick round of rapid-fire questions. First—a food you like, from anywhere in the world.
Liao Heng: Food—I’ve eaten so much mixed-grain, health-food stuff lately, I honestly don’t know what I like anymore, or haven’t thought about it in a long time.
Zhang Xiaojun: A little-known but important piece of knowledge—though you already gave one earlier, so let’s skip that. Based on everything you’ve read, can you recommend a few books?
Liao Heng: I’d recommend a few books. First, one we’ve also recommended internally—”The Idea Factory”—I think there’s a Chinese edition available on JD.com. It’s about the history of Bell Labs. What resonates with us especially is the period from 1940 to 1945. At that time, America’s economy—its GDP—was already very developed, but its science and technology still lagged behind the old academic strongholds of Europe—England, Germany. Back then, top professors mostly had to have studied abroad—either at Cambridge, or Göttingen, or some German university—that’s what made you a “top professor.”
But during those years, 1940 to 1945, Bell Labs produced some fantastic work—Bell Labs, incidentally, is right next to Princeton—draw a circle around it, 100 kilometers, and you’ll find, in that 50-to-100-kilometer radius, work emerged with extremely far-reaching impact. Computing had von Neumann at Princeton, information theory had Shannon, and the transistor had Shockley and Bardeen—that’s not a coincidence, it’s the result of some larger historical convergence of factors bringing them together. California hadn’t yet risen as a tech hub—the real tech center at the time was genuinely right around Princeton, within that 100km radius.
This is exactly why I said earlier that when I was at Princeton, I didn’t learn all that much—it was only in my forties or fifties that I suddenly came to appreciate how much brilliance, how much human talent, had once gathered in that specific place—and I feel I missed something, a bit of regret there.
Why are these particular works so relevant? Because the first involves semiconductors—we’ve already talked a lot about that—China is right now going through an extremely difficult period, breaking through, being reborn from the cocoon. Second is computing—von Neumann’s work is right at the heart of that, and Turing himself was actually also a Princeton student. And the third is communications—China’s already not weak in that area.
So I think these elements, in a compressed five-year window, actually echo the macro factors of our own current decade. Because the Turing machine represents the most foundational theory of symbolism, while today’s AI represents connectionism—and connectionism is right now in a period of explosion. Second is semiconductors, right—we’ve discussed a lot about semiconductors—China is right in the middle of extreme difficulty, breaking through and being reborn. And third, communications—China isn’t weak there at all. Meanwhile China’s real economy is already extremely strong—arguably unmatched in the world. But our technology still needs to leap from being a “follower” to becoming, in certain pockets, even world-leading in original invention and creation. Maybe that will happen somewhere in Shanghai, maybe somewhere around Wudaokou, maybe some combination of the two.
I think this era—this moment—is quietly happening right now, just with roughly an 80-year gap in time and space from that earlier period. But that 80-year gap represents China’s moment, its opening era.
So I think—regardless of trade wars or other factors—I think all these elements combined—like having “all the right great ingredients” at exactly the right time, the right place—that combination will inevitably trigger a chemical reaction. So that’s one book worth recommending.
Of course, the other one I mentioned—”The Innovator’s Dilemma”—is maybe more suited for company leaders to read, right—to recognize that your past success won’t guarantee future success, or that the very thing that made you successful might become the biggest obstacle to your next success.
I also recommended another book to our colleagues, called “Rules of Work”—you can probably find it on JD too. I don’t know the Chinese title offhand. It’s actually a very simple little booklet, about how young people should appropriately understand their work environment—because most universities never teach you this. The most fundamental point is: unless you’re a purely individual contributor—maybe you’re a mathematician who can just lock yourself in a room and never need to interact with anyone—but the moment more than one person is involved, human-to-human interaction gets relatively complex.
So there’s a basic understanding needed—some rules for reaching consensus in a healthy way—neither escalating to the extreme of trying to “revolutionize” and tear down the organization, nor swallowing excessive personal grievance. I think this book—I’ve given it to a lot of colleagues—has helped bring them a sense of relief and comfort. So I think “Rules of Work” is worth a look, especially for younger listeners.
While preparing for this, I also thought of a biography of Edison. Edison is, first of all, a legendary figure—a hundred years ago, Edison was basically the Elon Musk, or Steve Jobs, of that era. I think these figures might share some kind of gift from heaven—perhaps repeated “resurrections”—who knows, maybe the same soul being reborn again and again. But I think they share some common qualities.
Why do I recommend this particular Edison biography? You can find it free on gutenberg.org—I couldn’t find a physical copy, it’s too old, too rare—it was written by someone who worked closely with him, so it has a strong sense of authenticity. I think Edison—especially people who did really well academically—often confuse him with Einstein, thinking they’re both cut from the same cloth of great historical minds in the eyes of engineering-minded people. Actually they’re two completely different types. Edison probably didn’t even finish middle school—came from a very poor family, maybe only completed elementary or early middle school before starting to work—selling newspapers on trains.
But he had this innate “engineer” quality—and that engineering quality boils down to a few basic principles. First: you must build something useful—never waste time on flashy things with no substance, always invent something with real utility. Second: you must respect reality—you first need to know what problems are actually worth solving, and not waste time on meaningless ones.
The book is full of examples of what he invented—and mostly, he used what you might call “brute-force” methods. He wasn’t some kind of super-genius—his “extraordinary” quality was really about identifying what humanity needed. No electric bulb—invent the light bulb. No film—invent a way to record moving images. Even the first movie ever might’ve been filmed by Edison—there are a lot of interesting stories in there. I believe engineering-minded people, or people who lean toward engineering, will find a lot of meaningful reference material in it.
Another book I read early in my career is called “The Soul of a New Machine.” It’s about—I mentioned this earlier—that same era of Digital Equipment Corporation, and there was another company called Data General, working on minicomputers at the same time. “The Soul of a New Machine” describes an engineering team, maybe in the late ‘70s, designing a brand-new generation of machine.
I think this book might be more directly relevant for those of us actually working in computing—it’s a very real, almost day-to-day account of every individual, how they participated in the project, what difficulties they ran into at the factory, how they solved things even when they were having a bad day. I actually had the chance to meet one of the interns described in that book—his name was Bob, I think—by the time I met him, he was already an EMC Fellow, a white-haired old man.
So I want to say—this book is about the story of engineers, especially the story of an engineering community, and this story just keeps repeating generation after generation. Maybe you’re building a minicomputer, maybe we’re building a chip, or an AI super-node, or the next model, or the next hopeful super-app, right—for me, it’s still meaningful, because you see that whatever you’re going through, the people before you already went through it too. And it’s especially meaningful because engineers, honestly, are a fundamentally boring bunch, right—mostly introverted, not great talkers, and definitely not the type to write their own memoirs. So this is a rare kind of book—describing the real, almost first-person experience of engineers. I think it’s worth reading for the engineering community.
Zhang Xiaojun: When you talk about engineers passing things down generation to generation—it makes me think, everyone now says software engineers are about to be replaced by AI coding. If that happens, maybe people can pivot toward hardware instead—though maybe that’s not quite the right takeaway either.
Liao Heng: Since we’re on the topic of coding, I think there are a few things to consider. First—I don’t think people need to worry too much about being replaced, or rather, if you’re someone with your own ideas, unwilling to be easily replaced, you’ll definitely come up with new needs, invent the next interesting thing. Whether for your own company, or for yourself, you’ll be able to define something you previously couldn’t do. Now, thanks to AI’s boost, you might suddenly have the development capacity of ten people, even a hundred people—so couldn’t you go build something entirely new that’s never existed before? Maybe you won’t earn more salary from it, but—I think what you just described, or what a lot of people worry about, is really more about whether you’ll end up lying flat, or whether you have the internal drive to keep creating more useful things, services, or products.
So what I said earlier—whether within our own small organizational scope, or more broadly—I hope everyone has that inner drive to keep pushing forward on their own, rather than needing someone to tell them what to do before they act.
Zhang Xiaojun: Based on your current understanding, what’s one key, important bet you’re making right now?
Liao Heng: Betting on China doesn’t mean betting against America.


