Inside China’s AI World: The DeepSeek Transcript
Four hours of Liang Wenfeng on Nvidia, Huawei and the compute gap.
Grüezi!
In late July 2026, the transcript of a fascinating investor call surfaced on Chinese media. The fascination was down to the man holding the investor call – Liang Wenfeng, the CEO of DeepSeek.
Liang was pitching, but in the course of a few hours he dropped a whole host of interesting views and insights on how he sees the current AI space. Most of them are worth looking at closely because – well – they tell you a lot about the most important geopolitical rivalry of the 21C.
1 Death Cab for CUDA
His biggest claim was about US chip giant Nvidia. And essentially, Liang says, Nvidia is headed for trouble. Why? Not because it doesn’t make great chips. It’s in trouble partly because one of the things that’s kept it going for so long is not silicon but a software platform called CUDA.
So why is CUDA one of Nvidia’s secret weapons? Nvidia made its name making graphics processing units, or GPUs, chips that have thousands of tiny processors performing millions of calculations at once. It made computer games run faster and look cooler – but it’s absolutely perfect for the kind of matrix multiplications that are at the heart of AI.
But raw silicon alone is useless without the software that tells those processors what to do. And that’s where CUDA comes in. Nvidia launched it back in 2007, and it’s a platform and programming model that lets developers write exactly those kinds of instructions.
A kernel or an operator is a little program that performs one specific operation. It could be a matrix multiplication or an attention calculation on a GPU. Over nearly two decades, Nvidia has built up an enormous bank of pre-optimised kernels and higher-level libraries. For deep learning it has cuDNN; for linear algebra cuBLAS; NCCL for communicating between GPUs; and for inference, TensorRT. Nvidia is so embedded in the architecture that every major AI framework – PyTorch, TensorFlow, JAX – was built on CUDA first.
So Nvidia’s moat has never been just its chips. It’s been this patiently constructed software ecosystem and the millions of hours of time that developers have sunk into it.
A rival chipmaker could try and match Nvidia’s transistors, but a customer that’s built up years of CUDA code faces a really high rewriting cost. AMD has struggled as a competitor despite having very respectable hardware, because the switching costs were all about the software, not so much about chips.
2 The ugly, the bad, and the good
Now, this is where it gets interesting. Liang says Nvidia’s moat is drying up, and in the call he puts up three reasons why he thinks that’s the case.
The first is simply that AI can now write code. So if you want to construct a CUDA-equivalent ecosystem, it’s much, much easier than it used to be.
How does that claim stack up? Well, it’s right in the way that coding is headed. It’s still a little overstated on where we are now.
Models are now writing kernels. But code only earns its keep if it improves on what it replaces, and the standard test for that is KernelBench. In early 2025 it found that frontier reasoning models working by themselves failed to clear that bar in more than 80 per cent of cases. Results from 2026 are a little better, but nearly half the AI-written kernels that do actually work correctly are still slower than the code they were meant to improve on. AI can’t yet go toe-to-toe with Nvidia’s 19 years of hand-tuning an ecosystem all by itself.
The second reason Liang gives for dunking on Nvidia is TileLang, which he describes as a technology that DeepSeek produced. That’s a stretch. TileLang was an open-source project published back in April 2025 by a team that spanned Peking University, Imperial College London and Microsoft Research.
Why does he think TileLang hurts Nvidia? If you write a fast kernel in CUDA you’re really writing two things at once. First there’s the dataflow – the maths you want done. Then there’s the schedule. You have to detail every fiddly hardware-specific decision about how to actually do it – how to slice the data so it fits in the cache, which threads get which work, and how the memory is laid out. Those two are tangled up together in the same code, and the second half is written for a very particular Nvidia chip. Which is why changing chip means rewriting.
Nvidia’s two-decade library of hand-tuned kernels is only a moat for as long as kernels have to be written for Nvidia.
TileLang pulls those pieces apart. The dataflow – the maths – gets written once. The schedule – the hardware part – is separate and swappable. Point the same kernel at different silicon and in theory you only have to rewrite the schedule rather than the whole thing.
So DeepSeek both uses TileLang and benefits from it, but Liang’s claim is overstretched, because the schedule is where you get nearly all the performance. Separating it out doesn’t hand you a fast kernel on Huawei silicon. It just means you’re having to rewrite the hard half of the code instead of the whole thing.
There’s an Nvidia irony in Liang’s account of DeepSeek’s V3 training. To reconfigure Nvidia’s H800 streaming multiprocessors, DeepSeek allocated 20 of the chip’s 132 processing units to the communication between servers. To manage it they wrote in PTX, Nvidia’s assembly-level instruction layer, which is about as un-portable as code gets. That meant going deeper into Nvidia’s proprietary stack rather than around it. DeepSeek, far from being independent of Nvidia, was actually trying to hack Nvidia’s own architecture to make it work harder. So take Liang’s ‘Nvidia is headed for the rocks’ argument with a pinch of salt.
The third claim, though, is probably Liang’s strongest. He argues that the compute GPU market is now bigger than the gaming GPU market. So, AI accelerator chips are decoupling from the gaming chips that CUDA was built to serve, and purpose-built AI chips won’t need CUDA compatibility with its gaming legacy.
Liang’s got a real point here. Nvidia’s dominance was partly an accident, because gaming GPUs turned out to be very, very good for AI researchers. If you’re designing accelerator chips from scratch for transformer workloads, you don’t have to inherit a long history of converting gaming GPUs into AI chips.
And on the software side, things are becoming more portable. OpenAI’s Triton lets developers write kernels in Python that can compile to either Nvidia or AMD hardware, and AMD’s ROCm 7 delivers up to three and a half times better inference performance than previous versions.
And the analyst consensus agrees with Liang that the software library part of Nvidia’s moat is indeed starting to disappear.
But the networking layer, NCCL, tied to NVLink and InfiniBand – the one that lets chips talk to each other – remains hard to get off, because any lab porting its code will hit that dependency the moment it runs distributed training across many machines.
The most concrete evidence for Liang’s third argument is Huawei’s decision to open-source its CANN toolchain, a CUDA-equivalent for its own Ascend chip. At an industry conference in Beijing in August 2025, Huawei’s Xu Zhijun announced that CANN and the ‘Mind’ series of toolkits would be open-sourced by year end. Chinese media reported it as a direct shot across the bows of Nvidia.
By 2026, Huawei was claiming that CANN supported PyTorch, vLLM, SGLang, Triton and TileLang, and had aligned more than 2,300 common PyTorch APIs so that mainstream models could run on Ascend silicon with barely any modification.
3 Four times everything
Besides the Nvidia dunks, there were also some very specific hardware numbers dropped in the investor call.
Liang says Huawei allocates DeepSeek capacity for about 16,000 cards, whilst China’s internet giants get hundreds of thousands. And, Liang says, this may already be all the capacity Huawei has. He says Huawei is two years behind Nvidia and, per chip, four times worse.
But Huawei’s strategy is not to fight silicon with silicon. It unveiled its Ascend roadmap in September 2025, and its 950 series introduced a proprietary high-bandwidth memory (HBM).
Huawei knows Nvidia’s Blackwell chip is way ahead – Nvidia’s GB300 delivers around 15 petaflops of low-precision compute per chip – and so it has bet instead on lashing together enormous numbers of weaker chips. Huawei’s Atlas 950 SuperPoD is designed to connect 8,192 of its Ascend chips. The firm showed off a machine in July 2026, with the promise it’ll ship the real thing in the fourth quarter.
It’s not rocket science. System for system, if you wire all those Huawei chips together tightly enough, they’ll finish exactly the same training run as Nvidia’s. You just need four times the silicon, four times the power, and four times the floor space.
Floor space is cheap in China. Land in the west goes to strategic projects for next to nothing – it was written into the ‘East Data, West Computing’ plan from the start, and Inner Mongolia now has a few gigawatts’ worth of projects, with Huawei and ByteDance among the tenants. So model training can move west, because unlike inference it doesn’t care about latency, but not much of it has.
China’s energy edge is a little narrower than is claimed – the Oxford Institute for Energy Studies reckons Chinese power prices are lower than Europe’s but broadly comparable to those in the United States, because of excess capacity and under-used assets.
Where the crunch really comes is still in the silicon, and most importantly in the memory inside it. Four times the chips means four times the number of HBM stacks, and HBM is what China can’t yet manufacture at volume. But give it time.
Training a frontier model requires thousands of chips working together as one computer, constantly exchanging data back and forth. The bottleneck for that is often the speed at which they can talk to one another and to their memory, rather than their real processing power.
A supernode is a densely interconnected cluster that’s engineered so that hundreds or thousands of chips can behave, so far as the software is concerned, like one gigantic accelerator.
And this is why memory bandwidth and HBM matter so much, especially for inference. When a model is repeatedly generating text, it must read its enormous parameter set and its KV cache, which is the running memory of the conversation, from its HBM. If the arithmetic units are sitting idly waiting for their data to arrive, then the computing power goes to waste. High-bandwidth memory is what keeps everything moving.
China’s HBM dependence is a real chokepoint, which is why Huawei developing its own version, and the memory-chip ambitions of CXMT, the mainland’s largest mass producer of DRAM, matter so much for China’s AI development.
Liang admits that because of their higher power consumption Huawei chips depreciate faster – three years against five years for Nvidia. He also says Nvidia’s B200s can’t be bought for love or money in China, which reflects the US decision to block Chinese access to the company’s Blackwell chip suite.
Nvidia names its generations by scientist. H is for Grace Hopper – the H100, the H800 and the China-market H20. B is David Blackwell – the B200 and the GB300 – and it’s a generation ahead.
So how much Nvidia silicon is actually inside of China? Before April 2025’s restrictions kicked in, ByteDance, Alibaba and Tencent reportedly tried to order around a million H20s between them. The H20 was Nvidia’s deliberately cut-down China part, throttled to sit just under the export-control threshold, and a million of them would have been the company’s entire annual run of the chip. Then there’s the black market. Research group Epoch AI reckons about 660,000 H100-equivalents were smuggled in through the end of 2025, about a third of China’s total compute.
DeepSeek’s own 20,000 H-equivalents – everything it owns, normalised into Hopper-generation cards, most of them recently arrived – look pretty modest beside that. It’s consistent with a lab that historically ran on its hedge fund parent’s compute.
And the Huawei order doesn’t shift it much. Liang notes 16,000 Huawei 950s come to only about 4,000 Nvidia B-series cards. A Blackwell card is worth two or three Hoppers, so call it another 10,000 H-equivalents. Half as much again. Not a big amount, only enough to train this generation, he says. And the point of buying them, he says, is partly to help Huawei mature its own ecosystem.
4 The parameters that count
A model’s parameters are the numbers that it learns during training, its knowledge encoded as weights. For years, models were dense, which means that every single parameter was used to process every single word.
A mixture-of-experts (MoE) model contains many specialised sub-networks, and for each token a router activates only a small fraction of them. So, a model can have a vast total parameter count – its accumulated knowledge – whilst only activating a small subset per token, which determines its running cost. A little bit like how the human brain manages to run on 20W of power – most neurons are just sitting idle.
The gap between the two is quite dramatic. OpenAI’s gpt-oss-120b (which I run at home) has 116.8 billion total parameters and 5.1 billion active per token, which is a ratio of about 23 to 1. Kimi K3 has 2.8 trillion total parameters but only activates 16 of its 896 experts per token, which one technical write-up calls an unusually high sparsity ratio. DeepSeek V3 activated about 37 billion of its 671 billion. This is why a trillion-parameter model isn’t as expensive to run as it sounds. You’re only really paying for what gets activated.
Liang claimed on the call that the largest US models now activate around 800 billion parameters, whilst Chinese models sit at tens of billions. That puts China an order of magnitude behind. Is that 800 billion number plausible?
Nobody outside OpenAI and Anthropic really knows, and the published estimates are a mess. The guesses on total parameter counts for GPT-5 class models are anywhere from one and a half trillion to ten trillion and upwards – and the most frequently cited estimate comes from a paper whose own reviewers concluded their method wasn’t reliable. Numbers on active parameters are even flakier.
What we can assess more accurately are the models we’re allowed to peer inside of. OpenAI’s gpt-oss-120b activates one parameter in 23. GLM 5.2 runs 40 billion active out of 744 billion. DeepSeek V3 was 37 out of 671. Sparsity ratios of 15 or 20 to one are normal. Apply that to a frontier model of two or three trillion parameters and you land at active parameters in the low hundreds of billions.
So 800 billion looks like a high estimate of the American lead. But Liang’s conclusion is what’s interesting. Training a model that activates only 800 billion parameters would require, on his estimate, around 50,000 Nvidia GB300s or 200,000 Huawei 950s.
Even spending its entire RMB 50bn ($7.4bn) raise, DeepSeek couldn’t afford to train or run such a model.
So DeepSeek will stay at the tens-of-billions activation layer and push later on towards 150 to 250 billion, with a 150 billion model starting training late in 2026.
This is what the compute gap really means. China can build very good models. It can’t yet build the very biggest ones. So, whilst China’s talent gap in AI is close to zero, in Liang’s view the gap with the US comes down to one thing: resources.
5 Cheap tokens, expensive PhDs
Liang also tells us about the economics of inference. When you train a model, it’s an enormous one-off capital cost. You’re putting thousands of chips to work for months. Training is the cash furnace keeping OpenAI and Anthropic from being profitable.
Inference is different. That’s the ongoing cost of running a finished model to answer a user’s queries. A token is roughly a word fragment, and models are billed per million tokens of input and output. And inference served at scale is increasingly very profitable.
Liang says that DeepSeek prices its API so that its hardware pays back in 10 months, which means a six-fold profit over a five-year lifespan. At this pricing, third parties can’t profitably re-host DeepSeek’s own open-weight models.
SemiAnalysis estimates inference margins at the major labs rose from less than 40 per cent to more than 70 per cent during 2025–26. Anthropic’s inference gross margins are reportedly in the mid-60s today, up from 38 per cent in 2025, and minus 94 per cent in 2024. DeepSeek itself has claimed more than 80 per cent inference margin on R1.
How does this affect strategy? In May 2026, DeepSeek cut its flagship V4-Pro prices by 75 per cent, taking it to $0.435 per million input tokens and $0.87 per million output. That makes it roughly 11.5 times cheaper than GPT-5.5 on input and 34.5 times cheaper on output.
It also made a grab for agentic workflows, cutting its cache-hit pricing to $0.003625 per million. What’s cache-hit pricing? When an AI agent works through a long task, it repeatedly resends the exact same context – a codebase or a document. A cache hit is that repeated input, and billing it at a fraction of the full rate transforms the economics of long-running agentic workflows.
Set this beside Liang’s claim that demand is price-inelastic, and that doubling prices would nearly double revenue, and cutting prices is a very deliberate strategic choice.
DeepSeek is buying developer mindshare and setting a market floor to bleed out competitors rather than maximising its near-term margins.
The other half of the cost story goes the other way. Liang says that US and Chinese data annotation costs are about the same, and that China doesn’t have a cost advantage in high-end data. He also says that half his core researchers annotate data themselves.
This is one of his most counter-intuitive claims. The frontier of AI training no longer runs on Amazon Mechanical Turk workers drawing boxes around images. It runs on expert judgement for reinforcement learning.
Surge AI, which is a major data provider for Anthropic’s Claude, crossed one billion dollars in revenue with a network of some 50,000 expert contractors. Built In reports its rates running from the low hundreds of dollars for medical fellows and management consultants up to four figures for venture capital partners.
Mercor, a rival marketplace, pays doctors and lawyers comparable professional rates, and high-quality RLHF annotations – reinforcement learning from human feedback, where an expert rates what the model produced – run on the order of $100 each. The scarce input here is PhD-level human expertise. And Chinese PhDs in medicine, law and mathematics aren’t dramatically cheaper than the global market for this kind of work, nor can crowd labour really substitute for it.
The old cheap Chinese labour assumption doesn’t apply to frontier data.
6 Open weights, closed ranks
It’s easy to confuse open weights with open source. Open-source software ships its source code, which anyone can read and modify. Open-weight models only ship their final trained parameters, an enormous grid of numbers, for anyone to download, run and fine-tune. The training code and the training data stay private.
In other words, you can use a model and build on top of it, but you can’t fully reconstruct how it was made. This is how DeepSeek, Alibaba’s Qwen, Z.ai, MiniMax and Moonshot roll. And it’s what’s driving companies like Anthropic to get open-weight models restricted.
Moonshot’s Kimi K3, which was released on 16 July 2026 (with full weights promised by 27 July), is a 2.8-trillion-parameter MoE model, the first open model in the 3-trillion-parameter class.
It scored 57.1 on the composite Intelligence Index run by Artificial Analysis, an independent benchmarking outfit, third behind Claude Fable 5 on 59.9 and GPT-5.6 Sol on 58.9. It took first place for building web interfaces on LMArena with 1,679 points, ahead of Fable on 1,631 and GPT-5.6 Sol on 1,618, the first open-weight model to top that ranking. LMArena runs blind – which means that users vote without being told which model was involved – so K3’s victory counts for more than a typical benchmaxxed score.
Three days after launching, Moonshot had to pause new subscriptions, saying that massive demand had pushed it close to capacity.
That’s also a lesson in the scale gap. Chinese labs can build frontier models but they can’t yet afford to serve everyone who wants them.
K3 also broke the pricing pattern that made Chinese models so attractive. It launched at three dollars per million input tokens and 15 dollars per million output, which was triple K2.6’s rates, making it the most expensive release to date from a Chinese lab. The compute gap is starting to hit the price wall as well as serving capacity. If DeepSeek’s next flagship follows K3 upwards, then the ‘overcapacity’ reading of China’s AI strategy becomes considerably weaker.
Washington’s response has been to focus on distillation, which is a training technique where a smaller model learns from the outputs of a larger one. It does that by querying the strong model millions of times and training on its answers.
After Kimi K3’s release, White House science and technology policy chief Michael Kratsios took to social media to cry foul:
“We have information that Moonshot AI distilled Anthropic’s Fable for the development of its K3 model … large-scale, covert industrial distillation aimed at stealing proprietary US technology and undermining American research is unacceptable.”
In June 2026, Anthropic had told the Senate Banking Committee that Alibaba’s Qwen lab was behind thousands of fraudulent accounts that spent weeks generating millions of responses from Claude, targeting software engineering and agentic reasoning. Alibaba denied the claim. This followed allegations back in February which pointed to distillation campaigns from MiniMax, Moonshot and DeepSeek.
Congressional committees have already probed US firms, reportedly including Airbnb and Cursor, over their use of Chinese AI models.
Of course, there are some pretty important counter-arguments. The first is enforceability. Once open weights are published and downloaded worldwide, then banning them is close to impossible, and some analysts reckon that published weights may even qualify as protected expression under the First Amendment.
Secondly, there’s the definition. Distillation is a very standard technique, and American labs routinely distil their own large models into smaller ones.
Last, and most pointedly, there’s the charge of hypocrisy, which even Microsoft’s chief executive, Satya Nadella, has nodded towards. Frontier labs claim fair-use rights to train their models on the world’s public data, and then want to clamp down on distillation of their own outputs.
Whether extracting a model’s behaviour through its API is theft or fair use remains contested. But the concept is doing some heavy political lifting and the decision will likely be fought in the lobby shops of K Street.
The diplomatic half of the strategy was unmissable at the World AI Conference (WAIC) in Shanghai from 17 to 20 July 2026, where Xi Jinping delivered his first keynote in the conference’s nine-year history.
Xi called on countries to embrace the historic opportunity of open-source AI. The day before WAIC opened, 29 countries – including Russia, Brazil, Indonesia, Pakistan, Kazakhstan and South Africa, but no G7 country – signed the founding agreement of the World Artificial Intelligence Cooperation Organisation (WAICO), a China-led intergovernmental body headquartered in Shanghai.
China is positioning open-source AI as a global public good for the Global South, and it’s claiming WAICO as an alternative to Western frameworks like America’s Pax Silica, the State Department’s AI and supply-chain coalition launched in December 2025.
Whether that translates into real governance clout is up for debate. Rules that are written in Shanghai without G7 participation face a fairly tough adoption test. But Beijing’s open-weight strategy and its diplomatic strategy are one and the same.
If China’s models are free to download, then its governance bloc will set the norms for most of the world’s countries, and the American lead in raw capability will matter less.
7 Involution and its discontents
DeepSeek’s first external round closed in June 2026, and raised around RMB 50bn, giving a post-money valuation of over RMB 417bn. Liang’s company is famously secretive, and the figure only surfaced through a regulatory filing by someone who’d invested indirectly through a fund. Liang personally chipped in around RMB 20bn of that, and he retains control through a limited partnership that he manages. External investors reportedly don’t receive voting rights and face a five-year lockup.
His backers include Tencent, at around RMB 10bn, the battery maker CATL, JD, NetEase and IDG. China’s state-backed National AI Industry Investment Fund also put in a modest RMB 1bn directly and was, according to The Information, the only investor granted direct equity and voting rights.
A second round of RMB 50bn at a RMB 500bn valuation was reported within weeks. And DeepSeek is reportedly preparing to list on Shanghai’s STAR Market, the mainland’s Nasdaq-style board for tech companies.
When Liang claims the Chinese government won’t give him a dime, he’s not quite telling the whole truth. But when Beijing wants to invest – even 2 per cent of your round – it’s smart for Chinese entrepreneurs to give the government good terms.
Liang’s diffidence is probably more revealing as a statement of his self-image. He built DeepSeek on his own hedge fund fortune, and he pushes back on claims that the company is a state project. Still, you can’t afford to ignore the government in white-hot technology spaces – as Anthropic learned on the other side of the geopolitical black mirror.
Liang’s entanglement with the state industrial project runs the other way too. When CXMT priced its own STAR Market listing in July 2026 – raising around $8.6bn in Asia’s largest IPO of the year and the biggest Chinese A-share semiconductor offering ever – Liang’s High-Flyer Quant had more than 150 fund products bidding. China’s National Social Security Fund and Alibaba Cloud were in the strategic placement alongside him. The man who can’t buy high-bandwidth memory abroad is underwriting China’s best shot at making it.
Liang Wenfeng, with a personal fortune of $36bn, is simultaneously a Chinese AI founder, a major quant investor in China’s chip supply chain, and a somewhat grudging recipient of government money.
China’s AI strategy has been compared to its strategy on solar panels, steel and EVs – make technology cheap enough to dominate the space globally. The analogy is not unreasonable. In solar, Chinese firms scaled up aggressively behind state and local government support, and the country now produces over 80 per cent of the world’s panels. Overcapacity has collapsed margins and led to three years of falling factory gate prices and a slew of bankruptcies.
The Chinese government’s response was a campaign against self-defeating price competition – ‘anti-involution’ – which it’s since extended to EVs, batteries and steel. That involves capacity control, a reported RMB 50bn fund to retire excess capacity, and coordinated production cuts. By late 2025, polysilicon prices had recovered around 50 per cent, and leading firms like GCL had swung back into a modest profit. But for countries that had seen their nascent solar industries wiped out, there was nothing.
This analogy captures something very real about AI. There is a price war, near-identical frontier labs have proliferated, and there’s been a willingness to sacrifice margin for market share. DeepSeek’s Liang argues that China has too many foundation model labs whilst America has just three.
The analogy can’t be stretched too far. Solar panels are a commodity – the product is physically identical and the competition is purely on cost. AI models differ in capability, and what they have in common is scarce compute and scarce talent. A solar factory that overbuilds piles up unsellable panels. An AI lab that overbuilds turns out better models.
But there is also an ‘anti-involution’ logic – that China will eventually manage a consolidation to three or four national champions, through either state intervention or market pressure.
What makes Liang’s investor call outpouring so interesting is how differently he sees the world from policymakers in Washington or Beijing.
Yes, he overstates how fast Nvidia’s moat is draining, and DeepSeek’s distance from Nvidia’s tech – and TileLang isn’t his, and his own engineers burrowed further into Nvidia’s stack than almost anyone has.
He gripes about having to fight for what outsiders assume Beijing has given him. He lets slip that DeepSeek itself can’t afford to train the very largest models and won’t be able to for some time.
And yet export controls haven’t stopped Chinese labs reaching for today’s frontier – Kimi K3, GLM 5.2 and DeepSeek V4 Pro are out, and the first topped a leaderboard this month.
All they have imposed is a tax. And for now, China can pay it.
Thanks for reading!
Bis bald,
Adrian



