DeepSeek and Huawei Are Coming for CUDA, Not Nvidia’s Chips
DeepSeek’s latest release on Oct. 1 was code, not a model. Through WeChat, the Chinese lab announced support for Huawei’s Ascend chips in TileLang, a kernel programming language, plus six open-source software modules for the Ascend platform. According to The Next Web, which draws on Reuters, Bloomberg and the South China Morning Post, the modules include compute and communication libraries. They “mirror the open-source tools it had already published for Nvidia chips.” DeepSeek said Huawei gave “full support during development.” The two companies also built a “supernode” from 128 Ascend 950 chips. Bloomberg reported separately that DeepSeek plans to deploy at least 160,000 Huawei accelerators in an Inner Mongolia data centre.
The 160,000 figure will get the attention. The software matters more.
The moat is a compiler
Nvidia’s hardest advantage to copy is CUDA, the proprietary programming platform it has shipped since 2007. The GPUs matter, but two decades of kernels, libraries, profilers and tutorials are written for CUDA. The New York Times has estimated, as cited by TNW, that it has about four million developers. If an AI team moves to another chip, it has to rewrite or re-tune the low-level code its models depend on. Then it has to trust that code in production.
AMD shows how hard that is even with good hardware. In December 2024, SemiAnalysis benchmarked AMD’s MI300X against Nvidia’s H100 and H200 and concluded that “the CUDA moat has yet to be crossed by AMD”. It blamed software quality, not the chip.
DeepSeek’s own statement, in Reuters’ rendering, says the same thing in industrial-policy language: “To build a new generation of independent, self-controlled GPU software ecosystems, the first priority is establishing a high-level language that is universal, easy to program.”
What TileLang actually is
TileLang isn’t new, and it isn’t only DeepSeek’s. Its GitHub project describes it as a “domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels.” It uses Python-style syntax on top of the TVM compiler stack and started at Peking University, with work done during a Microsoft Research internship. Its primary backend is still Nvidia CUDA. It also supports AMD’s ROCm and Apple’s Metal. External adapters for Ascend have existed in the project since September 2025.
DeepSeek brings demand. TNW reports that the lab uses TileLang as its “core tool,” and DeepSeek already publishes a kernel library written in TileLang next to its CUDA-side projects like DeepGEMM and the DeepEP expert-parallel communication library. If a widely studied open-model lab writes its kernels in a language that compiles to both Nvidia and Huawei hardware, switching chips becomes a recompile rather than a rewrite. That’s the theory, at least.
China’s full-stack push
Huawei is making the same argument with hardware. It has promised an Atlas 950 SuperPoD of up to 8,192 Ascend 950DT chips in the fourth quarter of 2026. It also pledged to open-source its CANN software stack by the end of 2025. Next to that, DeepSeek’s 128-chip node is modest.
Alibaba is moving the same way. On Sept. 22 its T-Head chip unit unveiled the Zhenwu V900. It has 216GB of memory and native FP8 and FP4 support, and Alibaba says it delivers three times the performance of the earlier M890. TrendForce reports that mass production is scheduled for the first quarter of 2027. Alibaba Cloud is also targeting more than 20 gigawatts of data-centre capacity by 2032. Neither report says how developers will program the V900, which is the same gap DeepSeek is now addressing for Ascend.
Will it matter outside China?
Our read: not soon. Ascend tooling would need to reach the places developers already work. That means upstream support in the major frameworks and inference servers, Ascend hardware that people outside China can actually rent, and enough independent benchmarks to show the six modules hold up outside DeepSeek’s own clusters. None of that has been announced.
Inside China, the calculation is different. The labs there have reasons to want an alternative to Nvidia, and now one of the most closely watched open-model labs has published its tools for Huawei’s chips. Still, TileLang’s README lists Nvidia CUDA as its primary backend.
Sources
- The Next Web: DeepSeek and Huawei release open-source Ascend tools to rival CUDA
- Tom's Hardware: DeepSeek and Huawei release open-source Ascend AI programming tools
- GitHub: tile-ai/tilelang
- GitHub: deepseek-ai
- Wikipedia: CUDA
- SemiAnalysis: MI300X vs H100 vs H200 benchmark, part 1: training
- Huawei: Huawei Connect 2025 keynote, Ascend roadmap
- TrendForce: Alibaba unveils AI chip Zhenwu V900 for 1Q27 mass production
- TechNode: Alibaba unveils Zhenwu V900 AI chip as it targets 20GW of cloud data centers by 2032
James Whitfield covers hardware for prompt/power: chips, semiconductors, laptops, components and the benchmarks behind the launch-day claims. He thinks the most important number on any spec sheet is usually the one in the footnote.
Latest from prompt/power
- DeepSeek V4.1 Flash Cut the US AI Lead to 3%. What LiveBench MeasuresOct 6
- Ben Affleck Calls AI Job Fears ‘Propaganda.’ He Sold an AI Firm to NetflixOct 6
- Cohere’s North 2 Puts AI Agents on a Budget. Toronto Bets on BoringOct 6
- How to Stop ChatGPT, Claude, Gemini and Meta AI From Training on Your ChatsOct 6
- Musk Is a Trillionaire Again After SpaceX Stock Jumps 7.6%Oct 6
Leave a Reply