DeepSeek’s logo and Huawei’s Building F1 at its Bantian campus in Shenzhen, photographed in February 2025. Image: RoadMaster19, CC BY 4.0 (DeepSeek logo); Liuxingy, CC BY-SA 4.0 (Huawei Building F1, Shenzhen, 2025)

DeepSeek has open-sourced Huawei Ascend versions of the low-level software it uses to train and run its AI models, matching almost every tool it had previously built for Nvidia’s chips. The Chinese lab announced the release on Wednesday on its official WeChat account, saying it had worked with Huawei on the project.

The headline piece is TileLang, a programming language for writing the small, performance-critical programs (kernels) that AI chips run. DeepSeek said that to build “a new generation of independent, self-controlled GPU software ecosystem,” the first priority is a high-level language, and that TileLang offers a simpler programming model than Nvidia’s CUDA.

What DeepSeek released

Every component corresponds one-to-one with a tool DeepSeek had already published for Nvidia hardware. All of them target Huawei’s new Ascend 950 chips:

  • TileLang: the open-source language, started by a Peking University team, now “officially supports Huawei Ascend 950 NPUs with native code generation,” according to its project page.
  • DeepGEMM Ascend: the maths library that does the heavy matrix multiplication. Its documentation says it is “fully API-compatible” with the Nvidia version and supports the FP8 and FP4 formats used in modern training.
  • DeepEP Ascend: the communication library that shuffles data between chips when a model is split across many of them.
  • FlashMLA: DeepSeek’s attention kernels, which now power its DeepSeek V4.1 model “on NVIDIA GPUs and Huawei Ascend NPUs.”
  • TileKernels and DeepSelect: smaller libraries for everyday data operations and for the TopK step in DeepSeek’s sparse attention.

The point is portability. TileKernels now ships a second backend that is picked automatically at runtime, “so the same Python APIs run on both NVIDIA GPUs and Huawei NPUs.” For a developer, switching from Nvidia to Huawei could, in theory, mean installing a different package rather than rewriting code.

The numbers

DeepSeek published benchmark figures alongside the code, and they are its own measurements, not independent tests. On the Ascend 950, its FlashMLA attention kernels reach up to 410 teraflops during prefill, which it says is 95% of the hardware’s peak, and 360 teraflops, or 83%, during decoding.

The DeepEP Ascend documentation says sustained data transfer reaches “roughly 90 to 95% of the physical payload bandwidth limit” for jobs spread across up to 32 chips, with larger setups still being optimised. DeepSeek said it and Huawei had tuned the software for a 128-chip Ascend 950 “supernode.”

Huawei also gets a thank-you in the code itself. The TileKernels project “gratefully” acknowledges Huawei “for its technical support and engineering expertise.”

Why CUDA matters so much

Nvidia’s grip on AI is as much about software as silicon. CUDA has had nearly two decades to become the default way to program AI chips, and most of the world’s AI code is written for it. Rival chipmakers, from AMD to Chinese firms, have struggled because developers don’t want to rewrite everything to switch.

This release is DeepSeek’s attempt to remove that excuse, at least for its own style of model. It follows founder Liang Wenfeng’s bet that training on Huawei chips “must succeed”, and comes as Beijing weighs whether to let its tech giants buy Nvidia’s newest chip and Alibaba pushes its own AI chip. Huawei hasn’t published its own statement on the release.

Why it matters

US export controls were meant to slow China’s AI progress by cutting off Nvidia’s best chips. If DeepSeek’s tools make Huawei’s hardware nearly as easy to use as Nvidia’s, the hardest part of switching away from Nvidia gets much easier, leaving the supply of Huawei chips as the main bottleneck.

Sources: DeepSeek’s official WeChat announcement (Sep 30), TileLang, DeepGEMM Ascend, DeepEP Ascend, FlashMLA, TileKernels and DeepSelect on GitHub

Latest More Labs news

More More Labs news