DeepSeek is taking its most direct shot yet at Nvidia's software moat. The Chinese AI company announced today, in a post on its official WeChat account, that it is open-sourcing a full programming stack for Huawei's Ascend AI chips — an Ascend port of its TileLang language plus the compute and communication libraries it uses to train its own models. Reuters reported the announcement on Wednesday, September 30.

The stakes: CUDA is why Nvidia is so hard to leave. Developers have a decade of optimized code written against it, and moving to another accelerator means rewriting that stack by hand. DeepSeek is offering to erase that cost for Chinese developers willing to jump to Ascend — with a toolkit that, DeepSeek says, maps one-to-one to the versions it previously released for Nvidia hardware.

What's in the release#

The package covers the whole AI-kernel pipeline. The centerpiece is the Ascend version of TileLang, a high-level programming language and compiler for writing high-performance AI kernels, now wrapping Huawei's low-level Ascend C instructions. Around it sit the libraries: DeepGEMM for general matrix operations, DeepEP for large-scale, high-speed communication between chips, TileKernels for vector computation and memory access, FlashMLA for sparse attention, and DeepSelect for data filtering.

TileLang is positioned as the wedge against CUDA. "To build a new generation of independent, self-controlled GPU software ecosystems, the first priority is establishing a high-level language that is universal, easy to program, and still capable of reaching the hardware's full performance potential," DeepSeek said. "TileLang was created precisely to meet this need," and offers "a simpler programming model" than Nvidia's CUDA.

Illustration of an AI accelerator chip on a motherboard with glowing data streams
Illustration: AI Frontier Post (AI-generated).

Notably, DeepSeek says every TileLang operator used in training its own V4 models now has a high-performance Ascend implementation — this is the stack behind its own workloads, not a demo port. The same Python interfaces support Nvidia GPUs and Huawei NPUs, with the backend selected automatically. A community-maintained Ascend adapter, tilelang-ascend, has existed since September 2025; this release is DeepSeek's own official port.

The 128-chip Ascend 950 supernode#

Software is only half the story. Huawei provided full support in developing the programming infrastructure, and the two companies jointly advanced a "supernode" solution based on 128 Ascend 950 chips, optimizing both computation and communication across the cluster.

A supernode wires dozens to hundreds of accelerators into one tightly integrated system — the kind of architecture that matters for training and serving frontier models, where inter-chip communication is as decisive as raw per-chip FLOPs. DeepEP, the large-scale communication library in the release, is built for exactly that job.

Illustration of a server supernode with glowing accelerator modules linked by fiber
Illustration: AI Frontier Post (AI-generated).

Why software decides the chip war#

Hardware alone has never been enough. Nvidia's real moat is CUDA: the frameworks, libraries, documentation and developer habits that make GPUs practical. Huawei has been building its own software layer — its Compute Architecture for Neural Networks (CANN), with support for PyTorch, Triton and vLLM — and DeepSeek's release adds the missing application-level tooling.

This is also a mirror of DeepSeek's own engineering culture. The company that open-sourced FlashMLA, DeepGEMM and DeepEP for Nvidia hardware has now rebuilt that toolkit from scratch for Huawei silicon. If it works, a Chinese developer team can keep its high-level code and let TileLang target whichever accelerator is available — which is precisely the portability that has kept everyone on CUDA.

Timing and context#

The announcement came two weeks after Huawei unveiled its next generation of AI processors and supernode computing systems, which Huawei says it expects to see widely used for model training next year. DeepSeek's V4 model had already shipped with support for Huawei's Ascend processors — today's release turns that one-off adaptation into reusable public infrastructure.

And the collaboration is deepening beyond software: on the same day, DeepSeek was also reported to be preparing a Shanghai IPO, with CITIC beginning due diligence — the company is positioning itself as national infrastructure, not just a model vendor.

What to watch#

All of the claims above come from DeepSeek's own announcement; none have been independently benchmarked. "A simpler programming model than CUDA" is a promise, not a measurement — and the "independent, self-controlled GPU software ecosystem" framing reflects Beijing's push for technology sovereignty as much as developer convenience. But this is not a whitepaper: it's the actual toolkit DeepSeek uses to train its models, released in full. The test will be whether other Chinese labs and cloud providers start shipping Ascend workloads with it.