Software

DeepSeek Open-Sources Six Module Software Stack to Challenge Nvidia CUDA on Huawei Ascend

The Hangzhou lab released an Ascend 950 version of its TileLang language plus five companion libraries, betting that simpler programming can loosen Nvidia's grip on AI software.

T
By TechQuire Daily Staff TechQuire Daily Staff
October 3, 2026 / 7 min read

DeepSeek released six software modules built specifically for Huawei Technologies and its Ascend line of AI chips on September 30, 2026, publishing the code as open source through its official WeChat account and framing the work as a step toward a new generation of independent and controllable GPU software ecosystem. The modules map one to one onto an earlier set of tools the Hangzhou based lab had released for Nvidia hardware, a deliberate pairing that makes the competitive target of the release unmistakable.

The headline component is an Ascend compatible version of TileLang, a programming language and compiler meant to simplify the creation of high performance kernels, the small programs that carry out the core computations inside AI models on GPUs and CPUs. TileLang had treated Nvidia as its primary backend. The new release adds support for the Huawei Ascend 950 accelerator, bringing native code generation, automatic scheduling and synchronization, according to an update to the project page on GitHub.

Timing matters here. The release lands in a market reshaped by American export controls that bar Chinese companies from buying Nvidia's most advanced processors. Huawei has widened the Ascend family to fill the gap, and DeepSeek has been moving toward that silicon for most of 2026. Earlier in the year the lab published a preview of its V4 model adapted for Ascend processors, a departure from its previous reliance on Nvidia chips, and Huawei said at the time that it had coordinated with DeepSeek to keep V4 compatible across its Ascend products.

TileLang itself is not a DeepSeek invention. Researchers at Peking University developed the language and published it as an academic paper in April 2025, catalogued as arXiv:2504.17577v1 by Wang and colleagues. DeepSeek has been using it for roughly a year and now describes it as a core tool for its artificial general intelligence research. The language already lists Nvidia CUDA, AMD ROCm and Apple Metal among its backends, which means the Ascend 950 support makes Huawei the fourth hardware target rather than the only one.

Key Facts

The South China Morning Post reported on September 30 that DeepSeek published six software modules that correspond one for one with the tooling it had previously open sourced for Nvidia AI chips. The six are TileLang, DeepGEMM-Ascend, DeepEP-Ascend, TileKernels, FlashMLA and DeepSelect. The stated objective is a new generation of independent, controllable GPU software ecosystem, with the Ascend version of TileLang presented as an example that could serve a broader range of AI chips.

The supporting modules address the unglamorous work of running large models. DeepGEMM-Ascend handles matrix multiplication kernels and keeps compatibility with the original DeepGEMM application programming interface while supporting BF16, FP8 and FP4 number formats tuned for Ascend characteristics; the project team says a single kernel is designed to approach the hardware performance limit. DeepEP-Ascend manages communication between large numbers of chips, the bottleneck that constrains mixture of experts architectures. TileKernels covers vector operations and memory access, FlashMLA targets long context processing and supports the extended context window that V4 introduced in April, and DeepSelect filters data. ForkLog reported on October 1 that Bloomberg described the package as China's answer to Nvidia's CUDA.

Quartz reported on September 30 that DeepSeek disclosed a partnership with Huawei covering libraries for both computation and communication workloads, and that the two companies jointly developed a supernode system built from 128 Ascend 950 chips, optimized for compute and for data exchange between accelerators. The Next Web reported on October 1 that DeepSeek said Huawei provided its full support during development, and that Reuters journalists Che Pan and Farah Master described the release as including compute and communication libraries for the Ascend platform.

Scale makes the ambition concrete. The Next Web reported on October 1 that DeepSeek plans to deploy at least 160,000 Huawei accelerators at a data center in Inner Mongolia, and that the company is raising money at a valuation close to 500 billion yuan, or about 74 billion US dollars, having reached an annualized revenue run rate of 1 billion US dollars in the week before the launch. Huawei separately announced in September that its next generation Ascend 960DT chip will arrive in the first quarter of 2027, three quarters earlier than originally planned, alongside a roadmap running to 2029. Rotating chairman Eric Xu said Huawei believes its Ascend chips have taken more than Nvidia's share of China's AI chip market, but he offered no data to support the claim.

The competitive reference point is formidable. TechTimes reported on October 2 that Nvidia said at GTC 2026 that its CUDA developer base has passed 6 million people, up from 4.5 million in 2025, supported by more than 400 GPU accelerated libraries including cuDNN, cuBLAS and NCCL. CUDA was first released to the public in 2007. An earlier estimate from The New York Times, in reporting by Steve Lohr, put the CUDA developer community at roughly 4 million, a reminder that the installed base is both enormous and still growing.

Analysis

The bigger picture here is that DeepSeek is attacking the part of Nvidia's dominance that outlasts silicon: the software habit. CUDA's advantage is not that it is fast in the abstract but that millions of developers already know it, thousands of libraries assume it, and switching costs are measured in retraining and rewriting rather than in dollars per chip. A credible alternative does not need to beat CUDA on every benchmark. It needs to make a second option cheap enough that buyers and developers keep it on the table.

Portability is the mechanism. TechTimes reported on October 2 that a developer who adopts TileLang on Nvidia hardware is positioned to run the same code on Huawei Ascend processors without rewriting a single kernel. That property cuts two ways. It gives Huawei a migration path that does not require its customers to abandon years of work, and it gives developers a hedge that also covers AMD ROCm and Apple Metal, both of which TileLang already supports as backends. What this really means is that the strategic target is not Huawei's hardware at all but the assumption that AI code is written once, for one vendor, forever.

The caveats deserve equal billing. DeepSeek is a motivated party: it wants its own models, including V4, to run well on hardware it can actually buy at scale, and it plans to install at least 160,000 Huawei accelerators in Inner Mongolia. Every module it released serves that goal, which does not make the engineering less real but does mean that independent benchmarks will matter more than launch claims. The performance assertions, including the statement that some supernode test cases run close to the hardware limit, come from the project side. The 128 chip supernode was optimized jointly with Huawei, so it reflects a tightly controlled configuration rather than the messy heterogeneity of a typical customer cluster.

Nvidia's counterweight is depth. Six million CUDA developers, more than 400 accelerated libraries and close to two decades of accumulated tooling since the 2007 launch are not displaced by six modules and a compiler, however well built. The relevant question is marginal rather than absolute: how much of the next wave of AI software, especially software written outside the United States, gets written in a language that treats CUDA as one backend among four.

Why It Matters

For Chinese AI companies, the release is a hedge against a supply chain they do not control. Export controls have kept the most advanced Nvidia processors out of reach, and a software stack that treats Ascend as a first class target reduces the cost of building on domestic silicon. Huawei has said it sees its Ascend share of the Chinese AI chip market above Nvidia's, a claim Eric Xu made without figures, but the direction of travel is clear: the binding constraint has shifted from chips to the software and tooling that make chips usable.

For Nvidia, the risk is not a single quarter of lost orders. It is the erosion of a default. TileLang is explicitly positioned as a simpler programming model, and simplicity is how ecosystems change hands over a decade. If a meaningful slice of new kernels are written in a language that compiles to Ascend, ROCm and Metal as easily as to CUDA, Nvidia's lock in weakens even in markets where its chips keep selling.

For the wider market, the release tests whether open source can do for accelerator software what it did for operating systems and databases, commoditizing the layer that sits just above the hardware. The tools are free to download, the code is public, and the backend list is plural by design. That combination is unusual enough that its success will be judged less by DeepSeek's own deployment than by whether outside developers adopt it.

Next Up

DeepSeek and Huawei have said they plan to keep developing Ascend tooling together, and the next visible milestone is Huawei's Ascend 960DT, slated for the first quarter of 2027 with a roadmap stretching to 2029. DeepSeek's V4 model and its successor work will be the first serious proving ground, since every new training run and every long context workload puts the new libraries under load.

The harder test is adoption. Watch whether developers outside China pick up TileLang for its multi backend portability, whether other accelerator vendors treat it as a supported target, and how Nvidia answers at its next developer gathering. Six modules and a supernode of 128 Ascend 950 chips will not settle a platform war that began in 2007, but they change what a developer can assume about the cost of leaving it.

Tagged

Comments (0)

No comments yet. Be the first to share your thoughts.