NEWS

DeepSeek Releases Six Tools Letting Huawei Chips Replace Nvidia

The DeepSeek whale logo on a smartphone screen in front of a larger blurred DeepSeek logo, as DeepSeek open-sources six components for Huawei Ascend chips
DeepSeek published six open-source components on September 30, 2026 that run its models on Huawei Ascend 950 accelerators. Source: Santage
Quick answer: DeepSeek released six open-source software components on September 30, 2026 that let its models train and run on Huawei Ascend 950 accelerators instead of Nvidia GPUs. The release covers the kernel compiler, matrix multiplication, cross-chip communication, attention kernels, vector operators and data selection, and is published under the MIT licence.
TLDR

DeepSeek published six open-source software components on September 30 that run its models on Huawei's Ascend 950 accelerators, rebuilding on Chinese silicon the libraries that have tied the company's training and inference to Nvidia hardware.

Six components rebuild the CUDA libraries on Ascend silicon

The release maps almost one to one onto the Nvidia stack DeepSeek has published over the past two years. TileLang, the kernel language DeepSeek uses to write high-performance operators, now emits Ascend C instructions directly. DeepGEMM-Ascend handles matrix multiplication across BF16, FP8 and FP4, with grouped variants for mixture-of-experts routing. DeepEP-Ascend carries the all-to-all traffic that expert-parallel models generate between chips. FlashMLA supplies the sparse attention kernels, TileKernels the vector and memory operators, and DeepSelect the data filtering.

Diagram mapping five layers of the Nvidia CUDA software stack to the six DeepSeek components released for Huawei Ascend 950 chips, covering kernel language, matrix multiplication, cross-chip communication, attention kernels and vector operators
Each layer of DeepSeek's Nvidia stack now has an Ascend counterpart. Chart: Santage. Source: DeepSeek open-source repositories on GitHub, September 30, 2026.

The performance figures in the repositories are the substance of the claim. DeepGEMM-Ascend documents dense matrix multiplication reaching up to 99.8% of the hardware limit across data types, validated on Ascend 950 series parts. The communication library reports similar headroom.

“Approximately 90 to 95% of the physical payload bandwidth limit for EP sizes up to 32.”

DeepEP-Ascend performance notes, DeepSeek repository, September 30, 2026
What DeepSeek released
ComponentsSix, published under the MIT licence
Matrix multiplicationUp to 99.8% of hardware limit on BF16 dense, 98.3% on FP4
Dispatch bandwidth373 to 375 GB/s at 8 ranks, 313 to 320 GB/s at 128 ranks
Validated onAscend 950 series, CANN 9.20 toolkit
Commercial releaseMid-October 2026, through Huawei channels
Source: DeepGEMM-Ascend and DeepEP-Ascend repositories, September 30, 2026.

Nvidia's firmest hold on AI developers runs through software

Export controls have restricted which chips reach Chinese labs for three years, and Huawei has answered with silicon. The Ascend 950DT carries 144 GB of Huawei's own HiZQ 2.0 memory at 4 TB/s, delivers 1 PFLOPS in FP8 and 2 PFLOPS in MXFP4, and scales to 8,192 chips in the Atlas 950 SuperPoD, according to Huawei. Silicon alone left the harder problem untouched. Every kernel, every collective operation and every attention implementation that labs rely on was written against CUDA, and rewriting them for a new architecture costs months of specialist engineering that most teams will not spend.

DeepSeek has now spent it and given the result away. The FlashMLA release supports Ascend 950 alongside Nvidia's Blackwell-class parts and drops Hopper support entirely, a sequencing choice that says which platform the company expects to matter. Any Chinese lab building on open weights inherits a validated path onto domestic accelerators without doing the porting work itself.

Quick quiz
What had kept Chinese AI labs tied to Nvidia even as domestic chips improved?

The gap that matters is no longer the one measured in teraflops. Huawei has had credible accelerators for a year, and what kept them idle in training clusters was the absence of the kernels, collectives and compiler tooling that make a chip usable at scale. DeepSeek closed that gap in public, under a permissive licence, four weeks before Huawei's commercial toolkit arrives.

In short: DeepSeek open-sourced six components on September 30, 2026 that port its entire model infrastructure from Nvidia CUDA to Huawei Ascend 950 chips, reporting up to 99.8% of hardware limit on matrix multiplication and 375 GB/s on expert dispatch. The release removes the software dependency, rather than the hardware one, that has kept Chinese AI labs on American accelerators.

Santage is committed to independent, transparent journalism. This article is produced in accordance with Santage's Editorial Standards and aims to provide accurate and timely information. Readers are encouraged to verify information independently.