- Six components went live on DeepSeek's GitHub on September 30 under the MIT licence: an Ascend build of the TileLang compiler, DeepGEMM-Ascend, DeepEP-Ascend, TileKernels, FlashMLA and DeepSelect.
- DeepGEMM-Ascend reports reaching 99.8% of the hardware limit on BF16 dense matrix multiplication and 98.3% on FP4, measured on Ascend 950 parts running CANN 9.20.
- DeepEP-Ascend measures 373 to 375 GB/s dispatch bandwidth across eight ranks and holds 313 to 320 GB/s at 128 ranks, with a commercial release through Huawei channels due in mid-October.
DeepSeek published six open-source software components on September 30 that run its models on Huawei's Ascend 950 accelerators, rebuilding on Chinese silicon the libraries that have tied the company's training and inference to Nvidia hardware.
Six components rebuild the CUDA libraries on Ascend silicon
The release maps almost one to one onto the Nvidia stack DeepSeek has published over the past two years. TileLang, the kernel language DeepSeek uses to write high-performance operators, now emits Ascend C instructions directly. DeepGEMM-Ascend handles matrix multiplication across BF16, FP8 and FP4, with grouped variants for mixture-of-experts routing. DeepEP-Ascend carries the all-to-all traffic that expert-parallel models generate between chips. FlashMLA supplies the sparse attention kernels, TileKernels the vector and memory operators, and DeepSelect the data filtering.
The performance figures in the repositories are the substance of the claim. DeepGEMM-Ascend documents dense matrix multiplication reaching up to 99.8% of the hardware limit across data types, validated on Ascend 950 series parts. The communication library reports similar headroom.
“Approximately 90 to 95% of the physical payload bandwidth limit for EP sizes up to 32.”
DeepEP-Ascend performance notes, DeepSeek repository, September 30, 2026
| Components | Six, published under the MIT licence |
| Matrix multiplication | Up to 99.8% of hardware limit on BF16 dense, 98.3% on FP4 |
| Dispatch bandwidth | 373 to 375 GB/s at 8 ranks, 313 to 320 GB/s at 128 ranks |
| Validated on | Ascend 950 series, CANN 9.20 toolkit |
| Commercial release | Mid-October 2026, through Huawei channels |
Nvidia's firmest hold on AI developers runs through software
Export controls have restricted which chips reach Chinese labs for three years, and Huawei has answered with silicon. The Ascend 950DT carries 144 GB of Huawei's own HiZQ 2.0 memory at 4 TB/s, delivers 1 PFLOPS in FP8 and 2 PFLOPS in MXFP4, and scales to 8,192 chips in the Atlas 950 SuperPoD, according to Huawei. Silicon alone left the harder problem untouched. Every kernel, every collective operation and every attention implementation that labs rely on was written against CUDA, and rewriting them for a new architecture costs months of specialist engineering that most teams will not spend.
DeepSeek has now spent it and given the result away. The FlashMLA release supports Ascend 950 alongside Nvidia's Blackwell-class parts and drops Hopper support entirely, a sequencing choice that says which platform the company expects to matter. Any Chinese lab building on open weights inherits a validated path onto domestic accelerators without doing the porting work itself.
The gap that matters is no longer the one measured in teraflops. Huawei has had credible accelerators for a year, and what kept them idle in training clusters was the absence of the kernels, collectives and compiler tooling that make a chip usable at scale. DeepSeek closed that gap in public, under a permissive licence, four weeks before Huawei's commercial toolkit arrives.
Santage is committed to independent, transparent journalism. This article is produced in accordance with Santage's Editorial Standards and aims to provide accurate and timely information. Readers are encouraged to verify information independently.