DeepSeek Opens Six Ascend Tools vs CUDA

Introduction
Competition showed up with receipts. On Sep 30, 2026, DeepSeek’s official WeChat said it was open-sourcing Ascend-platform infrastructure — a high-level language plus compute and communication libraries — with Huawei’s full support, including joint work on an Ascend 950 × 128-card supernode. RoundtableSpace framed it as six tools to “cut Nvidia out of the picture entirely.” The ship is real; the total-eclipse headline is spin. What actually landed is a mirrored open stack for Huawei’s NPU path — DeepSeek’s answer to “CUDA isn’t just silicon, it’s the software people already know.”
What DeepSeek Actually Opened
Chinese press quoting the WeChat post (and the live GitHub set) name six pieces that correspond one-to-one with DeepSeek’s earlier NVIDIA-oriented open tools:
ProjectRoleNotes
High-level kernel DSL
DeepSeek: simpler programming model than CUDA, still aiming at hardware ceiling; Ascend path wraps Ascend C
Matrix multiply (GEMM)
Ascend-dedicated repo; API aligned to NVIDIA DeepGEMM
Expert-parallel / cluster comms
Ascend-dedicated; MoE dispatch/combine style workloads
TileLang operator library
Cross-platform collection; Ascend support in this wave
Multi-head Latent Attention kernels
Flagship DeepSeek attention path; Ascend in the package
TopK / sparse select
Sparse-attention + sampling helpers; Ascend kernels called out
There’s also a dedicated Ascend adapter line in the TileLang ecosystem (tilelang-ascend). Reuters and SCMP independently attribute the announcement to DeepSeek’s WeChat and the Huawei partnership — this is not a screenshot-only leak.
Huawei, Ascend 950, and the CUDA Angle
DeepSeek’s own framing (via WeChat paraphrases): building a new independent, controllable GPU-class software ecosystem starts with a universal, easy-to-program high-level language that can still hit peak hardware. TileLang is that bet — first proven on mature NVIDIA gear (DeepSeek says most V4 training ops rode TileLang), now carrying Ascend backends so the same mental model travels.
Huawei didn’t just wave a flag: DeepSeek says the teams co-optimized compute and communication for the 128-card Ascend 950 supernode. Vendor benches claim compute/comm performance “near the hardware ceiling.” Treat those numbers as first-party until someone outside the loop reproduces them.
Here’s the honest competition read: Nvidia’s moat was never only transistors — it was CUDA + libraries + four million habits. Shipping an open Ascend stack that rhymes with DeepSeek’s NVIDIA open stack is exactly how you chip at that moat. Calling CUDA “gone” overnight is fanfiction. Calling this a welcome, concrete alternative ecosystem move is fair.
What Creators Should (and Shouldn’t) Expect
- If you train/serve on Ascend / CANN land: this is the interesting week — language, GEMM, EP, MLA, TopK in public repos.
- If you live on CUDA workstations and ArtRealmAI Gen: nothing flips in your UI tomorrow. This is infra politics + open weights of the stack, not a new image/video model drop.
- Perf claims: README/WeChat tables (e.g. DeepGEMM utilization, DeepEP bandwidth at EP8–EP128) are useful signals, not a SemiAnalysis bakeoff. Early DeepEP docs even flag hardware/firmware (HDK) caveats for full bandwidth.
Original Source
Primary confirmations: Reuters, SCMP, and the GitHub set above (start with DeepGEMM-Ascend + TileLang).
Conclusion
Six open Ascend tools, Huawei on the co-sign, TileLang pitched as the friendlier-than-CUDA kernel lane for Ascend 950 clusters — confirmed, same-day wire + live repos. Enjoy the competitive heat; just don’t confuse a real software-stack landing with Nvidia vanishing from the map.
—Aurelia ♡
