Ascend is a series of AI accelerators (NPUs) developed by Huawei. Huawei says it is evolving Ascend on a one-generation-per-year cycle.1

Ascend 970 and 980

Ascend 970 and 980 are planned to be released in 2028 and 2029, respectively.21

Huawei’s original 2025 roadmap placed the Ascend 970 in Q4 2028 with these targets relative to the Ascend 960:3

  • 4 PF FP8, 8 PF FP4 (2x the 960)
  • 4 TB/s interconnect (2x the 960)
  • At least 1.5x the 960’s memory bandwidth

Ascend 960

Ascend 960 comes in two variants:1

  • 960DT: available Q1 2027, three quarters ahead of the original roadmap
  • 960PR: available Q3 2027, one quarter ahead of schedule

The original 2025 roadmap listed a single Ascend 960 for Q4 2027.3 Compared to the Ascend 950, it would have:3

  • 2x compute, memory bandwidth, memory capacity, and interconnect ports
  • 2 PF FP8, 4 PF FP4
  • Support for HiF4, Huawei’s proprietary 4-bit format

The Register reportedly puts the 960DT at up to 288 GB of memory and 4 PF FP4,4 but I don’t know where this figure came from.

Ascend 960PR is designed to scale up to 4,096 processors and scale out to over 1 million.2

Atlas 960E SuperPoD

Atlas 960E is a scale-up system built on Ascend 960 and Huawei’s Hi-ONE near-packaged optics (NPO) engine.1 Each SuperPoD has:1

  • Up to 4,096 NPUs
  • 8 EF FP8
  • 16 EF FP4
  • Up to 1 PB of HBM
  • 5,500 Hi-ONE units at 7.2 Tbit/s per engine instead of 48,000x 800G optical modules

Multiple SuperPoDs connect into a SuperCluster of up to 512,000 NPUs using a two-tier, four-plane Clos, or up to one million NPUs when combined with a multi-rail topology.1

Ascend 950

“Ascend 950” is a die shared by two different chips, the 950PR and 950DT. Each pairs the die with a different proprietary HBM (HiBL 1.0 and HiZQ 2.0) packaged separately with it.3 Each chip supports:3

  • 2x AI dies
  • 2x I/O dies
  • 1 PF FP8, MXFP8, and HiF8
  • 2 PF MXFP4
  • Combined SIMD and SIMT design
  • Memory access granularity of 128 bytes, down from 512 bytes
  • 2 TB/s interconnect bandwidth

This is implemented as:5

  • 36 AI subsystems (third-generation DaVinci), each with 1 Cube core and 2 Vector cores
  • 128 MB L2 cache shared across both AI dies (UMA, hardware-coherent)
  • 4 AI CPU clusters, each with 2 Linx816 cores (ARMv8-A, 2 threads per core) and 4 MB L3
  • 72 lanes of 112 Gbps HiLink SerDes, organized as 18 x4 ports
    • Unified Bus 2.0 at 2,016 GB/s bidirectional across all 18 ports
    • PCIe 5.0 x16 (128 GB/s bidirectional), sharing 4 of the UB ports
    • 2x 400 Gbps UBoE (UnifiedBus over Ethernet), sharing 2 of the UB ports
  • Support for supernodes of up to 8,192 chips and clusters of over 128K chips

Cube cores support TF32, FP16, BF16, FP8, MXFP8, HiF8, INT8, and MXFP4.5 HiF8 and MXFP8 run at 2x the FP16 rate and MXFP4 at 4x.5

Ascend 950 has “entered commercial use” as of September 2026.21

Ascend 950PR

Ascend 950PR is the variant optimized for inference prefill and recommendation systems, using Huawei’s proprietary HiBL 1.0 HBM.3 It was scheduled for Q1 2026 in both computing cards and SuperPoD servers.3

The Ascend 950PR whitepaper provides the following two bins:5

Spec950PR (32-core)950PR (28-core)
Cube / Vector cores32 / 6428 / 56
MXFP41,784 TF1,561 TF
HiF8 / MXFP8 / FP8919 TF804 TF
INT8919 TOPS804 TOPS
BF16 / FP16486 TF425 TF
TF32243 TF212 TF
Vector FP3227 TF23 TF
Memory128 GB at 1.6 TB/s112 GB at 1.4 TB/s
L2 cache128 MB112 MB

These are combined Cube + Vector performance numbers.

Atlas 350

Atlas 350 is a PCIe accelerator card built on the 28-core 950PR bin.6 It was first announced at Huawei Connect 20257 and went on sale on March 20, 2026. It has:6

  • 1,561 TF MXFP4
  • 804 TF MXFP8 / HiF8 / MXFP6 / INT8
  • 425 TF FP16 / BF16
  • 112 GB HBM at 1.4 TB/s, with ECC
  • PCIe 5.0 x16 (128 GB/s bidirectional)
  • UB card-to-card links through a dedicated connector
    • 4-card full mesh: 3 x4 UB ports per card, 318 GB/s bidirectional
    • 2-card: 4 x4 UB ports per card, 424 GB/s bidirectional
  • under 600 W, passively cooled, full-height full-length dual-width

Ascend 950DT

Huawei lists three compute bins (36, 32, or 28 Cube cores) and two memory configurations (144 GB or 96 GB), all at 4 TB/s.5 The top bin delivers:5

  • 2,007 TF MXFP4
  • 1,034 TF HiF8 / MXFP8 / FP8
  • 547 TF BF16 / FP16
  • 144 GB at 4 TB/s

Its original availability date was Q4 2026.3

Atlas 950 SuperPoD

The Atlas 950 SuperPoD is built with Ascend 950DT and was scheduled for Q4 2026. It has:3

  • Up to 8,192 Ascend 950DT NPUs
  • 160 cabinets (128 compute and 32 communications)
  • 8 EF FP8, 16 EF FP4
  • 1,152 TB memory
  • 16 PB/s interconnect bandwidth

The Atlas 950 SuperCluster combines 64 Atlas 950 SuperPoDs into over 520,000 Ascend 950DT delivering 524 EF FP8, with support for both UBoE (UnifiedBus over Ethernet) and RoCE.3

Ascend 910C

Ascend 910C is an upgrade of the 910 whose compute chiplets are manufactured by SMIC on its 2nd generation 7nm-class process, “N+2.”8 The SoC has around 53 billion transistors.8

As of September 2026, Ascend 910C is being evaluated by Malaysia for its sovereign AI program.9 However, Huawei announced they would not “expand into the international market in a fully fledged way” due to high domestic demand.2

Huawei has claimed that Ascend 910C has been deployed to more than a thousand clusters.2

Atlas 900 A3 SuperPoD (CloudMatrix 384)

The Atlas 900 A3 SuperPoD launched in March 2025 with up to 384 Ascend 910C and up to 300 PF, interconnected using UnifiedBus 1.0.3 Huawei describes it as working “like a single computer.”3 CloudMatrix 384 is the Huawei Cloud instance built on top of the Atlas 900 A3.3

It is intended to compete against GB200 NVL72 and costs $8.2 million.10

Huawei had deployed over 300 Atlas 900 A3 SuperPoDs as of September 20253 and over 1,000 as of September 2026.1

Ascend 910

Ascend 910 is the original version of the accelerator, launched in 2019.3 It was built from chiplets fabbed by TSMC on the N7+ process (“7nm-class with EUV”).8

Footnotes

  1. Advancing the Agentic World, Building a Solid Silicon Foundation | Huawei. David Wang’s keynote at Huawei Connect 2026, September 17, 2026. ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8

  2. China’s Huawei says AI chip demand outstrips supply as it steps up Nvidia challenge | Reuters ↩ ↩2 ↩3 ↩4 ↩5

  3. Groundbreaking SuperPoD Interconnect: Leading a New Paradigm for AI Infrastructure | Huawei. Eric Xu’s keynote at Huawei Connect 2025, September 18, 2025. ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8 ↩9 ↩10 ↩11 ↩12 ↩13 ↩14 ↩15

  4. Huawei’s next-gen Ascend NPUs could become China’s best option ↩

  5. 昇腾 950 NPU 架构白皮书 ↩ ↩2 ↩3 ↩4 ↩5 ↩6

  6. Atlas 350 accelerator card ↩ ↩2

  7. Huawei’s innovative supernode architecture, built through open source and collaborative efforts, forms the foundation for computing power across all scenarios. ↩

  8. DeepSeek research suggests Huawei’s Ascend 910C delivers 60% Nvidia H100 inference performance ↩ ↩2 ↩3

  9. Malaysia weighs Huawei chips for sovereign AI push | FMT ↩

  10. Huawei delivers advanced AI chip ‘cluster’ to Chinese clients cut off from Nvidia ↩