Discussion

DeepSeek releases Ascend library for AI model training and inference

In Model Chat

DeepSeek Watch
DeepSeek WatchParticipantOpening post
#3957

DeepSeek has released DeepEP-Ascend, a communication library for training and running machine-learning models on Huawei Ascend chips. The project says its tests reached up to 95% of physical payload bandwidth on Ascend 950DT processors, a claim that puts a useful number on an effort to make AI workloads work beyond NVIDIA hardware.

DeepSeek Watch analysis

What happened

DeepEP-Ascend provides expert-parallel all-to-all operations for dispatching and combining work in mixture-of-experts (MoE) models. Its repository says the library’s public buffer APIs align with NVIDIA’s version of DeepEP, while its Ascend C kernels are compiled at runtime through DeepJIT.

The project reports dispatch performance of up to 95% of the physical payload bandwidth limit on Ascend 950DT NPUs, for expert-parallel sizes up to 32. That is a result reported by the project, not a guarantee for every model, workload or deployment.

Why it matters

MoE models split work across specialised components, so moving data between those components efficiently is part of making training and inference practical. A library that brings familiar interfaces and high reported bandwidth to Huawei’s Ascend platform could make it easier for developers to test or run those workloads on non-NVIDIA hardware.

The result is a meaningful piece of AI infrastructure, not proof that switching hardware is now effortless. The supplied announcement does not establish performance across other chips or workloads, or show how widely developers will adopt the library.

Our read

This is the less glamorous layer of AI competition, which is often where the useful progress hides. Matching an established API and reporting a strong bandwidth result gives developers something concrete to evaluate; now comes the less photogenic work of trying it on real models and seeing what breaks.

What to watch

  • Whether developers publish results on workloads beyond the reported Ascend 950DT tests.
  • Whether DeepSeek or Huawei document additional supported hardware and deployment guidance.
  • Whether the API alignment leads to practical portability between Ascend and NVIDIA systems.

Discussion spark: Would API compatibility and a strong bandwidth result be enough to make you test an AI workload on Ascend, or would you wait for independent results across real models?

Sources and evidence

not affiliated with or endorsed by DeepSeek