deepseek-ai/DeepEP-Ascend
A high-performance communication library for machine learning training and inference on Huawei Ascend NPUs.
About the project
A high-performance communication library for ML training and inference on Huawei Ascend NPUs: expert-parallel all-to-all for MoE dispatch/combine, plus PP, CP/DP and Engram. Public buffer APIs are aligned with NVIDIA DeepEP.
Useful for
- Run MoE training with expert-parallel dispatch/combine on an Ascend NPU cluster
- Speed up MoE inference with FP8 dispatch and deferred epilogues
- Port code from NVIDIA DeepEP to Ascend without changing the buffer API
README summarized by DeepSeek V4.1 Flash. Details may be inaccurate.
Why it’s trending
- The repository is 0 days old and already has 104 stars. With less than two weeks of history, there's no usual pace to compare the spike against yet.
- Top new repositories this week: #141.
- About 45 forks a day — people are taking the code.
- Recent forks include notable developers: @gmh5225 (1,277 followers).
Stars per day
Bars are daily stars, the line is the usual pace. Red marks spike days.
Numbers
- Total stars
- 104
- Today
- 0 · ≈ 0 by evening
- Forks
- 5
- Issues and pull requests
- 0
- Watchers
- 0
- Language
- C++
- Created
- September 30, 2026
- Last push
- September 30, 2026
Star trust
Growth looks organic: forks and discussion are in line with active projects, and stars arrive unevenly, the way people give them.
These are heuristics, not a verdict: we judge by the repository’s behavior, not by a list of stargazers.
Spotted in
- September 30, 2026Top new repositories this week: #141
Similar by description
-
1.6
vllm-project/vllm-ascend
A hardware plugin for vLLM that runs LLM inference on Ascend NPUs (Atlas A2/A3). It enables deployment of Transformer, MoE, embedding and multimodal models on Huawei Ascend hardware.
-
5.8
deepseek-ai/DeepGEMM-Ascend
A port of DeepGEMM to Huawei Ascend: GEMM kernels (BF16, FP8, FP4, MQA logits, MegaMoE) with an API compatible with DeepGEMM, targeting peak NPU performance.
-
1.7
microsoft/onnxruntime
Cross-platform accelerator for ML inference and training. Runs models from PyTorch, TensorFlow, scikit-learn and others via the ONNX format with graph optimizations and hardware acceleration.