Seismograph

What’s gaining stars on GitHub right now

deepseek-ai/DeepGEMM-Ascend

DeepGEMM-Ascend: clean and efficient matrix multiplication kernel library for Huawei Ascend NPUs

Language modelsC++
5.6 Early signal Magnitude out of 10 — how fast interest is growing, not a quality score. Star growth looks organic. Data as of September 30, 2026, 04:42 UTC.
Open on Seismograph Open on GitHub

About the project

A port of DeepGEMM to Huawei Ascend: GEMM kernels (BF16, FP8, FP4, MQA logits, MegaMoE) with an API compatible with DeepGEMM, targeting peak NPU performance.

Useful for

README summarized by DeepSeek V4.1 Flash. Details may be inaccurate.

Why it’s trending

Stars per day

075150September 27, 2026September 30, 2026

Bars are daily stars, the line is the usual pace. Red marks spike days.

Numbers

Total stars
127
Today
0 · ≈ 85 by evening
Forks
5
Issues and pull requests
0
Watchers
0
Language
C++
License
MIT
Created
September 29, 2026
Last push
September 30, 2026

Star trust

Growth looks organic: forks and discussion are in line with active projects, and stars arrive unevenly, the way people give them.

These are heuristics, not a verdict: we judge by the repository’s behavior, not by a list of stargazers.

Spotted in

Share

README badge

Paste it into your README — the badge shows the current magnitude and links to this page.

Seismograph: 5.6
[![Seismograph](https://gitnova.dev/badge/deepseek-ai/DeepGEMM-Ascend.svg?lang=en)](https://gitnova.dev/en/r/deepseek-ai/DeepGEMM-Ascend)

More in this category

  1. 7.7
    firelex/jeff

    Small fine-tuned Qwen3.5 and Gemma 4 models for zero-shot classification: given a situation description and a list of options, they return a calibrated probability for each option in a single forward pass.

    Early signalLanguage modelsPython+0 stars today, ≈ 299 by evening

  2. 7.7
    VectifyAI/PageIndex

    A vectorless RAG engine that builds a hierarchical tree index of a document and retrieves by LLM reasoning over that tree, the way a human navigates a report. Aimed at long professional PDFs such as financial, legal…

    BreakoutLanguage modelsPython+0 stars today, ≈ 664 by evening

  3. 7.4
    PostHog/jeeves

    Jeeves is a reasoning model built on Qwen3.5-9B (LoRA + pointer head), trained with SFT and CISPO, that thinks before answering and returns calibrated probabilities for yes/no, multiple-choice and rating questions via…

    Early signalLanguage modelsPython+0 stars today, ≈ 193 by evening

  4. 7.2
    Niko1221/Strata

    A C++ local inference engine that runs the 125B MoE model Qwen3.8-Flash-Next on a regular PC with one NVIDIA GPU (12-24 GB) and 64 GB RAM, exposing an OpenAI/Anthropic-compatible API on localhost.

    BreakoutLanguage modelsC+++0 stars today, ≈ 454 by evening

  5. 5.7
    ninjahawk/livenerf

    A benchmark for tracking whether a frontier model quietly degrades after release: it runs a frozen question panel daily and statistically measures drift in accuracy and token counts against the launch-week baseline.

    Early signalLanguage modelsPython+0 stars today, ≈ 147 by evening

  6. 5.4
    ashhart/TensorFold

    LLM inference server for Apple Silicon (MLX) and NVIDIA GPUs with an OpenAI-compatible API and exact speculative decoding. Supports Nemotron, Qwen3.8, GLM-5.3 and Gemma 4 families with per-family kernels and draft…

    Early signalLanguage modelsPython+0 stars today, ≈ 63 by evening