# deepseek-ai/DeepGEMM-Ascend > A port of DeepGEMM to Huawei Ascend: GEMM kernels (BF16, FP8, FP4, MQA logits, MegaMoE) with an API compatible with DeepGEMM, targeting peak NPU performance. - Magnitude: 5.6 out of 10 — Early signal - Stars: 127 total · +0 stars today, ≈ 85 by evening - Star trust: star growth looks organic - Category: Language models · Language: C++ · License: MIT · Created: 2026-09-29 · Last push: 2026-09-30 - GitHub: https://github.com/deepseek-ai/DeepGEMM-Ascend · Page: https://gitnova.dev/en/r/deepseek-ai/DeepGEMM-Ascend ## Useful for - Build and install the package to run GEMM kernels on Ascend 950 via torch_npu - Reuse the DeepGEMM API unchanged when porting code from NVIDIA to Ascend - Cross-check custom kernel performance against the library's ACLNN reference GEMMs ## Why it’s here - 0 stars so far today, about 85 expected by the end of the day. - The spike has held for 2 days in a row — not a one-off blip. - The repository is 1 day old and already has 127 stars. With less than two weeks of history, there's no usual pace to compare the spike against yet. - Top new repositories this week: #98. - About 49 forks a day — people are taking the code. ## Star trust Star growth looks organic. Star-trust labels are heuristics based on the repository’s behavior, not a check of every stargazer. ## Numbers - Forks: 5 - Issues and pull requests: 0 - Watchers: 0 - Average over the last week: 108 per day - Usual pace: too little history (under two weeks) - Stars in the last hour (measured): 23 ## Stars per day, last 4 days (oldest → newest, today is partial) 2026-09-27 … 2026-09-30: 0, 0, 131, 0 ## Spotted in now - Top new repositories this week: #98 ## More in this category 1. **firelex/jeff** — 7.7 · Early signal · Language models · Python · +0 stars today, ≈ 299 by evening Small fine-tuned Qwen3.5 and Gemma 4 models for zero-shot classification: given a situation description and a list of options, they return a calibrated probability for each option in a single forward pass. Full card: https://gitnova.dev/en/r/firelex/jeff.md 2. **VectifyAI/PageIndex** — 7.7 · Breakout · Language models · Python · +0 stars today, ≈ 664 by evening A vectorless RAG engine that builds a hierarchical tree index of a document and retrieves by LLM reasoning over that tree, the way a human navigates a report. Aimed at long professional PDFs such as financial, legal and technical documents. Full card: https://gitnova.dev/en/r/VectifyAI/PageIndex.md 3. **PostHog/jeeves** — 7.4 · Early signal · Language models · Python · +0 stars today, ≈ 193 by evening Jeeves is a reasoning model built on Qwen3.5-9B (LoRA + pointer head), trained with SFT and CISPO, that thinks before answering and returns calibrated probabilities for yes/no, multiple-choice and rating questions via a Jev-compatible API. Full card: https://gitnova.dev/en/r/PostHog/jeeves.md 4. **Niko1221/Strata** — 7.2 · Breakout · Language models · C++ · +0 stars today, ≈ 454 by evening A C++ local inference engine that runs the 125B MoE model Qwen3.8-Flash-Next on a regular PC with one NVIDIA GPU (12-24 GB) and 64 GB RAM, exposing an OpenAI/Anthropic-compatible API on localhost. Full card: https://gitnova.dev/en/r/Niko1221/Strata.md 5. **ninjahawk/livenerf** — 5.7 · Early signal · Language models · Python · +0 stars today, ≈ 147 by evening A benchmark for tracking whether a frontier model quietly degrades after release: it runs a frozen question panel daily and statistically measures drift in accuracy and token counts against the launch-week baseline. Full card: https://gitnova.dev/en/r/ninjahawk/livenerf.md --- Magnitude (0–10) measures how fast and how unusually interest in a repository is growing right now. It is not a quality score. Days are UTC. “So far today” is a fact; “expected by the end of the day” is a forecast. Summaries and use cases are written by an LLM (DeepSeek V4.1 Flash) from the README and may be inaccurate: verify specific claims (benchmarks, speed, hardware) in the repository itself. Data as of 2026-09-30 04:42 UTC, updated every 30 minutes.