# deepseek-ai/DeepGEMM > CUDA tensor core kernel library: GEMM (FP8, FP4, BF16), fused MoE, MQA scoring for the indexer. Kernels compile at runtime via DeepJIT, no CUDA build at install time. - Magnitude: 5.7 out of 10 — Breakout - Stars: 8,576 total · +47 stars measured 2026-10-06, 11:46–14:09 UTC, ≈ 97 by evening - Star trust: star growth looks organic - Category: Language models · Language: Cuda · License: MIT · Created: 2025-02-13 · Last push: 2026-09-30 - GitHub: https://github.com/deepseek-ai/DeepGEMM · Page: https://gitnova.dev/en/r/deepseek-ai/DeepGEMM ## Useful for - Run FP8 GEMM on H800 GPU to speed up LLM inference - Build grouped GEMM for MoE layers with shared N and K - Use masked grouped GEMM during CUDA graph decoding ## Why it’s here - Star-counter measurements on 2026-10-06 (UTC), 11:46–14:09: 8529 → 8576 stars (+47). This is the change over that interval. - Estimated end-of-day forecast: about +97 stars, using observed gains and the previous day. - Over the last two days the pace is 13× that of the previous week and a half. - Elevated gains were measured on the last 2 completed UTC days in a row. Continued growth is not guaranteed. - GitHub Trending today: #9, +363 stars. - About 19 forks a day — people are taking the code. ## Star trust Star growth looks organic. Star-trust labels are heuristics based on the repository’s behavior, not a check of every stargazer. ## Numbers - Forks: 1,364 - Issues and pull requests: 464 - Watchers: 70 - Average over the last week: 103 per day - Usual pace: 4 per day - Stars in the last hour (measured): 27 - Latest release: nv_dev_f8e8fb5 (2026-07-20) ## Stars per day, last 30 days (oldest → newest, today is partial) 2026-09-07 … 2026-10-06: 6, 2, 7, 16, 7, 4, 2, 6, 4, 8, 6, 1, 4, 6, 2, 3, 6, 3, 2, 4, 6, 4, 11, 14, 42, 24, 9, 328, 204, 47 ## Spotted in now - GitHub Trending today: #9, +363 stars ## Similar by description 1. **deepseek-ai/DeepGEMM-Ascend** — 1.4 · Cooling · Language models · C++ · +4 stars measured 2026-10-06, 00:28–14:30 UTC, ≈ 7 by evening A port of DeepGEMM to Huawei Ascend: GEMM kernels (BF16, FP8, FP4, MQA logits, MegaMoE) with an API compatible with DeepGEMM, targeting peak NPU performance. Full card: https://gitnova.dev/en/r/deepseek-ai/DeepGEMM-Ascend.md --- Magnitude (0–10) measures how fast and how unusually interest in a repository is growing right now. It is not a quality score. Days are UTC. “So far today” is a fact; “expected by the end of the day” is a forecast. Summaries and use cases are written by an LLM (DeepSeek V4.1 Flash) from the README and may be inaccurate: verify specific claims (benchmarks, speed, hardware) in the repository itself. Data as of 2026-10-06 14:31 UTC, updated every 30 minutes.