# NVIDIA/cutlass > CUTLASS is a collection of CUDA C++ templates and Python DSLs (CuTe DSL) for implementing high-performance linear algebra, primarily GEMM, on NVIDIA GPUs. It targets researchers and performance engineers writing optimized Tensor Core kernels. - Magnitude: 1.5 out of 10 — Steady - Stars: 10,474 total · +6 stars today, ≈ 9 by evening - Star trust: star growth looks organic - Category: Language models · Language: C++ · Created: 2017-11-30 · Last push: 2026-09-23 - GitHub: https://github.com/NVIDIA/cutlass · Homepage: https://docs.nvidia.com/cutlass/index.html · Page: https://gitnova.dev/en/r/NVIDIA/cutlass ## Useful for - Write a custom GEMM kernel in CUDA C++ with specific tiling and data types - Prototype a kernel in CuTe DSL using Python without deep C++ expertise - Run mixed-precision FP8/FP4 on Tensor Cores of Hopper or Blackwell architectures ## Why it’s here - 6 stars so far today, about 9 expected by the end of the day. - GitHub Trending C++ today: #15, +6 stars. - Recent forks include notable developers: @FindHao (325 followers). ## Star trust Star growth looks organic. Star-trust labels are heuristics based on the repository’s behavior, not a check of every stargazer. ## Numbers - Forks: 2,098 - Issues and pull requests: 3,314 - Watchers: 121 - Average over the last week: 6 per day - Usual pace: 6 per day - Stars in the last hour (measured): -1 - Latest release: v4.8.0 (2026-09-22) ## Stars per day, last 30 days (oldest → newest, today is partial) 2026-08-25 … 2026-09-23: 9, 19, 12, 6, 11, 4, 6, 3, 9, 6, 3, 6, 6, 14, 2, 5, 7, 5, 4, 7, 1, 11, 4, 9, 3, 7, 4, 4, 3, 6 ## Spotted in now - GitHub Trending C++ today: #15, +6 stars ## More in this category 1. **jaredpalmer/kev** — 9.0 · Breakout · Language models · Python · +464 stars today, ≈ 1,121 by evening kev is a LoRA adapter with a small readout head on top of Qwen2.5-0.5B that answers many typed questions about a document in a single forward pass, returning calibrated probabilities instead of text. Full card: https://gitnova.dev/en/r/jaredpalmer/kev.md 2. **mizorewww/laya-mlx** — 7.1 · Peaking · Language models · Python · +134 stars today, ≈ 464 by evening Native MLX runtime for Laya typed decision models: returns probabilities for choices, scores or truth values without text generation, running locally on Apple Silicon. Full card: https://gitnova.dev/en/r/mizorewww/laya-mlx.md 3. **volotat/mini-AGI** — 6.8 · Early signal · Language models · Python · +32 stars today, ≈ 108 by evening A byte-level continual-learning language model that assembles its own architecture, trains from scratch on a single 8 GB VRAM GPU, and keeps learning from everything it reads. An experiment showing continual learning without catastrophic… Full card: https://gitnova.dev/en/r/volotat/mini-AGI.md 4. **TheoLeeCJ/SemIf** — 6.4 · Peaking · Language models · Python · +68 stars today, ≈ 211 by evening An open reproduction of the Jev-style semantic decision interface: reads typed option probabilities directly from a 4B model's logits without generating text. Runs on a single RTX 3090. Full card: https://gitnova.dev/en/r/TheoLeeCJ/SemIf.md 5. **nokia-applied-research/AnyJev** — 6.2 · Early signal · Language models · Python · +122 stars today, ≈ 213 by evening Turns any LLM into a decision model with typed questions (choice, yes/no, score) that return real probabilities read from the next-token distribution, without training. L0 removes position and label-prior bias; L1 adds temperature… Full card: https://gitnova.dev/en/r/nokia-applied-research/AnyJev.md --- Magnitude (0–10) measures how fast and how unusually interest in a repository is growing right now. It is not a quality score. Days are UTC. “So far today” is a fact; “expected by the end of the day” is a forecast. Summaries and use cases are written by an LLM (DeepSeek V4.1 Flash) from the README and may be inaccurate: verify specific claims (benchmarks, speed, hardware) in the repository itself. Data as of 2026-09-23 12:53 UTC, updated every 30 minutes.