# tile-ai/tilelang > A Pythonic domain-specific language built on TVM for writing high-performance GPU/CPU/NPU kernels (GEMM, FlashAttention, etc.) with low-level optimizations. - Magnitude: 6.1 out of 10 — Breakout - Stars: 7,989 total · +57 stars today, ≈ 122 by evening - Star trust: star growth looks organic - Category: Developer tools · Language: Python · Created: 2024-10-03 · Last push: 2026-09-30 - GitHub: https://github.com/tile-ai/tilelang · Homepage: https://tilelang.com/ · Page: https://gitnova.dev/en/r/tile-ai/tilelang ## Useful for - Write a GEMM kernel in TileLang for GPU tensor cores - Port a FlashAttention-like kernel to the Ascend 950 NPU - Debug compiler lowering with IR Lower Trace and Pass Visualizer ## Why it’s here - 57 stars so far today, about 122 expected by the end of the day. The usual pace is 8 per day, so that's 16× as much. - Over the last two days the pace is 28× that of the previous week and a half. - The spike has held for 3 days in a row — not a one-off blip. - GitHub Trending today: #11, +157 stars. - GitHub Trending Python today: #1, +157 stars. - About 17 forks a day — people are taking the code. - Recent forks include notable developers: @ochafik (434 followers). ## Star trust Star growth looks organic. Star-trust labels are heuristics based on the repository’s behavior, not a check of every stargazer. ## Numbers - Forks: 806 - Issues and pull requests: 3,325 - Watchers: 43 - Average over the last week: 84 per day - Usual pace: 8 per day - Stars in the last hour (measured): 12 - Latest release: v0.1.15 (2026-09-30) ## Stars per day, last 30 days (oldest → newest, today is partial) 2026-09-02 … 2026-10-01: 21, 6, 6, 3, 6, 7, 5, 13, 5, 5, 3, 4, 3, 8, 4, 11, 6, 5, 8, 15, 7, 10, 4, 1, 4, 5, 4, 219, 230, 57 ## Spotted in now - GitHub Trending today: #11, +157 stars - GitHub Trending Python today: #1, +157 stars - GitHub Trending this week: #13, +480 stars ## Similar by description 1. **openxla/xla** — 0.4 · Quiet · Language models · C++ · +0 stars today, ≈ 0 by evening XLA (Accelerated Linear Algebra) is an open-source ML compiler for GPUs, CPUs, and ML accelerators. It takes models from PyTorch, TensorFlow, and JAX and optimizes them for high-performance execution across different hardware. Full card: https://gitnova.dev/en/r/openxla/xla.md 2. **NVIDIA/cutlass** — 1.2 · Steady · Language models · C++ · +1 star today, ≈ 2 by evening CUTLASS is a collection of CUDA C++ templates and Python DSLs (CuTe DSL) for implementing high-performance linear algebra, primarily GEMM, on NVIDIA GPUs. It targets researchers and performance engineers writing optimized Tensor Core… Full card: https://gitnova.dev/en/r/NVIDIA/cutlass.md --- Magnitude (0–10) measures how fast and how unusually interest in a repository is growing right now. It is not a quality score. Days are UTC. “So far today” is a fact; “expected by the end of the day” is a forecast. Summaries and use cases are written by an LLM (DeepSeek V4.1 Flash) from the README and may be inaccurate: verify specific claims (benchmarks, speed, hardware) in the repository itself. Data as of 2026-10-01 13:52 UTC, updated every 30 minutes.