tile-ai/tilelang
Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels
About the project
A Pythonic domain-specific language built on TVM for writing high-performance GPU/CPU/NPU kernels (GEMM, FlashAttention, etc.) with low-level optimizations.
Useful for
- Write a GEMM kernel in TileLang for GPU tensor cores
- Port a FlashAttention-like kernel to the Ascend 950 NPU
- Debug compiler lowering with IR Lower Trace and Pass Visualizer
README summarized by DeepSeek V4.1 Flash. Details may be inaccurate.
Why it’s trending
- 44 stars so far today, about 113 expected by the end of the day. The usual pace is 8 per day, so that's 15× as much.
- Over the last two days the pace is 27× that of the previous week and a half.
- The spike has held for 3 days in a row — not a one-off blip.
- GitHub Trending today: #11, +157 stars.
- GitHub Trending Python today: #1, +157 stars.
- About 17 forks a day — people are taking the code.
- Recent forks include notable developers: @ochafik (434 followers).
Stars per day
Bars are daily stars, the line is the usual pace. Red marks spike days.
Numbers
- Total stars
- 7,979
- Today
- 44 · ≈ 113 by evening
- Forks
- 805
- Issues and pull requests
- 3,324
- Watchers
- 43
- Language
- Python
- Latest release
- v0.1.15 · September 30, 2026
- Created
- October 3, 2024
- Last push
- September 30, 2026
Star trust
Growth looks organic: forks and discussion are in line with active projects, and stars arrive unevenly, the way people give them.
These are heuristics, not a verdict: we judge by the repository’s behavior, not by a list of stargazers.
Spotted in
- October 1, 2026GitHub Trending today: #11, +157 stars; GitHub Trending Python today: #1, +157 stars; GitHub Trending this week: #13, +480 stars
Similar by description
-
0.4
openxla/xla
XLA (Accelerated Linear Algebra) is an open-source ML compiler for GPUs, CPUs, and ML accelerators. It takes models from PyTorch, TensorFlow, and JAX and optimizes them for high-performance execution across different…
-
1.2
NVIDIA/cutlass
CUTLASS is a collection of CUDA C++ templates and Python DSLs (CuTe DSL) for implementing high-performance linear algebra, primarily GEMM, on NVIDIA GPUs. It targets researchers and performance engineers writing…