# Inferact/tpu-megakernels > A collection of fused megakernels for LLM inference on TPU, targeting Kimi K3 and Qwen3.8-27B with speculative decoding, reaching up to ~2x the decode throughput of a GB200 baseline at batch sizes 1 to 8. - Magnitude: 3.3 out of 10 — Early signal - Stars: 95 total · +0 stars today, ≈ 22 by evening - Star trust: star growth looks organic - Category: Language models · Language: Python · License: Apache-2.0 · Created: 2026-09-23 · Last push: 2026-09-25 - GitHub: https://github.com/Inferact/tpu-megakernels · Page: https://gitnova.dev/en/r/Inferact/tpu-megakernels ## Useful for - Start an OpenAI-compatible Qwen3.8-27B server on a single host with eight TPUs for chat inference - Run the CPU correctness tests on 32 virtual devices without TPU hardware - Benchmark Kimi K3 decode throughput with DSpark against the GB200 baseline ## Why it’s here - 0 stars so far today, about 22 expected by the end of the day. - The spike has held for 3 days in a row — not a one-off blip. - The repository is 2 days old and already has 95 stars. With less than two weeks of history, there's no usual pace to compare the spike against yet. - Top new repositories this week: #163. ## Star trust Star growth looks organic. Star-trust labels are heuristics based on the repository’s behavior, not a check of every stargazer. ## Numbers - Forks: 2 - Issues and pull requests: 1 - Watchers: 1 - Average over the last week: 39 per day - Usual pace: too little history (under two weeks) - Stars in the last hour (measured): 0 ## Stars per day, last 6 days (oldest → newest, today is partial) 2026-09-20 … 2026-09-25: 0, 0, 0, 68, 27, 0 ## Spotted in now - Top new repositories this week: #163 ## Similar by description 1. **MiaAI-Lab/DeepSeek-v4.1-Flash-EXL3-2x-DGX-Sparks** — 1.1 · Cooling · Language models · Python · +0 stars today, ≈ 3 by evening A local EXL3 2.9 bpw checkpoint of DeepSeek-V4.1-Flash (196 GiB, 39 shards) served via an OpenAI-compatible vLLM API on a 2x NVIDIA GB10 (DGX Spark) kit with tensor-parallel 2. Includes DSpark speculative decoding and memory tuning for… Full card: https://gitnova.dev/en/r/MiaAI-Lab/DeepSeek-v4.1-Flash-EXL3-2x-DGX-Sparks.md 2. **MiaAI-Lab/DeepSeek-v4.1-Flash-DGX-Sparks** — 2.1 · Steady · Language models · Python · +0 stars today, ≈ 7 by evening Scripts and patches to serve DeepSeek-V4.1-Flash with SGLang across a 3–4 node NVIDIA DGX Spark cluster, using MXFP4/FP8, speculative decoding and an OpenAI-compatible endpoint. Full card: https://gitnova.dev/en/r/MiaAI-Lab/DeepSeek-v4.1-Flash-DGX-Sparks.md --- Magnitude (0–10) measures how fast and how unusually interest in a repository is growing right now. It is not a quality score. Days are UTC. “So far today” is a fact; “expected by the end of the day” is a forecast. Summaries and use cases are written by an LLM (DeepSeek V4.1 Flash) from the README and may be inaccurate: verify specific claims (benchmarks, speed, hardware) in the repository itself. Data as of 2026-09-25 02:11 UTC, updated every 30 minutes.