sqliteai/warp
Run the full 2.78-trillion-parameter Kimi K3 model, DeepSeek V4.1 Flash or GLM-5.3-Flash beyond available RAM by streaming activated weights directly from NVMe. A dependency-free, embeddable C inference engine.
About the project
An embeddable dependency-free C inference engine that runs huge MoE models (Kimi K3, DeepSeek-V4.1-Flash, GLM-5.3-Flash) on consumer hardware by streaming activated experts from NVMe and using spare RAM as a bounded expert cache.
Useful for
- Run Kimi K3 on a 64 GB MacBook Pro with weights on SSD and experts cached in RAM
- Serve DeepSeek-V4.1-Flash or GLM-5.3-Flash locally without cloud on a low-memory machine
- Tune num_experts_per_token in the container manifest to trade quality for decode speed
README summarized by DeepSeek V4.1 Flash. Details may be inaccurate.
Why it’s trending
- 3 stars so far today, about 4 expected by the end of the day.
- The repository is 52 days old and already has 2,432 stars.
- Hacker News: “Show HN: Warp – Run DeepSeek v4.1 Flash with 5 GB of RAM at 3.77 tok/s” — 13 points, 3 days ago.
Stars per day
Bars are daily stars, the line is the usual pace. Red marks spike days.
Numbers
- Total stars
- 2,432
- Stars in a day
- 4
- Forks
- 181
- Issues and pull requests
- 74
- Watchers
- 9
- Language
- C
- License
- Apache-2.0
- Latest release
- v0.8.1 · September 18, 2026
- Created
- July 28, 2026
- Last push
- September 18, 2026
Star trust
Growth looks organic: forks and discussion are in line with active projects, and stars arrive unevenly, the way people give them.
These are heuristics, not a verdict: we judge by the repository’s behavior, not by a list of stargazers.
Hacker News discussions
- Show HN: Warp – Run DeepSeek v4.1 Flash with 5 GB of RAM at 3.77 tok/s13 points, September 15, 2026
Spotted in
- September 18, 2026Spotted on Hacker News
- September 17, 2026Spotted on Hacker News
- September 16, 2026Spotted on Hacker News
Similar projects
-
8.6
TheoLeeCJ/SemIf
An open reproduction of the Jev-style semantic decision interface: reads typed option probabilities directly from a 4B model's logits without generating text. Runs on a single RTX 3090.
-
7.1
vinnylarouge/jevlike
A small trainable model that picks one option from a changing list of text options in a single forward pass, an open alternative to TypeSafe's Jev. Aimed at developers who need fast classification or menu selection…
-
6.9
Continuum-AI-Corp/OrcaBonsai-27B-Uncensored
A tool for runtime refusal ablation in the compressed Ternary Bonsai 2 27B LLM: it projects the residual stream to remove the refusal direction without modifying or re-quantizing weights. Runs on Apple Silicon via MLX.
-
6.4
TianyuCodings/NanoJev
A nano replica of Jev built on Qwen3-0.6B that outputs probability distributions over candidates in a single forward pass with no token decoding, plus a training pipeline and game demos (maze, Snake).
-
6.3
zhengkid/Dream-RSI
Research project on recursive self-improvement: past discovery histories become a replay simulator used to evaluate and refine exploration policies. Code is still being prepared; paper and demo are available.
-
5.0
ekzhang/openjev-sglang
HTTP server implementing the TypeSafe/Jev API for structured text classification on Qwen3.6-35B-A3B via SGLang; returns answer probabilities without generating a chain of thought.