ashhart/TensorFold
Fast, exact LLM decoding on Apple Silicon (MLX) behind an OpenAI-compatible endpoint
About the project
LLM inference server for Apple Silicon (MLX) and NVIDIA GPUs with an OpenAI-compatible API and exact speculative decoding. Supports Nemotron, Qwen3.8, GLM-5.3 and Gemma 4 families with per-family kernels and draft models.
Useful for
- Serve a local OpenAI-compatible endpoint on a Mac for 4-bit Qwen3.8-27B
- Pre-download a checkpoint and draft model with tensorfold pull before starting the server
- Verify drafted output matches serial output by comparing with draft: false
README summarized by DeepSeek V4.1 Flash. Details may be inaccurate.
Why it’s trending
- 10 stars so far today, about 65 expected by the end of the day. The usual pace is 0 per day, so that's 65× as much.
- Before this, the repository barely got any stars — about 2 per day.
- The spike has held for 3 days in a row — not a one-off blip.
- GitHub Trending Python today: #4, +160 stars.
- About 18 forks a day — people are taking the code.
Stars per day
Bars are daily stars, the line is the usual pace. Red marks spike days.
Numbers
- Total stars
- 534
- Today
- 10 · ≈ 65 by evening
- Forks
- 52
- Issues and pull requests
- 66
- Watchers
- 9
- Language
- Python
- License
- MIT
- Latest release
- v0.3.6.1 · September 28, 2026
- Created
- June 19, 2026
- Last push
- September 28, 2026
Star trust
Growth looks organic: forks and discussion are in line with active projects, and stars arrive unevenly, the way people give them.
These are heuristics, not a verdict: we judge by the repository’s behavior, not by a list of stargazers.
Spotted in
- September 28, 2026GitHub Trending Python today: #4, +160 stars
Similar by description
-
0.5
MiaAI-Lab/DeepSeek-v4.1-Flash-EXL3-2x-DGX-Sparks
A local EXL3 2.9 bpw checkpoint of DeepSeek-V4.1-Flash (196 GiB, 39 shards) served via an OpenAI-compatible vLLM API on a 2x NVIDIA GB10 (DGX Spark) kit with tensor-parallel 2. Includes DSpark speculative decoding and…
-
1.6
NVIDIA/TensorRT-LLM
NVIDIA library for optimizing inference of large language models and visual generative models on GPUs, with a Python API, specialized kernels, and an efficient C++/Python runtime.
-
4.3
NVIDIA/Model-Optimizer
NVIDIA library for compressing and accelerating models: quantization, pruning, distillation, NAS and speculative decoding with export to TensorRT-LLM, vLLM, SGLang. For ML engineers preparing models for deployment.
-
2.5
incoai/splash
A local LLM inference engine for Apple silicon, specialized per model: serves OpenAI- and Anthropic-compatible APIs to coding agents on a single Mac.