# ashhart/TensorFold > LLM inference server for Apple Silicon (MLX) and NVIDIA GPUs with an OpenAI-compatible API and exact speculative decoding. Supports Nemotron, Qwen3.8, GLM-5.3 and Gemma 4 families with per-family kernels and draft models. - Magnitude: 5.6 out of 10 — Early signal - Stars: 534 total · +10 stars today, ≈ 58 by evening - Star trust: star growth looks organic - Category: Language models · Language: Python · License: MIT · Created: 2026-06-19 · Last push: 2026-09-28 - GitHub: https://github.com/ashhart/TensorFold · Page: https://gitnova.dev/en/r/ashhart/TensorFold ## Useful for - Serve a local OpenAI-compatible endpoint on a Mac for 4-bit Qwen3.8-27B - Pre-download a checkpoint and draft model with tensorfold pull before starting the server - Verify drafted output matches serial output by comparing with draft: false ## Why it’s here - 10 stars so far today, about 58 expected by the end of the day. The usual pace is 0 per day, so that's 58× as much. - Before this, the repository barely got any stars — about 2 per day. - The spike has held for 3 days in a row — not a one-off blip. - GitHub Trending Python today: #4, +160 stars. - About 16 forks a day — people are taking the code. ## Star trust Star growth looks organic. Star-trust labels are heuristics based on the repository’s behavior, not a check of every stargazer. ## Numbers - Forks: 53 - Issues and pull requests: 67 - Watchers: 9 - Average over the last week: 78 per day - Usual pace: 0 per day - Stars in the last hour (measured): 4 - Latest release: v0.3.6.1 (2026-09-28) ## Stars per day, last 30 days (oldest → newest, today is partial) 2026-08-30 … 2026-09-28: 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 26, 252, 207, 10 ## Spotted in now - GitHub Trending Python today: #4, +160 stars ## Similar by description 1. **MiaAI-Lab/DeepSeek-v4.1-Flash-EXL3-2x-DGX-Sparks** — 0.5 · Steady · Language models · Python · +0 stars today, ≈ 0 by evening A local EXL3 2.9 bpw checkpoint of DeepSeek-V4.1-Flash (196 GiB, 39 shards) served via an OpenAI-compatible vLLM API on a 2x NVIDIA GB10 (DGX Spark) kit with tensor-parallel 2. Includes DSpark speculative decoding and memory tuning for… Full card: https://gitnova.dev/en/r/MiaAI-Lab/DeepSeek-v4.1-Flash-EXL3-2x-DGX-Sparks.md 2. **NVIDIA/TensorRT-LLM** — 1.6 · Steady · Language models · Python · +5 stars today, ≈ 9 by evening NVIDIA library for optimizing inference of large language models and visual generative models on GPUs, with a Python API, specialized kernels, and an efficient C++/Python runtime. Full card: https://gitnova.dev/en/r/NVIDIA/TensorRT-LLM.md 3. **NVIDIA/Model-Optimizer** — 4.3 · Peaking · Language models · Python · +17 stars today, ≈ 55 by evening NVIDIA library for compressing and accelerating models: quantization, pruning, distillation, NAS and speculative decoding with export to TensorRT-LLM, vLLM, SGLang. For ML engineers preparing models for deployment. Full card: https://gitnova.dev/en/r/NVIDIA/Model-Optimizer.md 4. **incoai/splash** — 2.5 · Cooling · Language models · C++ · +14 stars today, ≈ 32 by evening A local LLM inference engine for Apple silicon, specialized per model: serves OpenAI- and Anthropic-compatible APIs to coding agents on a single Mac. Full card: https://gitnova.dev/en/r/incoai/splash.md --- Magnitude (0–10) measures how fast and how unusually interest in a repository is growing right now. It is not a quality score. Days are UTC. “So far today” is a fact; “expected by the end of the day” is a forecast. Summaries and use cases are written by an LLM (DeepSeek V4.1 Flash) from the README and may be inaccurate: verify specific claims (benchmarks, speed, hardware) in the repository itself. Data as of 2026-09-28 13:02 UTC, updated every 30 minutes.