jayleaton/glm53-tensorfold-spark
GLM-5.3-Flash (abliterated EXL3) on 2x NVIDIA DGX Spark with the TensorFold engine: 1.8x faster decode than vLLM, 4x256k concurrent threads, byte-exact speculative decoding. Work in progress.
About the project
Experimental setup for serving the 4-bit GLM-5.3-Flash checkpoint across two NVIDIA DGX Sparks via the TensorFold engine, with an OpenAI-compatible API and a set of patches for faster decoding.
Useful for
- Stand up an OpenAI-compatible GLM-5.3-Flash server on a pair of DGX Sparks
- Benchmark TensorFold vs vLLM decode speed on your own prompts
- Enable tool-calling and structured-output patches via config/prod.env
README summarized by DeepSeek V4.1 Flash. Details may be inaccurate.
Why it’s trending
- 0 stars so far today, about 22 expected by the end of the day.
- The spike has held for 3 days in a row — not a one-off blip.
- The repository is 3 days old and already has 87 stars. With less than two weeks of history, there's no usual pace to compare the spike against yet.
- Top new repositories this week: #152.
Stars per day
Bars are daily stars, the line is the usual pace. Red marks spike days.
Numbers
- Total stars
- 87
- Today
- 0 · ≈ 22 by evening
- Forks
- 5
- Issues and pull requests
- 14
- Watchers
- 1
- Language
- Python
- License
- Apache-2.0
- Created
- September 28, 2026
- Last push
- October 1, 2026
Star trust
Growth looks organic: forks and discussion are in line with active projects, and stars arrive unevenly, the way people give them.
These are heuristics, not a verdict: we judge by the repository’s behavior, not by a list of stargazers.
Spotted in
- October 1, 2026Top new repositories this week: #152
Similar by description
-
5.3
ashhart/TensorFold
LLM inference server for Apple Silicon (MLX) and NVIDIA GPUs with an OpenAI-compatible API and exact speculative decoding. Supports Nemotron, Qwen3.8, GLM-5.3 and Gemma 4 families with per-family kernels and draft…
-
3.8
MiaAI-Lab/Qwen3.8-Flash-Next-Single-DGX-Spark-TensorFold
Scripts to serve Qwen3.8 Flash Next on a single NVIDIA DGX Spark via an OpenAI-compatible API, with image and video input and a 262,144-token context.