MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks-TensorFold
GLM-5.3-Flash EXL3 on 2x DGX Spark with TensorFold
About the project
Scripts to serve GLM-5.3-Flash on two NVIDIA DGX Sparks via TensorFold with an OpenAI-compatible API, 1M-token context, and image/video input.
Useful for
- Deploy a local OpenAI-compatible GLM-5.3-Flash endpoint on two DGX Sparks
- Serve requests with up to 1M-token context and image input via the API
- Benchmark prefill and decode throughput on your own DGX Spark pair
README summarized by DeepSeek V4.1 Flash. Details may be inaccurate.
Why it’s trending
- 98 stars so far today, about 131 expected by the end of the day.
- The repository is 1 day old and already has 91 stars. With less than two weeks of history, there's no usual pace to compare the spike against yet.
- Top new repositories this week: #161.
- About 71 forks a day — people are taking the code.
- Recent forks include notable developers: @d3vilbug (302 followers).
Stars per day
Bars are daily stars, the line is the usual pace. Red marks spike days.
Numbers
- Total stars
- 91
- Today
- 98 · ≈ 131 by evening
- Forks
- 7
- Issues and pull requests
- 7
- Watchers
- 3
- Language
- Shell
- License
- Apache-2.0
- Created
- September 30, 2026
- Last push
- October 1, 2026
Star trust
Growth looks organic: forks and discussion are in line with active projects, and stars arrive unevenly, the way people give them.
These are heuristics, not a verdict: we judge by the repository’s behavior, not by a list of stargazers.
Spotted in
- October 1, 2026Top new repositories this week: #161
Similar by description
-
2.9
jayleaton/glm53-tensorfold-spark
Experimental setup for serving the 4-bit GLM-5.3-Flash checkpoint across two NVIDIA DGX Sparks via the TensorFold engine, with an OpenAI-compatible API and a set of patches for faster decoding.
-
3.6
MiaAI-Lab/Qwen3.8-Flash-Next-Single-DGX-Spark-TensorFold
Scripts to serve Qwen3.8 Flash Next on a single NVIDIA DGX Spark via an OpenAI-compatible API, with image and video input and a 262,144-token context.
-
0.6
architectds/collabosm
A set of scripts to run Qwen3.8-Flash-Next (125B MoE) on a single Colab A100-80GB High-RAM with an OpenAI-compatible endpoint and measured throughput.
-
0.0
mmastrac/djev-spark
Container recipe for running DiffusionGemma 26B-A4B (NVFP4) on a DGX Spark with vLLM and a structured-decision server exposing Jev's /v1/systemone API.