# 0xSero/DeepSeek-V4.1-Flash-Two-Sparks > Scripts to serve DeepSeek-V4.1-Flash on two NVIDIA DGX Spark nodes with vLLM (tensor parallel 2 over RoCE), using EXL3-quantized experts, and to connect it to the Pi coding agent. - Magnitude: 3.3 out of 10 — Early signal - Stars: 72 total · +3 stars measured 2026-10-05, 10:59–12:35 UTC, ≈ 12 by evening - Star trust: star growth looks organic - Category: Language models · Language: Shell · License: MIT · Created: 2026-10-03 · Last push: 2026-10-03 - GitHub: https://github.com/0xSero/DeepSeek-V4.1-Flash-Two-Sparks · Page: https://gitnova.dev/en/r/0xSero/DeepSeek-V4.1-Flash-Two-Sparks ## Useful for - Deploy a local OpenAI-compatible DeepSeek-V4.1-Flash endpoint on two DGX Sparks - Connect the Pi coding agent to your own model server via pi/install.sh - Run smoke tests and throughput benchmarks with scripts/smoke.py and bench.py ## Why it’s here - Star-counter measurements on 2026-10-05 (UTC), 10:59–12:35: 69 → 72 stars (+3). This is the change over that interval. - Estimated end-of-day forecast: about +12 stars, using observed gains and the previous day. - The repository is 2 days old and already has 72 stars. With less than two weeks of history, there's no usual pace to compare the spike against yet. - Recent forks include notable developers: @gmh5225 (1,283 followers). ## Star trust Star growth looks organic. Star-trust labels are heuristics based on the repository’s behavior, not a check of every stargazer. ## Numbers - Forks: 5 - Issues and pull requests: 2 - Watchers: 0 - Average over the last week: 26 per day - Usual pace: too little history (under two weeks) - Stars in the last hour (measured): 4 ## Stars per day, last 9 days (oldest → newest, today is partial) 2026-09-27 … 2026-10-05: 0, 0, 0, 0, 0, 0, 34, 32, 3 ## Similar by description 1. **MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks-TensorFold** — 2.5 · Cooling · Language models · Shell · +8 stars measured 2026-10-05, 00:08–12:37 UTC, ≈ 16 by evening Scripts to serve GLM-5.3-Flash on two NVIDIA DGX Sparks via TensorFold with an OpenAI-compatible API, 1M-token context, and image/video input. Full card: https://gitnova.dev/en/r/MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks-TensorFold.md 2. **jayleaton/glm53-tensorfold-spark** — 0.8 · Steady · Language models · Python · +0 stars measured 2026-10-05, 00:14–11:59 UTC, ≈ 1 by evening Experimental setup for serving the 4-bit GLM-5.3-Flash checkpoint across two NVIDIA DGX Sparks via the TensorFold engine, with an OpenAI-compatible API and a set of patches for faster decoding. Full card: https://gitnova.dev/en/r/jayleaton/glm53-tensorfold-spark.md --- Magnitude (0–10) measures how fast and how unusually interest in a repository is growing right now. It is not a quality score. Days are UTC. “So far today” is a fact; “expected by the end of the day” is a forecast. Summaries and use cases are written by an LLM (DeepSeek V4.1 Flash) from the README and may be inaccurate: verify specific claims (benchmarks, speed, hardware) in the repository itself. Data as of 2026-10-05 12:40 UTC, updated every 30 minutes.