# antirez/ds4 > DwarfStar (ds4) is a native C inference engine for running a few specific large LLMs (DeepSeek V4 Flash/PRO, GLM 5.x, Qwen3.8 Flash Next) on Metal, CUDA and ROCm. Targets consumer hardware like MacBooks, DGX Spark and Strix Halo, with SSD streaming and multi-GPU support. - Magnitude: 5.1 out of 10 — Breakout - Stars: 23,297 total · +84 stars today, ≈ 156 by evening - Star trust: star growth looks organic - Category: Language models · Language: C · License: MIT · Created: 2026-05-06 · Last push: 2026-09-20 - GitHub: https://github.com/antirez/ds4 · Page: https://gitnova.dev/en/r/antirez/ds4 ## Useful for - Run DeepSeek V4 Flash locally on a 96+ GB MacBook via Metal - Set up a multi-user LLM server on several L40S cards via CUDA - Build a two-Mac RDMA cluster with tensor parallelism for 4-bit models ## Why it’s here - 84 stars so far today, about 156 expected by the end of the day. The usual pace is 31 per day, so that's 5.1× as much. - Over the last two days the pace is 6.6× that of the previous week and a half. - The spike has held for 3 days in a row — not a one-off blip. - GitHub Trending today: #15, +211 stars. - About 23 forks a day — people are taking the code. ## Star trust Star growth looks organic. Star-trust labels are heuristics based on the repository’s behavior, not a check of every stargazer. ## Numbers - Forks: 2,239 - Issues and pull requests: 1,173 - Watchers: 176 - Average over the last week: 92 per day - Usual pace: 31 per day - Stars in the last hour (measured): 16 ## Stars per day, last 30 days (oldest → newest, today is partial) 2026-09-05 … 2026-10-04: 16, 23, 31, 32, 20, 37, 58, 41, 43, 53, 31, 41, 24, 22, 25, 38, 50, 34, 27, 18, 20, 17, 24, 27, 23, 30, 34, 168, 209, 84 ## Spotted in now - GitHub Trending today: #15, +211 stars ## Similar by description 1. **0xShug0/audio.cpp** — 1.7 · Steady · Generative media · C++ · +14 stars today, ≈ 22 by evening A native C++ inference engine for audio models built on ggml: TTS, speech recognition, VAD, voice cloning, music generation. Runs without Python on CPU, CUDA, ROCm, Vulkan and Metal. Full card: https://gitnova.dev/en/r/0xShug0/audio.cpp.md 2. **MiaAI-Lab/Qwen3.8-Flash-Next-Single-DGX-Spark-TensorFold** — 1.8 · Cooling · Language models · Shell · +6 stars today, ≈ 11 by evening Scripts to serve Qwen3.8 Flash Next on a single NVIDIA DGX Spark via an OpenAI-compatible API, with image and video input and a 262,144-token context. Full card: https://gitnova.dev/en/r/MiaAI-Lab/Qwen3.8-Flash-Next-Single-DGX-Spark-TensorFold.md 3. **JustVugg/colibri** — 3.7 · Steady · Language models · C · +66 stars today, ≈ 135 by evening A pure-C, zero-dependency inference engine for running large MoE models (744B–2.8T parameters) on consumer hardware by treating VRAM, RAM, and storage as a single multitier hierarchy and streaming experts from disk. Full card: https://gitnova.dev/en/r/JustVugg/colibri.md 4. **Niko1221/Strata** — 7.5 · Breakout · Language models · C++ · +543 stars today, ≈ 1,113 by evening A C++ local inference engine that runs the 125B MoE model Qwen3.8-Flash-Next on a regular PC with one NVIDIA GPU (12-24 GB) and 64 GB RAM, exposing an OpenAI/Anthropic-compatible API on localhost. Full card: https://gitnova.dev/en/r/Niko1221/Strata.md --- Magnitude (0–10) measures how fast and how unusually interest in a repository is growing right now. It is not a quality score. Days are UTC. “So far today” is a fact; “expected by the end of the day” is a forecast. Summaries and use cases are written by an LLM (DeepSeek V4.1 Flash) from the README and may be inaccurate: verify specific claims (benchmarks, speed, hardware) in the repository itself. Data as of 2026-10-04 13:56 UTC, updated every 30 minutes.