# mw00/project-maya > Local engine, server and dashboard to run the 321B MoE model GLM-5.3-Flash on one or two NVIDIA GPUs, offloading experts to RAM and NVMe. - Magnitude: 4.9 out of 10 — Early signal - Stars: 77 total · +5 stars measured 2026-10-08, 04:56–06:22 UTC, ≈ 50 by evening - Star trust: star growth looks organic - Category: Language models · Language: C++ · License: MIT · Created: 2026-10-06 · Last push: 2026-10-08 - GitHub: https://github.com/mw00/project-maya · Page: https://gitnova.dev/en/r/mw00/project-maya ## Useful for - Run GLM-5.3-Flash locally on one or two NVIDIA GPUs without cloud - Deploy an OpenAI-compatible API for chat and image generation on own hardware - Benchmark decode and prefill speed on your machine via ./maya.sh --bench ## Why it’s here - Star-counter measurements on 2026-10-08 (UTC), 04:56–06:22: 72 → 77 stars (+5). This is the change over that interval. - Estimated end-of-day forecast: about +50 stars, using observed gains and the previous day. - The repository is 2 days old and already has 77 stars. With less than two weeks of history, there's no usual pace to compare the spike against yet. - Top new repositories this week: #190. - About 18 forks a day — people are taking the code. ## Star trust Star growth looks organic. Star-trust labels are heuristics based on the repository’s behavior, not a check of every stargazer. ## Numbers - Forks: 11 - Issues and pull requests: 7 - Watchers: 0 - Average over the last week: 43 per day - Usual pace: too little history (under two weeks) - Stars in the last hour (measured): 4 - Latest release: v1.0.6 (2026-10-08) ## Stars per day, last 5 days (oldest → newest, today is partial) 2026-10-04 … 2026-10-08: 0, 0, 0, 78, 5 ## Spotted in now - Top new repositories this week: #190 ## Similar by description 1. **sybil-solutions/dsv41-flash-offload** — 1.4 · Cooling · Language models · Python · +1 star measured 2026-10-08, 00:05–06:28 UTC, ≈ 4 by evening Serves DeepSeek-V4.1-Flash (EXL3 3.0 bpw) on a single 24 GB RTX 3090 by offloading experts to DDR4 and NVMe, exposing an OpenAI-compatible API. Aimed at running a large MoE model locally on consumer hardware. Full card: https://gitnova.dev/en/r/sybil-solutions/dsv41-flash-offload.md 2. **MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks-TensorFold** — 3.2 · Early signal · Language models · Shell · +7 stars measured 2026-10-08, 00:04–06:23 UTC, ≈ 25 by evening Scripts to serve GLM-5.3-Flash on two NVIDIA DGX Sparks via TensorFold with an OpenAI-compatible API, 1M-token context, and image/video input. Full card: https://gitnova.dev/en/r/MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks-TensorFold.md 3. **0xSero/DeepSeek-V4.1-Flash-Two-Sparks** — 1.1 · Cooling · Language models · Shell · +1 star measured 2026-10-08, 00:10–06:32 UTC, ≈ 3 by evening Scripts to serve DeepSeek-V4.1-Flash on two NVIDIA DGX Spark nodes with vLLM (tensor parallel 2 over RoCE), using EXL3-quantized experts, and to connect it to the Pi coding agent. Full card: https://gitnova.dev/en/r/0xSero/DeepSeek-V4.1-Flash-Two-Sparks.md 4. **jayleaton/glm53-tensorfold-spark** — 0.0 · Steady · Language models · Python · Not enough measurements today Experimental setup for serving the 4-bit GLM-5.3-Flash checkpoint across two NVIDIA DGX Sparks via the TensorFold engine, with an OpenAI-compatible API and a set of patches for faster decoding. Full card: https://gitnova.dev/en/r/jayleaton/glm53-tensorfold-spark.md 5. **Niko1221/Strata** — 5.7 · Peaking · Language models · C++ · +443 stars measured 2026-10-08, 00:01–06:21 UTC, ≈ 1,486 by evening A C++ local inference engine that runs the 125B MoE model Qwen3.8-Flash-Next on a regular PC with one NVIDIA GPU (12-24 GB) and 64 GB RAM, exposing an OpenAI/Anthropic-compatible API on localhost. Full card: https://gitnova.dev/en/r/Niko1221/Strata.md --- Magnitude (0–10) measures how fast and how unusually interest in a repository is growing right now. It is not a quality score. Days are UTC. “So far today” is a fact; “expected by the end of the day” is a forecast. Summaries and use cases are written by an LLM (DeepSeek V4.1 Flash) from the README and may be inaccurate: verify specific claims (benchmarks, speed, hardware) in the repository itself. Data as of 2026-10-08 06:33 UTC, updated every 30 minutes.