# lithos-ai/lithos-metal > Open-source LLM inference engine for Apple silicon that compiles Metal megakernels and runs a local server with DSpark speculative decoding. - Magnitude: 4.5 out of 10 — Early signal - Stars: 78 total · +4 stars measured 2026-10-09, 03:36–05:00 UTC, ≈ 49 by evening - Star trust: star growth looks organic - Category: Language models · Language: Python · License: Apache-2.0 · Created: 2026-10-05 · Last push: 2026-10-08 - GitHub: https://github.com/lithos-ai/lithos-metal · Page: https://gitnova.dev/en/r/lithos-ai/lithos-metal ## Useful for - Run a local Qwen3.8-27B server on a Mac via lithos-metal serve - Connect a coding agent (opencode, claude, codex) to the local model - Call the Chat Completions API with curl to generate text ## Why it’s here - Star-counter measurements on 2026-10-09 (UTC), 03:36–05:00: 74 → 78 stars (+4). This is the change over that interval. - Estimated end-of-day forecast: about +49 stars, using observed gains and the previous day. - The repository is 4 days old and already has 78 stars. With less than two weeks of history, there's no usual pace to compare the spike against yet. - Top new repositories this week: #196. - Recent forks include notable developers: @zhaochenyang20 (2,932 followers). ## Star trust Star growth looks organic. Star-trust labels are heuristics based on the repository’s behavior, not a check of every stargazer. ## Numbers - Forks: 4 - Issues and pull requests: 7 - Watchers: 1 - Average over the last week: 25 per day - Usual pace: too little history (under two weeks) - Stars in the last hour (measured): 2 - Latest release: v0.1.2 (2026-10-05) ## Stars per day, last 6 days (oldest → newest, today is partial) 2026-10-04 … 2026-10-09: 0, 5, 2, 3, 68, 4 ## Spotted in now - Top new repositories this week: #196 ## More in this category 1. **SamsungLabs/LittleBit** — 6.3 · Early signal · Language models · Python · +20 stars measured 2026-10-09, 00:06–04:59 UTC, ≈ 110 by evening Official implementation of LittleBit (NeurIPS 2025) and LittleBit-2 (ICML 2026) — sub-1-bit LLM quantization methods that factorize weights into low-rank latent factors, binarize them, and restore magnitude with learned scales. Full card: https://gitnova.dev/en/r/SamsungLabs/LittleBit.md 2. **Albert-Weasker/niubigeo** — 5.7 · Breakout · Language models · TypeScript · +395 stars measured 2026-10-09, 00:12–05:00 UTC, ≈ 969 by evening Open-source tool for tracking brand visibility and competitors in AI answers: given a domain, it shows how models describe your product, whom they recommend, and which sources they cite. Full card: https://gitnova.dev/en/r/Albert-Weasker/niubigeo.md 3. **Niko1221/Strata** — 5.4 · Peaking · Language models · C++ · +306 stars measured 2026-10-09, 00:05–05:00 UTC, ≈ 1,286 by evening A C++ local inference engine that runs the 125B MoE model Qwen3.8-Flash-Next on a regular PC with one NVIDIA GPU (12-24 GB) and 64 GB RAM, exposing an OpenAI/Anthropic-compatible API on localhost. Full card: https://gitnova.dev/en/r/Niko1221/Strata.md 4. **mw00/project-maya** — 5.2 · Early signal · Language models · C++ · +20 stars measured 2026-10-09, 00:05–05:00 UTC, ≈ 79 by evening Local engine, server and dashboard to run the 321B MoE model GLM-5.3-Flash on one or two NVIDIA GPUs, offloading experts to RAM and NVMe. Full card: https://gitnova.dev/en/r/mw00/project-maya.md 5. **ymcrcat/rgpu** — 4.5 · Early signal · Language models · C++ · +13 stars measured 2026-10-09, 00:09–05:00 UTC, ≈ 41 by evening Runs PyTorch operations and holds tensors on a remote NVIDIA GPU while the application stays on the client, including on a Mac without CUDA. Offers two paths: an rgpu device for PyTorch and a CUDA shim for existing Linux programs. Full card: https://gitnova.dev/en/r/ymcrcat/rgpu.md --- Magnitude (0–10) measures how fast and how unusually interest in a repository is growing right now. It is not a quality score. Days are UTC. “So far today” is a fact; “expected by the end of the day” is a forecast. Summaries and use cases are written by an LLM (DeepSeek V4.1 Flash) from the README and may be inaccurate: verify specific claims (benchmarks, speed, hardware) in the repository itself. Data as of 2026-10-09 05:10 UTC, updated every 30 minutes.