# Niko1221/Strata > A C++ local inference engine that runs the 125B MoE model Qwen3.8-Flash-Next on a regular PC with one NVIDIA GPU (12-24 GB) and 64 GB RAM, exposing an OpenAI/Anthropic-compatible API on localhost. - Magnitude: 4.5 out of 10 — Early signal - Stars: 88 total · +54 stars today, ≈ 85 by evening - Star trust: star growth looks organic - Category: Language models · Language: C++ · Created: 2026-09-24 · Last push: 2026-09-25 - GitHub: https://github.com/Niko1221/Strata · Page: https://gitnova.dev/en/r/Niko1221/Strata ## Useful for - Run a private chat with a large model on your own PC without sending data to the cloud - Connect the local model as an OpenAI-compatible provider to a coding agent or script - Send a screenshot or scanned page to the model for recognition via image input ## Why it’s here - 54 stars so far today, about 85 expected by the end of the day. - The spike has held for 3 days in a row — not a one-off blip. - The repository is 2 days old and already has 88 stars. With less than two weeks of history, there's no usual pace to compare the spike against yet. - Top new repositories this week: #164. - About 10 forks a day — people are taking the code. ## Star trust Star growth looks organic. Star-trust labels are heuristics based on the repository’s behavior, not a check of every stargazer. ## Numbers - Forks: 10 - Issues and pull requests: 11 - Watchers: 2 - Average over the last week: 40 per day - Usual pace: too little history (under two weeks) - Stars in the last hour (measured): 7 - Latest release: v0.1.2 (2026-09-25) ## Stars per day, last 7 days (oldest → newest, today is partial) 2026-09-20 … 2026-09-26: 0, 0, 0, 0, 17, 19, 54 ## Spotted in now - Top new repositories this week: #164 ## Similar by description 1. **AtomicBot-ai/Atomic-Chat** — 0.9 · Steady · Language models · TypeScript · +3 stars today, ≈ 5 by evening Desktop app and local inference engine for running open-weight LLMs (Llama, Gemma, Qwen, Mistral, etc.) on your own machine, exposing an OpenAI-compatible API at localhost:1337/v1. Full card: https://gitnova.dev/en/r/AtomicBot-ai/Atomic-Chat.md 2. **incoai/splash** — 2.5 · Cooling · Language models · C++ · +14 stars today, ≈ 30 by evening A local LLM inference engine for Apple silicon, specialized per model: serves OpenAI- and Anthropic-compatible APIs to coding agents on a single Mac. Full card: https://gitnova.dev/en/r/incoai/splash.md 3. **Neroued/ninfer** — 1.9 · Cooling · Language models · C++ · +8 stars today, ≈ 20 by evening NInfer is a from-scratch C++/CUDA inference engine for Qwen3.5 Dense and MoE models on a single NVIDIA RTX 5090. It serves text, image, and video prompts via a local CLI or OpenAI-/Anthropic-compatible HTTP APIs. Full card: https://gitnova.dev/en/r/Neroued/ninfer.md 4. **architectds/collabosm** — 3.3 · Early signal · Language models · Python · +1 star today, ≈ 12 by evening A set of scripts to run Qwen3.8-Flash-Next (125B MoE) on a single Colab A100-80GB High-RAM with an OpenAI-compatible endpoint and measured throughput. Full card: https://gitnova.dev/en/r/architectds/collabosm.md 5. **EricLBuehler/mistral.rs** — 0.5 · Steady · Language models · Rust · +0 stars today, ≈ 0 by evening A Rust LLM inference engine with automatic model loading, multimodality (text, image, video, audio), quantization, and an OpenAI/Anthropic-compatible server. Full card: https://gitnova.dev/en/r/EricLBuehler/mistral.rs.md --- Magnitude (0–10) measures how fast and how unusually interest in a repository is growing right now. It is not a quality score. Days are UTC. “So far today” is a fact; “expected by the end of the day” is a forecast. Summaries and use cases are written by an LLM (DeepSeek V4.1 Flash) from the README and may be inaccurate: verify specific claims (benchmarks, speed, hardware) in the repository itself. Data as of 2026-09-26 12:21 UTC, updated every 30 minutes.