# 1314521gjy/ninfer-fusion-kvmem > A downstream fork of the NInfer inference engine targeting small VRAM with long context: a KVMem ring plus host-backed KV-cache reuse so context isn't refilled every turn. Early version with known bugs, personal use only. - Magnitude: 3.6 out of 10 — Early signal - Stars: 103 total · +4 stars measured 2026-10-06, 15:03–17:36 UTC, ≈ 8 by evening - Star trust: star growth looks organic - Category: Language models · Language: C++ · License: Apache-2.0 · Created: 2026-10-04 · Last push: 2026-10-06 - GitHub: https://github.com/1314521gjy/ninfer-fusion-kvmem · Page: https://gitnova.dev/en/r/1314521gjy/ninfer-fusion-kvmem ## Useful for - Run the engine with a 4,000-token KV pool and check answers on a 96K context - Build the engine from source following the CUDA 13.3 and MSVC guide - Verify KV-cache reuse on the second dialogue turn with the smoke script ## Why it’s here - Star-counter measurements on 2026-10-06 (UTC), 15:03–17:36: 99 → 103 stars (+4). This is the change over that interval. - Estimated end-of-day forecast: about +8 stars, using observed gains and the previous day. - The repository is 2 days old and already has 103 stars. With less than two weeks of history, there's no usual pace to compare the spike against yet. - Top new repositories this week: #136. ## Star trust Star growth looks organic. Star-trust labels are heuristics based on the repository’s behavior, not a check of every stargazer. ## Numbers - Forks: 11 - Issues and pull requests: 3 - Watchers: 1 - Average over the last week: 32 per day - Usual pace: too little history (under two weeks) - Stars in the last hour (measured): 0 - Latest release: engine-v0.11.0-kvmem-20261003 (2026-10-04) ## Stars per day, last 3 days (oldest → newest, today is partial) 2026-10-04 … 2026-10-06: 39, 49, 4 ## Spotted in now - Top new repositories this week: #136 ## More in this category 1. **Niko1221/Strata** — 7.2 · Breakout · Language models · C++ · +1,501 stars measured 2026-10-06, 00:17–17:34 UTC, ≈ 2,059 by evening A C++ local inference engine that runs the 125B MoE model Qwen3.8-Flash-Next on a regular PC with one NVIDIA GPU (12-24 GB) and 64 GB RAM, exposing an OpenAI/Anthropic-compatible API on localhost. Full card: https://gitnova.dev/en/r/Niko1221/Strata.md 2. **deepseek-ai/DeepGEMM** — 5.9 · Breakout · Language models · Cuda · +102 stars measured 2026-10-06, 11:46–17:34 UTC, ≈ 142 by evening CUDA tensor core kernel library: GEMM (FP8, FP4, BF16), fused MoE, MQA scoring for the indexer. Kernels compile at runtime via DeepJIT, no CUDA build at install time. Full card: https://gitnova.dev/en/r/deepseek-ai/DeepGEMM.md 3. **Hiteater-wzm/eeo** — 5.4 · Early signal · Language models · JavaScript · +61 stars measured 2026-10-06, 00:23–17:35 UTC, ≈ 84 by evening An open community and toolkit for "Everything Engine Optimization": brands publish machine-readable cards about themselves, and the tooling measures how visible a brand is in AI engine answers. Full card: https://gitnova.dev/en/r/Hiteater-wzm/eeo.md 4. **StayLameBro/backburner** — 5.4 · Early signal · Language models · Python · +77 stars measured 2026-10-06, 00:17–17:35 UTC, ≈ 108 by evening A llama.cpp fork that plugs an iPhone into a Mac over USB-C and splits the work: the Mac runs layers 1-40, the iPhone runs 41-64 on its GPU, speeding up prefill and allowing context up to 196k-229k tokens. Aimed at people running… Full card: https://gitnova.dev/en/r/StayLameBro/backburner.md 5. **empero-org/brewery-ai** — 4.7 · Early signal · Language models · Python · +7 stars measured 2026-10-06, 00:18–17:35 UTC, ≈ 15 by evening A console agent that walks users through the full fine-tuning cycle for language and image models via a chat with an AI guide: from picking a base model and data to training on a GPU and publishing to Hugging Face. Full card: https://gitnova.dev/en/r/empero-org/brewery-ai.md --- Magnitude (0–10) measures how fast and how unusually interest in a repository is growing right now. It is not a quality score. Days are UTC. “So far today” is a fact; “expected by the end of the day” is a forecast. Summaries and use cases are written by an LLM (DeepSeek V4.1 Flash) from the README and may be inaccurate: verify specific claims (benchmarks, speed, hardware) in the repository itself. Data as of 2026-10-06 17:51 UTC, updated every 30 minutes.