gauravapiscean/agentic-kv-cache
Reproducing agentic KV-cache policy claims on real traces. 68k requests from 393 Claude Code sessions. LRU is harder to beat than the papers suggest.
About the project
A prefix-cache simulator for LLM serving that replays real Claude Code and Mooncake traces to test whether published policies beat LRU. The author finds LRU is harder to beat than the papers claim.
Useful for
- Test your own KV-cache eviction policy on real agent-session traces
- Reproduce Mooncake's published hit-rate-vs-capacity curve
- Measure how much recompute comes from capacity pressure rather than TTL
README summarized by DeepSeek V4.1 Flash. Details may be inaccurate.
Why it’s trending
- 2 stars today.
- The repository is 3 days old and already has 21 stars.
- Hacker News: “LRU is harder to beat than the KV-cache papers suggest” — 108 points, 3 days ago.
Stars per day
Bars are daily stars, the line is the usual pace. Red marks spike days.
Numbers
- Total stars
- 21
- Stars in a day
- 2
- Forks
- 0
- Issues and pull requests
- 0
- Watchers
- 0
- Language
- Python
- License
- MIT
- Created
- September 10, 2026
- Last push
- September 10, 2026
Star trust
Growth looks organic: forks and discussion are in line with active projects, and stars arrive unevenly, the way people give them.
These are heuristics, not a verdict: we judge by the repository’s behavior, not by a list of stargazers.
Hacker News discussions
- LRU is harder to beat than the KV-cache papers suggest108 points, September 10, 2026
Spotted in
- September 13, 2026Spotted on Hacker News
Similar projects
-
6.1
JustVugg/colibri
A pure-C, zero-dependency inference engine for running large MoE models (744B–2.8T parameters) on consumer hardware by treating VRAM, RAM, and storage as a single multitier hierarchy and streaming experts from disk.
-
5.7
asgeirtj/system_prompts_leaks
A collection of extracted system prompts from Anthropic, OpenAI, Google, xAI and others — the hidden instructions chatbots receive before a user's first message.
-
5.1
kennethwolters/litelm
A lightweight litellm alternative: routes LLM calls across 19 providers and translates message formats in ~2,900 lines with two dependencies, without proxy, caching, or cost tracking.
-
5.0
MiaAI-Lab/DeepSeek-v4.1-Flash-EXL3-2x-DGX-Sparks
A local EXL3 2.9 bpw checkpoint of DeepSeek-V4.1-Flash (196 GiB, 39 shards) served via an OpenAI-compatible vLLM API on a 2x NVIDIA GB10 (DGX Spark) kit with tensor-parallel 2. Includes DSpark speculative decoding and…
-
4.8
unclecode/crawl4ai
Open-source web crawler and scraper that turns pages into clean, LLM-ready Markdown for RAG, agents and data pipelines. Runs via Python API, CLI and Docker with no API keys.
-
4.5
Edge0-AI/Edge0
An open-source streaming MoE inference framework: expert weights are offloaded from SSD on demand while a trained prerouter predicts routing ahead of time. Runs on Apple Silicon via MLX and ships with two ready-to-run…