# vllm-project/vllm > A high-throughput and memory-efficient inference and serving engine for large language models, using PagedAttention for KV-cache management. Supports 200+ model architectures and an OpenAI-compatible API. - Magnitude: 2.8 out of 10 — Steady - Stars: 93,326 total · +7 stars measured 2026-10-07, 11:04–12:42 UTC, ≈ 25 by evening - Star trust: star growth looks organic - Category: Language models · Language: Python · License: Apache-2.0 · Created: 2023-02-09 · Last push: 2026-10-07 - GitHub: https://github.com/vllm-project/vllm · Homepage: https://vllm.ai · Page: https://gitnova.dev/en/r/vllm-project/vllm ## Useful for - Deploy an OpenAI-compatible server for a local LLM - Speed up model inference with continuous batching and PagedAttention - Serve a multimodal or MoE Hugging Face model on your own GPUs ## Why it’s here - Star-counter measurements on 2026-10-07 (UTC), 11:04–12:42: 93319 → 93326 stars (+7). This is the change over that interval. - Estimated end-of-day forecast: about +25 stars, using observed gains and the previous day. - GitHub Trending Python today: #12, +73 stars. - About 35 forks a day — people are taking the code. ## Star trust Star growth looks organic. Star-trust labels are heuristics based on the repository’s behavior, not a check of every stargazer. ## Numbers - Forks: 23,081 - Issues and pull requests: 59,572 - Watchers: 599 - Average over the last week: 54 per day - Usual pace: 80 per day - Stars in the last hour (measured): 0 - Latest release: v0.31.0 (2026-10-05) ## Stars per day, last 30 days (oldest → newest, today is partial) 2026-09-08 … 2026-10-07: 96, 95, 81, 75, 66, 83, 108, 101, 88, 102, 80, 70, 77, 116, 82, 107, 81, 51, 50, 59, 78, 83, 55, 62, 41, 57, 57, 60, 72, 7 ## Spotted in now - GitHub Trending Python today: #12, +73 stars ## Similar by description 1. **sgl-project/sglang** — 2.1 · Steady · Language models · Python · +12 stars measured 2026-10-07, 00:24–12:46 UTC, ≈ 23 by evening · unusual star pattern (heuristic) SGLang is an open-source inference framework for fast serving of large language and multimodal models, optimized for agentic workloads, RL rollouts, and large-scale deployment. Full card: https://gitnova.dev/en/r/sgl-project/sglang.md 2. **ggml-org/llama.cpp** — 3.4 · Steady · Language models · C++ · +52 stars measured 2026-10-07, 00:21–12:40 UTC, ≈ 104 by evening LLM and VLM inference implemented in C/C++ with no dependencies, running on CPUs and GPUs across many hardware backends. Enables local model execution with minimal setup and quantization. Full card: https://gitnova.dev/en/r/ggml-org/llama.cpp.md 3. **kvcache-ai/Mooncake** — 1.4 · Steady · Language models · C++ · +5 stars measured 2026-10-07, 00:27–12:50 UTC, ≈ 8 by evening LLM serving platform built on a KVCache-centric disaggregated architecture: separates prefill and decode and transfers KV cache between nodes over RDMA. Powers Kimi in production at Moonshot AI. Full card: https://gitnova.dev/en/r/kvcache-ai/Mooncake.md 4. **FlashML-org/FreeToken** — 2.0 · Steady · Language models · Python · +7 stars measured 2026-10-07, 00:23–12:46 UTC, ≈ 17 by evening A Mixture-of-Experts inference engine for running frontier-scale open-weight MoE models (290B+) on consumer hardware — GPUs, CPUs, and host memory. Aimed at users who want to run frontier models locally. Full card: https://gitnova.dev/en/r/FlashML-org/FreeToken.md 5. **microsoft/onnxruntime** — 1.4 · Steady · Language models · C++ · +3 stars measured 2026-10-07, 00:26–12:51 UTC, ≈ 6 by evening Cross-platform accelerator for ML inference and training. Runs models from PyTorch, TensorFlow, scikit-learn and others via the ONNX format with graph optimizations and hardware acceleration. Full card: https://gitnova.dev/en/r/microsoft/onnxruntime.md --- Magnitude (0–10) measures how fast and how unusually interest in a repository is growing right now. It is not a quality score. Days are UTC. “So far today” is a fact; “expected by the end of the day” is a forecast. Summaries and use cases are written by an LLM (DeepSeek V4.1 Flash) from the README and may be inaccurate: verify specific claims (benchmarks, speed, hardware) in the repository itself. Data as of 2026-10-07 12:58 UTC, updated every 30 minutes.