# vllm-project/semantic-router > A programmable routing layer for Mixture-of-Models systems across heterogeneous LLM infrastructure: it selects or composes the right model path per request based on request signals, user preferences, and application policies. It helps manage quality, cost, latency, privacy, and safety without hard-coding routing logic into applications. - Magnitude: 2.7 out of 10 — Steady - Stars: 5,978 total · +4 stars today, ≈ 12 by evening - Star trust: star growth looks organic - Category: Language models · Language: Go · License: Apache-2.0 · Created: 2025-08-26 · Last push: 2026-09-29 - GitHub: https://github.com/vllm-project/semantic-router · Homepage: https://vllm-sr.ai · Page: https://gitnova.dev/en/r/vllm-project/semantic-router ## Useful for - Route sensitive-data requests to private or edge models and the rest to the cloud - Cut inference cost by sending simple requests to cheap models and hard ones to stronger models - Define per-user and per-workload model selection policies ## Why it’s here - 4 stars so far today, about 12 expected by the end of the day. - Over the last two days the pace is 2.8× that of the previous week and a half. - GitHub Trending Go today: #4, +22 stars. - Recent forks include notable developers: @strategist922 (894 followers), @ItsRoy69 (245 followers). ## Star trust Star growth looks organic. Star-trust labels are heuristics based on the repository’s behavior, not a check of every stargazer. ## Numbers - Forks: 990 - Issues and pull requests: 4,355 - Watchers: 63 - Average over the last week: 13 per day - Usual pace: 17 per day - Stars in the last hour (measured): 0 - Latest release: v0.4.0 (2026-09-27) ## Stars per day, last 30 days (oldest → newest, today is partial) 2026-08-31 … 2026-09-29: 27, 22, 23, 34, 32, 28, 28, 26, 27, 23, 18, 21, 20, 24, 23, 30, 11, 4, 8, 8, 9, 16, 10, 5, 8, 3, 3, 27, 31, 4 ## Spotted in now - GitHub Trending Go today: #4, +22 stars ## More in this category 1. **firelex/jeff** — 8.2 · Early signal · Language models · Python · +255 stars today, ≈ 486 by evening Small fine-tuned Qwen3.5 and Gemma 4 models for zero-shot classification: given a situation description and a list of options, they return a calibrated probability for each option in a single forward pass. Full card: https://gitnova.dev/en/r/firelex/jeff.md 2. **VectifyAI/PageIndex** — 7.4 · Breakout · Language models · Python · +285 stars today, ≈ 538 by evening A vectorless RAG engine that builds a hierarchical tree index of a document and retrieves by LLM reasoning over that tree, the way a human navigates a report. Aimed at long professional PDFs such as financial, legal and technical documents. Full card: https://gitnova.dev/en/r/VectifyAI/PageIndex.md 3. **Niko1221/Strata** — 7.2 · Early signal · Language models · C++ · +183 stars today, ≈ 387 by evening A C++ local inference engine that runs the 125B MoE model Qwen3.8-Flash-Next on a regular PC with one NVIDIA GPU (12-24 GB) and 64 GB RAM, exposing an OpenAI/Anthropic-compatible API on localhost. Full card: https://gitnova.dev/en/r/Niko1221/Strata.md 4. **PostHog/jeeves** — 6.6 · Early signal · Language models · Python · +83 stars today, ≈ 119 by evening Jeeves is a reasoning model built on Qwen3.5-9B (LoRA + pointer head), trained with SFT and CISPO, that thinks before answering and returns calibrated probabilities for yes/no, multiple-choice and rating questions via a Jev-compatible API. Full card: https://gitnova.dev/en/r/PostHog/jeeves.md 5. **ollaya-dev/ollaya** — 5.6 · Early signal · Language models · Rust · +40 stars today, ≈ 90 by evening A local runtime for decision models: pulls and serves classification and routing models behind a TypeSafe-compatible API. Like Ollama, but for models that return probabilities instead of text. Full card: https://gitnova.dev/en/r/ollaya-dev/ollaya.md --- Magnitude (0–10) measures how fast and how unusually interest in a repository is growing right now. It is not a quality score. Days are UTC. “So far today” is a fact; “expected by the end of the day” is a forecast. Summaries and use cases are written by an LLM (DeepSeek V4.1 Flash) from the README and may be inaccurate: verify specific claims (benchmarks, speed, hardware) in the repository itself. Data as of 2026-09-29 13:28 UTC, updated every 30 minutes.