vllm-project/semantic-router
A programmable Mixture-of-Models router for heterogeneous LLM inference
About the project
A programmable routing layer for Mixture-of-Models systems across heterogeneous LLM infrastructure: it selects or composes the right model path per request based on request signals, user preferences, and application policies. It helps manage quality, cost, latency, privacy, and safety without hard-coding routing logic into applications.
Useful for
- Route sensitive-data requests to private or edge models and the rest to the cloud
- Cut inference cost by sending simple requests to cheap models and hard ones to stronger models
- Define per-user and per-workload model selection policies
README summarized by DeepSeek V4.1 Flash. Details may be inaccurate.
Why it’s trending
- 4 stars so far today, about 13 expected by the end of the day.
- Over the last two days the pace is 2.8× that of the previous week and a half.
- GitHub Trending Go today: #4, +22 stars.
- Recent forks include notable developers: @strategist922 (894 followers), @ItsRoy69 (245 followers).
Stars per day
Bars are daily stars, the line is the usual pace. Red marks spike days.
Numbers
- Total stars
- 5,978
- Today
- 4 · ≈ 13 by evening
- Forks
- 989
- Issues and pull requests
- 4,350
- Watchers
- 63
- Language
- Go
- License
- Apache-2.0
- Latest release
- v0.4.0 · September 27, 2026
- Created
- August 26, 2025
- Last push
- September 29, 2026
Star trust
Growth looks organic: forks and discussion are in line with active projects, and stars arrive unevenly, the way people give them.
These are heuristics, not a verdict: we judge by the repository’s behavior, not by a list of stargazers.
Spotted in
- September 29, 2026GitHub Trending Go today: #4, +22 stars
More in this category
-
8.2
firelex/jeff
Small fine-tuned Qwen3.5 and Gemma 4 models for zero-shot classification: given a situation description and a list of options, they return a calibrated probability for each option in a single forward pass.
-
7.4
VectifyAI/PageIndex
A vectorless RAG engine that builds a hierarchical tree index of a document and retrieves by LLM reasoning over that tree, the way a human navigates a report. Aimed at long professional PDFs such as financial, legal…
-
7.2
Niko1221/Strata
A C++ local inference engine that runs the 125B MoE model Qwen3.8-Flash-Next on a regular PC with one NVIDIA GPU (12-24 GB) and 64 GB RAM, exposing an OpenAI/Anthropic-compatible API on localhost.
-
5.6
ollaya-dev/ollaya
A local runtime for decision models: pulls and serves classification and routing models behind a TypeSafe-compatible API. Like Ollama, but for models that return probabilities instead of text.
-
5.3
PostHog/jeeves
Jeeves is a reasoning model built on Qwen3.5-9B (LoRA + pointer head), trained with SFT and CISPO, that thinks before answering and returns calibrated probabilities for yes/no, multiple-choice and rating questions via…
-
4.9
deepfates/imp
A port of DSPy to the BEAM: declarative LM programs in Elixir with signatures, optimizers and agents running as OTP processes.