sgl-project/sglang
SGLang is a high-performance serving framework for large language models and multimodal models.
About the project
SGLang is an open-source inference framework for fast serving of large language and multimodal models, optimized for agentic workloads, RL rollouts, and large-scale deployment.
Useful for
- Launch a local inference server for Llama or DeepSeek using the lmsysorg/sglang Docker image
- Deploy a multimodal VLM model across multiple GPUs with load balancing
- Use SGLang as a rollout generation backend for RL training via Miles or verl
README summarized by DeepSeek V4.1 Flash. Details may be inaccurate.
Why it’s trending
- 13 stars so far today, about 31 expected by the end of the day.
- GitHub Trending Python today: #19, +39 stars.
- About 18 forks a day — people are taking the code.
- Recent forks include notable developers: @bvolpato (1,763 followers), @geelen (1,536 followers).
Stars per day
Bars are daily stars, the line is the usual pace. Red marks spike days.
Numbers
- Total stars
- 36,691
- Today
- 13 · ≈ 31 by evening
- Forks
- 9,250
- Issues and pull requests
- 41,552
- Watchers
- 184
- Language
- Python
- License
- Apache-2.0
- Latest release
- v0.5.20 · September 18, 2026
- Created
- January 8, 2024
- Last push
- October 1, 2026
Star trust
Unusual star pattern. The magnitude is lowered, not zeroed:
- Stars arrived in one batch on a single day (1,248) and stopped right after: 93 over the next two days. Real interest usually fades gradually.
These are heuristics, not a verdict: we judge by the repository’s behavior, not by a list of stargazers.
Spotted in
- October 1, 2026GitHub Trending Python today: #19, +39 stars
Similar by description
-
1.2
vllm-project/vllm-ascend
A hardware plugin for vLLM that runs LLM inference on Ascend NPUs (Atlas A2/A3). It enables deployment of Transformer, MoE, embedding and multimodal models on Huawei Ascend hardware.
-
0.9
kserve/kserve
KServe is a platform for deploying and serving generative and predictive AI models on Kubernetes, supporting multiple frameworks, autoscaling, and a standardized inference protocol.
-
1.7
microsoft/onnxruntime
Cross-platform accelerator for ML inference and training. Runs models from PyTorch, TensorFlow, scikit-learn and others via the ONNX format with graph optimizations and hardware acceleration.
-
1.5
amitshekhariitbhu/ai-system-design
A study guide to designing AI systems built on LLMs, RAG, and agents: from inference and GPUs to MCP, multi-agent systems, and interview preparation.
-
1.7
ml-explore/mlx
MLX is an array and machine learning framework from Apple, optimized for Apple silicon with unified memory and lazy computation. Its Python, C++, C, and Swift APIs mirror NumPy and PyTorch, simplifying model training…
-
0.4
openxla/xla
XLA (Accelerated Linear Algebra) is an open-source ML compiler for GPUs, CPUs, and ML accelerators. It takes models from PyTorch, TensorFlow, and JAX and optimizes them for high-performance execution across different…