EricLBuehler/mistral.rs
Fast, flexible LLM inference
About the project
A Rust LLM inference engine with automatic model loading, multimodality (text, image, video, audio), quantization, and an OpenAI/Anthropic-compatible server.
Useful for
- Run a local model via CLI for interactive chat or one-shot prompts
- Start an OpenAI-compatible API server with a built-in web UI
- Load a GGUF or quantized model and run inference on your own hardware
README summarized by DeepSeek V4.1 Flash. Details may be inaccurate.
Why it’s trending
- 1 star so far today, about 3 expected by the end of the day.
- GitHub Trending Rust today: #9, +4 stars.
- Recent forks include notable developers: @tobert (302 followers), @lizzz0523 (207 followers).
Stars per day
Bars are daily stars, the line is the usual pace. Red marks spike days.
Numbers
- Total stars
- 7,703
- Stars in a day
- 3
- Forks
- 708
- Issues and pull requests
- 2,384
- Watchers
- 46
- Language
- Rust
- License
- MIT
- Latest release
- v0.9.3 · September 7, 2026
- Created
- February 26, 2024
- Last push
- September 8, 2026
Star trust
Growth looks organic: forks and discussion are in line with active projects, and stars arrive unevenly, the way people give them.
These are heuristics, not a verdict: we judge by the repository’s behavior, not by a list of stargazers.
Spotted in
- September 19, 2026GitHub Trending Rust today: #9, +4 stars
Similar projects
-
7.7
TheoLeeCJ/SemIf
An open reproduction of the Jev-style semantic decision interface: reads typed option probabilities directly from a 4B model's logits without generating text. Runs on a single RTX 3090.
-
7.4
TianyuCodings/NanoJev
A nano replica of Jev built on Qwen3-0.6B that outputs probability distributions over candidates in a single forward pass with no token decoding, plus a training pipeline and game demos (maze, Snake).
-
6.4
Continuum-AI-Corp/OrcaBonsai-27B-Uncensored
A tool for runtime refusal ablation in the compressed Ternary Bonsai 2 27B LLM: it projects the residual stream to remove the refusal direction without modifying or re-quantizing weights. Runs on Apple Silicon via MLX.
-
6.1
jaredpalmer/kev
kev is a LoRA adapter with a small readout head on top of Qwen2.5-0.5B that answers many typed questions about a document in a single forward pass, returning calibrated probabilities instead of text.
-
6.0
githubnext/localjev
A local TypeScript/Bun HTTP service implementing a Jev-compatible POST /v1/systemone API on top of an OpenAI-compatible Chat Completions endpoint (DiffusionGemma via oMLX). It acts as a bridge for typed decisions…
-
6.0
higgsfield-ai/higgsfield
Fault-tolerant GPU orchestration and ML framework for distributed training of models with billions to trillions of parameters, including LLMs. Manages node access, experiment queueing, and GitHub Actions integration.