# NVIDIA/Model-Optimizer > NVIDIA library for compressing and accelerating models: quantization, pruning, distillation, NAS and speculative decoding with export to TensorRT-LLM, vLLM, SGLang. For ML engineers preparing models for deployment. - Magnitude: 2.7 out of 10 — Breakout - Stars: 3,902 total · +36 stars today, ≈ 56 by evening - Star trust: star growth looks organic - Category: Language models · Language: Python · License: Apache-2.0 · Created: 2024-04-23 · Last push: 2026-09-24 - GitHub: https://github.com/NVIDIA/Model-Optimizer · Homepage: https://nvidia.github.io/Model-Optimizer/ · Page: https://gitnova.dev/en/r/NVIDIA/Model-Optimizer ## Useful for - Quantize an LLM to FP8 or NVFP4 to speed up vLLM inference - Compress a model via pruning and distillation while keeping quality - Export an optimized checkpoint to TensorRT-LLM or SGLang ## Why it’s here - 36 stars so far today, about 56 expected by the end of the day. The usual pace is 10 per day, so that's 5.7× as much. - Over the last two days the pace is 6.4× that of the previous week and a half. - GitHub Trending today: #5, +22 stars. - GitHub Trending Python today: #3, +22 stars. ## Star trust Star growth looks organic. Star-trust labels are heuristics based on the repository’s behavior, not a check of every stargazer. ## Numbers - Forks: 621 - Issues and pull requests: 2,523 - Watchers: 32 - Average over the last week: 15 per day - Usual pace: 10 per day - Stars in the last hour (measured): 17 - Latest release: 0.47.0 (2026-09-23) ## Stars per day, last 30 days (oldest → newest, today is partial) 2026-08-26 … 2026-09-24: 7, 9, 32, 65, 35, 32, 22, 29, 17, 14, 6, 7, 5, 10, 8, 6, 7, 5, 6, 10, 5, 5, 5, 3, 5, 8, 7, 4, 21, 36 ## Spotted in now - GitHub Trending today: #5, +22 stars - GitHub Trending Python today: #3, +22 stars ## Similar by description 1. **microsoft/onnxruntime** — 1.6 · Steady · Language models · C++ · +2 stars today, ≈ 6 by evening Cross-platform accelerator for ML inference and training. Runs models from PyTorch, TensorFlow, scikit-learn and others via the ONNX format with graph optimizations and hardware acceleration. Full card: https://gitnova.dev/en/r/microsoft/onnxruntime.md 2. **NVIDIA/TensorRT-LLM** — 1.6 · Steady · Language models · Python · +3 stars today, ≈ 6 by evening NVIDIA library for optimizing inference of large language models and visual generative models on GPUs, with a Python API, specialized kernels, and an efficient C++/Python runtime. Full card: https://gitnova.dev/en/r/NVIDIA/TensorRT-LLM.md 3. **MiaAI-Lab/DeepSeek-v4.1-Flash-EXL3-2x-DGX-Sparks** — 1.1 · Cooling · Language models · Python · +2 stars today, ≈ 4 by evening A local EXL3 2.9 bpw checkpoint of DeepSeek-V4.1-Flash (196 GiB, 39 shards) served via an OpenAI-compatible vLLM API on a 2x NVIDIA GB10 (DGX Spark) kit with tensor-parallel 2. Includes DSpark speculative decoding and memory tuning for… Full card: https://gitnova.dev/en/r/MiaAI-Lab/DeepSeek-v4.1-Flash-EXL3-2x-DGX-Sparks.md 4. **hao-ai-lab/FastVideo** — 1.1 · Steady · Generative media · Python · +3 stars today, ≈ 5 by evening A framework for accelerated video generation: post-training (distillation, LoRA, sparse attention) and fast inference of diffusion video models on GPU and Apple Silicon. Full card: https://gitnova.dev/en/r/hao-ai-lab/FastVideo.md 5. **deepseek-ai/DeepSpec** — 1.4 · Steady · Language models · Python · +6 stars today, ≈ 10 by evening Full-stack pipeline for training and evaluating draft models for speculative decoding: data preparation, training against a target-model cache, and acceptance measurement on benchmarks. Supports DSpark, DFlash and Eagle3. Full card: https://gitnova.dev/en/r/deepseek-ai/DeepSpec.md --- Magnitude (0–10) measures how fast and how unusually interest in a repository is growing right now. It is not a quality score. Days are UTC. “So far today” is a fact; “expected by the end of the day” is a forecast. Summaries and use cases are written by an LLM (DeepSeek V4.1 Flash) from the README and may be inaccurate: verify specific claims (benchmarks, speed, hardware) in the repository itself. Data as of 2026-09-24 13:15 UTC, updated every 30 minutes.