# lyogavin/airllm > A library for inference and training of large language models on low-VRAM GPUs: it streams weights layer by layer from disk, letting 70B–671B models run on 4–12 GB of VRAM without quantization or distillation. - Magnitude: 4.9 out of 10 — Breakout - Stars: 36,117 total · +29 stars measured 2026-10-11, 11:04–12:27 UTC, ≈ 146 by evening - Star trust: star growth looks organic - Category: Language models · Language: Jupyter Notebook · License: Apache-2.0 · Created: 2023-06-12 · Last push: 2026-10-11 - GitHub: https://github.com/lyogavin/airllm · Page: https://gitnova.dev/en/r/lyogavin/airllm ## Useful for - Run Llama 3 70B or DeepSeek-V3 locally on a single consumer GPU with 4–12 GB of VRAM - Speed up inference up to 3x by enabling 4bit/8bit block-wise quantization via bitsandbytes - Fine-tune a 125B model on an RTX 3060 Ti by keeping adapters on GPU and streaming frozen layers ## Why it’s here - Star-counter measurements on 2026-10-11 (UTC), 11:04–12:27: 36088 → 36117 stars (+29). This is the change over that interval. - Estimated end-of-day forecast: about +146 stars, using observed gains and the previous day. - Over the last two days the pace is 7.4× that of the previous week and a half. - GitHub Trending today: #7, +541 stars. - About 41 forks a day — people are taking the code. ## Star trust Star growth looks organic. Star-trust labels are heuristics based on the repository’s behavior, not a check of every stargazer. ## Numbers - Forks: 3,791 - Issues and pull requests: 332 - Watchers: 302 - Average over the last week: 113 per day - Usual pace: 61 per day - Stars in the last hour (measured): 21 - Latest release: v4.0.0 (2026-09-05) ## Stars per day, last 30 days (oldest → newest, today is partial) 2026-09-12 … 2026-10-11: 53, 73, 55, 37, 55, 69, 51, 44, 44, 58, 41, 44, 33, 170, 94, 102, 51, 51, 44, 27, 25, 57, 47, 44, 41, 35, 28, 36, 463, 29 ## Spotted in now - GitHub Trending today: #7, +541 stars ## More in this category 1. **pytorch/pytorch** — 5.4 · Breakout · Language models · Python · +131 stars measured 2026-10-11, 00:39–12:27 UTC, ≈ 237 by evening PyTorch is a library for tensor computation with strong GPU acceleration and dynamic neural networks built on tape-based autograd. It is used as a GPU-ready NumPy replacement and as a deep learning research platform. Full card: https://gitnova.dev/en/r/pytorch/pytorch.md 2. **huggingface/transformers** — 5.3 · Breakout · Language models · Python · +237 stars measured 2026-10-11, 00:38–12:27 UTC, ≈ 446 by evening Model-definition framework for state-of-the-art pretrained ML models (text, vision, audio, multimodal) for inference and training, compatible with most training and inference engines. Full card: https://gitnova.dev/en/r/huggingface/transformers.md 3. **NandhaKishorM/vegaml** — 5.2 · Early signal · Language models · Python · +29 stars measured 2026-10-11, 02:45–12:27 UTC, ≈ 71 by evening Library for typed decisions from a frozen language model and a small physics engine: returns a choice, score, or probability with calibration, conformal sets, and an abstain flag. Supports 73k context and images, runs locally. Full card: https://gitnova.dev/en/r/NandhaKishorM/vegaml.md 4. **Albert-Weasker/niubigeo** — 5.0 · Breakout · Language models · TypeScript · +182 stars measured 2026-10-11, 00:38–12:27 UTC, ≈ 429 by evening Open-source tool for tracking brand visibility and competitors in AI answers: given a domain, it shows how models describe your product, whom they recommend, and which sources they cite. Full card: https://gitnova.dev/en/r/Albert-Weasker/niubigeo.md 5. **microsoft/BitNet** — 5.0 · Breakout · Language models · C++ · +9 stars measured 2026-10-11, 11:12–12:27 UTC, ≈ 62 by evening Official inference framework for 1-bit LLMs (BitNet b1.58) with optimized kernels for CPU and GPU. Runs ternary models fast and losslessly, including a 100B model on a single CPU. Full card: https://gitnova.dev/en/r/microsoft/BitNet.md --- Magnitude (0–10) measures how fast and how unusually interest in a repository is growing right now. It is not a quality score. Days are UTC. “So far today” is a fact; “expected by the end of the day” is a forecast. Summaries and use cases are written by an LLM (DeepSeek V4.1 Flash) from the README and may be inaccurate: verify specific claims (benchmarks, speed, hardware) in the repository itself. Data as of 2026-10-11 12:38 UTC, updated every 30 minutes.