# SamsungLabs/LittleBit > Official implementation of LittleBit (NeurIPS 2025) and LittleBit-2 (ICML 2026) — sub-1-bit LLM quantization methods that factorize weights into low-rank latent factors, binarize them, and restore magnitude with learned scales. - Magnitude: 2.9 out of 10 — Steady - Stars: 107 total · +12 stars measured 2026-10-08, 16:51–18:02 UTC, ≈ 15 by evening - Star trust: star growth looks organic - Category: Language models · Language: Python · Created: 2025-11-18 · Last push: 2026-05-06 - GitHub: https://github.com/SamsungLabs/LittleBit · Page: https://gitnova.dev/en/r/SamsungLabs/LittleBit ## Useful for - Compress Llama 2 7B to 1.0 or 0.1 bits per weight via QAT with SmoothSign - Enable Joint-ITQ initialization with --use_itq to improve latent geometry - Evaluate a checkpoint on wikitext2/c4 perplexity and zeroshot tasks via eval.py ## Why it’s here - Star-counter measurements on 2026-10-08 (UTC), 16:51–18:02: 95 → 107 stars (+12). This is the change over that interval. - Estimated end-of-day forecast: about +15 stars, using observed gains and the previous day. - Hacker News: “Sub-1-Bit LLM Compression via Latent Factorization” — 58 points, 4 h ago. - Recent forks include notable developers: @alexdolbun (708 followers), @danzeeeman (263 followers). ## Star trust Star growth looks organic. Star-trust labels are heuristics based on the repository’s behavior, not a check of every stargazer. ## Numbers - Forks: 16 - Issues and pull requests: 21 - Watchers: 3 - Average over the last week: 2 per day - Usual pace: 0 per day - Stars in the last hour (measured): 10 ## Stars per day, last 30 days (oldest → newest, today is partial) 2026-09-09 … 2026-10-08: 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 12 ## Hacker News - Sub-1-Bit LLM Compression via Latent Factorization — 58 points, 8 comments: https://news.ycombinator.com/item?id=50005608 ## Spotted in now - Spotted on Hacker News ## More in this category 1. **Niko1221/Strata** — 5.7 · Peaking · Language models · C++ · +1,214 stars measured 2026-10-08, 00:01–17:58 UTC, ≈ 1,579 by evening A C++ local inference engine that runs the 125B MoE model Qwen3.8-Flash-Next on a regular PC with one NVIDIA GPU (12-24 GB) and 64 GB RAM, exposing an OpenAI/Anthropic-compatible API on localhost. Full card: https://gitnova.dev/en/r/Niko1221/Strata.md 2. **mw00/project-maya** — 4.9 · Early signal · Language models · C++ · +43 stars measured 2026-10-08, 04:56–17:58 UTC, ≈ 58 by evening Local engine, server and dashboard to run the 321B MoE model GLM-5.3-Flash on one or two NVIDIA GPUs, offloading experts to RAM and NVMe. Full card: https://gitnova.dev/en/r/mw00/project-maya.md 3. **Edge0-AI/Edge0** — 4.4 · Breakout · Language models · Python · +216 stars measured 2026-10-08, 00:03–17:59 UTC, ≈ 280 by evening An open-source streaming MoE inference framework: expert weights are offloaded from SSD on demand while a trained prerouter predicts routing ahead of time. Runs on Apple Silicon via MLX and ships with two ready-to-run model tiers (35B and… Full card: https://gitnova.dev/en/r/Edge0-AI/Edge0.md 4. **deepseek-ai/DeepGEMM** — 4.3 · Peaking · Language models · Cuda · +16 stars measured 2026-10-08, 00:01–17:59 UTC, ≈ 24 by evening CUDA tensor core kernel library: GEMM (FP8, FP4, BF16), fused MoE, MQA scoring for the indexer. Kernels compile at runtime via DeepJIT, no CUDA build at install time. Full card: https://gitnova.dev/en/r/deepseek-ai/DeepGEMM.md 5. **StayLameBro/backburner** — 4.2 · Cooling · Language models · Python · +40 stars measured 2026-10-08, 00:02–17:59 UTC, ≈ 53 by evening A llama.cpp fork that plugs an iPhone into a Mac over USB-C and splits the work: the Mac runs layers 1-40, the iPhone runs 41-64 on its GPU, speeding up prefill and allowing context up to 196k-229k tokens. Aimed at people running… Full card: https://gitnova.dev/en/r/StayLameBro/backburner.md --- Magnitude (0–10) measures how fast and how unusually interest in a repository is growing right now. It is not a quality score. Days are UTC. “So far today” is a fact; “expected by the end of the day” is a forecast. Summaries and use cases are written by an LLM (DeepSeek V4.1 Flash) from the README and may be inaccurate: verify specific claims (benchmarks, speed, hardware) in the repository itself. Data as of 2026-10-08 18:19 UTC, updated every 30 minutes.