# Dreamer-Toby/STEPQuant > A post-training quantization method for Delta-rule recurrent states: it allocates bits per state unit based on error lifetime and impact on the output. It reduces inference memory for LLMs with quantized states. - Magnitude: 0.5 out of 10 — Steady - Stars: 55 total · +0 stars today, ≈ 2 by evening - Star trust: star growth looks organic - Category: Language models · Language: Python · Created: 2026-09-29 · Last push: 2026-10-01 - GitHub: https://github.com/Dreamer-Toby/STEPQuant · Homepage: https://arxiv.org/abs/2609.38169 · Page: https://gitnova.dev/en/r/Dreamer-Toby/STEPQuant ## Useful for - Calibrate a state quantization plan for Qwen and serve it via SGLang - Evaluate long-generation task accuracy with 4- and 6-bit states versus FP32 - Compare memory usage of state quantization against INT8 and INT6 ## Why it’s here - 0 stars so far today, about 1 expected by the end of the day. - The repository is 4 days old and already has 55 stars. With less than two weeks of history, there's no usual pace to compare the spike against yet. - Top new repositories this week: #198. ## Star trust Star growth looks organic. Star-trust labels are heuristics based on the repository’s behavior, not a check of every stargazer. ## Numbers - Forks: 0 - Issues and pull requests: 0 - Watchers: 0 - Average over the last week: 11 per day - Usual pace: too little history (under two weeks) - Stars in the last hour (measured): 1 ## Stars per day, last 7 days (oldest → newest, today is partial) 2026-09-27 … 2026-10-03: 0, 0, 10, 26, 17, 2, 0 ## Spotted in now - Top new repositories this week: #198 ## More in this category 1. **Niko1221/Strata** — 8.1 · Breakout · Language models · C++ · +0 stars today, ≈ 1,164 by evening A C++ local inference engine that runs the 125B MoE model Qwen3.8-Flash-Next on a regular PC with one NVIDIA GPU (12-24 GB) and 64 GB RAM, exposing an OpenAI/Anthropic-compatible API on localhost. Full card: https://gitnova.dev/en/r/Niko1221/Strata.md 2. **StayLameBro/backburner** — 5.2 · Early signal · Language models · Objective-C++ · +0 stars today, ≈ 73 by evening A llama.cpp fork that plugs an iPhone into a Mac over USB-C and splits the work: the Mac runs layers 1-40, the iPhone runs 41-64 on its GPU, speeding up prefill and allowing context up to 196k-229k tokens. Aimed at people running… Full card: https://gitnova.dev/en/r/StayLameBro/backburner.md 3. **Kutuyyy/Leaked-System-Prompt-AI** — 5.2 · Early signal · Language models · +0 stars today, ≈ 63 by evening A repository that collects and documents system prompts, instructions, and tool definitions of major AI models and agent platforms (OpenAI, Anthropic, Google, Cursor, Devin, etc.) for studying their behavior. Full card: https://gitnova.dev/en/r/Kutuyyy/Leaked-System-Prompt-AI.md 4. **Vibra-Ingenn/Janus** — 5.0 · Early signal · Language models · Go · +0 stars today, ≈ 31 by evening Local LLM server in Go: runs .gguf models via llama.cpp (Vulkan or CPU) and exposes an OpenAI-compatible API with a web UI and built-in tools. Full card: https://gitnova.dev/en/r/Vibra-Ingenn/Janus.md 5. **ninjahawk/livenerf** — 4.3 · Early signal · Language models · Python · +0 stars today, ≈ 68 by evening A benchmark for tracking whether a frontier model quietly degrades after release: it runs a frozen question panel daily and statistically measures drift in accuracy and token counts against the launch-week baseline. Full card: https://gitnova.dev/en/r/ninjahawk/livenerf.md --- Magnitude (0–10) measures how fast and how unusually interest in a repository is growing right now. It is not a quality score. Days are UTC. “So far today” is a fact; “expected by the end of the day” is a forecast. Summaries and use cases are written by an LLM (DeepSeek V4.1 Flash) from the README and may be inaccurate: verify specific claims (benchmarks, speed, hardware) in the repository itself. Data as of 2026-10-03 03:19 UTC, updated every 30 minutes.