ninjahawk/livenerf
Benchmark for tracking model capability after release.
About the project
A benchmark for tracking whether a frontier model quietly degrades after release: it runs a frozen question panel daily and statistically measures drift in accuracy and token counts against the launch-week baseline.
Useful for
- Run the daily question panel through headless Claude Code to collect a baseline
- Check whether model accuracy dropped over a 10-day window after release
- Track falling output token counts as an early sign of reduced model effort
README summarized by DeepSeek V4.1 Flash. Details may be inaccurate.
Why it’s trending
- 0 stars so far today, about 30 expected by the end of the day.
- The spike has held for 3 days in a row — not a one-off blip.
- The repository is 6 days old and already has 84 stars. With less than two weeks of history, there's no usual pace to compare the spike against yet.
- Top new repositories this week: #157.
Stars per day
Bars are daily stars, the line is the usual pace. Red marks spike days.
Numbers
- Total stars
- 84
- Today
- 0 · ≈ 30 by evening
- Forks
- 4
- Issues and pull requests
- 5
- Watchers
- 5
- Language
- Python
- Created
- September 22, 2026
- Last push
- September 27, 2026
Star trust
Growth looks organic: forks and discussion are in line with active projects, and stars arrive unevenly, the way people give them.
These are heuristics, not a verdict: we judge by the repository’s behavior, not by a list of stargazers.
Spotted in
- September 28, 2026Top new repositories this week: #157
More in this category
-
6.9
Niko1221/Strata
A C++ local inference engine that runs the 125B MoE model Qwen3.8-Flash-Next on a regular PC with one NVIDIA GPU (12-24 GB) and 64 GB RAM, exposing an OpenAI/Anthropic-compatible API on localhost.
-
6.5
ollaya-dev/ollaya
A local runtime for decision models: pulls and serves classification and routing models behind a TypeSafe-compatible API. Like Ollama, but for models that return probabilities instead of text.
-
6.3
deepfates/imp
A port of DSPy to the BEAM: declarative LM programs in Elixir with signatures, optimizers and agents running as OTP processes.
-
4.7
Mapika/decider
A language model that does not generate text: from a single forward pass it returns calibrated probabilities for typed questions (Choice, Score, Noul) about a given state. An open reproduction of the "System One" model…
-
4.6
NVIDIA/Model-Optimizer
NVIDIA library for compressing and accelerating models: quantization, pruning, distillation, NAS and speculative decoding with export to TensorRT-LLM, vLLM, SGLang. For ML engineers preparing models for deployment.
-
4.2
InternLM/Intern-Decision
A multimodal decision model built on Qwen3.5 that turns a state, optional images, and typed questions (choice, score, yes/no) into decisions with probabilities. Ships with training, two inference backends, temperature…