# Simreal-AI/Simreal-MLBench > Open benchmark for evaluating ML research agents on 60 tasks with external scoring against real competition ground truth. Publishes the task catalog, evaluation protocol and scoring code, while the runtime is operated as a Simreal service. - Magnitude: 2.5 out of 10 — Early signal - Stars: 82 total · +0 stars today, ≈ 16 by evening - Star trust: star growth looks organic - Category: Language models · Language: Python · License: Apache-2.0 · Created: 2026-09-21 · Last push: 2026-09-24 - GitHub: https://github.com/Simreal-AI/Simreal-MLBench · Homepage: https://simreal.co · Page: https://gitnova.dev/en/r/Simreal-AI/Simreal-MLBench ## Useful for - Recompute a published task score from saved reference counts to verify it - Validate the task catalog and protocol configuration with mleb validate - Check the submission id, artifact hash and reference snapshot behind each reported number ## Why it’s here - 0 stars so far today, about 16 expected by the end of the day. - The spike has held for 3 days in a row — not a one-off blip. - The repository is 5 days old and already has 82 stars. With less than two weeks of history, there's no usual pace to compare the spike against yet. - Top new repositories this week: #163. ## Star trust Star growth looks organic. Star-trust labels are heuristics based on the repository’s behavior, not a check of every stargazer. ## Numbers - Forks: 12 - Issues and pull requests: 0 - Watchers: 9 - Average over the last week: 16 per day - Usual pace: too little history (under two weeks) - Stars in the last hour (measured): 2 ## Stars per day, last 7 days (oldest → newest, today is partial) 2026-09-20 … 2026-09-26: 0, 1, 19, 26, 14, 22, 0 ## Spotted in now - Top new repositories this week: #163 ## More in this category 1. **ollaya-dev/ollaya** — 6.9 · Early signal · Language models · Rust · +0 stars today, ≈ 129 by evening A local runtime for decision models: pulls and serves classification and routing models behind a TypeSafe-compatible API. Like Ollama, but for models that return probabilities instead of text. Full card: https://gitnova.dev/en/r/ollaya-dev/ollaya.md 2. **NVIDIA/Model-Optimizer** — 6.4 · Breakout · Language models · Python · +0 stars today, ≈ 248 by evening NVIDIA library for compressing and accelerating models: quantization, pruning, distillation, NAS and speculative decoding with export to TensorRT-LLM, vLLM, SGLang. For ML engineers preparing models for deployment. Full card: https://gitnova.dev/en/r/NVIDIA/Model-Optimizer.md 3. **nokia-applied-research/AnyJev** — 5.3 · Early signal · Language models · Python · +0 stars today, ≈ 144 by evening Turns any LLM into a decision model with typed questions (choice, yes/no, score) that return real probabilities read from the next-token distribution, without training. L0 removes position and label-prior bias; L1 adds temperature… Full card: https://gitnova.dev/en/r/nokia-applied-research/AnyJev.md 4. **Badtheorylabs/interference-search** — 4.7 · Early signal · Language models · Python · +0 stars today, ≈ 77 by evening A search method that reasons over explicit states instead of a language model's linear transcript: many branches expand at once, duplicates merge, dead ends are dropped by a trained judge, and survivors advance together. The repo ships… Full card: https://gitnova.dev/en/r/Badtheorylabs/interference-search.md 5. **Liuziyu77/Valen** — 4.4 · Early signal · Language models · Python · +0 stars today, ≈ 57 by evening A multimodal decision model built on a Qwen3.5-0.8B/2B backbone: it takes text, images and video with an instruction and returns probabilities over supplied candidates without generating answer tokens. The repo includes model code, data… Full card: https://gitnova.dev/en/r/Liuziyu77/Valen.md --- Magnitude (0–10) measures how fast and how unusually interest in a repository is growing right now. It is not a quality score. Days are UTC. “So far today” is a fact; “expected by the end of the day” is a forecast. Summaries and use cases are written by an LLM (DeepSeek V4.1 Flash) from the README and may be inaccurate: verify specific claims (benchmarks, speed, hardware) in the repository itself. Data as of 2026-09-26 03:18 UTC, updated every 30 minutes.