Seismograph

What’s gaining stars on GitHub right now

Simreal-AI/Simreal-MLBench

[Public preview] Externally scored agentic ML research benchmark: 60 tasks, real competition ground truth. Open protocol, operated evaluation.

Language modelsPython#agentic#benchmark#benchmarking#calibration#llm
2.5 Early signal Magnitude out of 10 — how fast interest is growing, not a quality score. Star growth looks organic. Data as of September 26, 2026, 02:37 UTC.
Open on Seismograph Open on GitHub

About the project

Open benchmark for evaluating ML research agents on 60 tasks with external scoring against real competition ground truth. Publishes the task catalog, evaluation protocol and scoring code, while the runtime is operated as a Simreal service.

Useful for

README summarized by DeepSeek V4.1 Flash. Details may be inaccurate.

Why it’s trending

Stars per day

01530September 20, 2026September 26, 2026

Bars are daily stars, the line is the usual pace. Red marks spike days.

Numbers

Total stars
81
Today
0 · ≈ 17 by evening
Forks
12
Issues and pull requests
0
Watchers
9
Language
Python
License
Apache-2.0
Created
September 21, 2026
Last push
September 24, 2026

Star trust

Growth looks organic: forks and discussion are in line with active projects, and stars arrive unevenly, the way people give them.

These are heuristics, not a verdict: we judge by the repository’s behavior, not by a list of stargazers.

Spotted in

Share

README badge

Paste it into your README — the badge shows the current magnitude and links to this page.

Seismograph: 2.5
[![Seismograph](https://gitnova.dev/badge/Simreal-AI/Simreal-MLBench.svg?lang=en)](https://gitnova.dev/en/r/Simreal-AI/Simreal-MLBench)

More in this category

  1. 6.9
    ollaya-dev/ollaya

    A local runtime for decision models: pulls and serves classification and routing models behind a TypeSafe-compatible API. Like Ollama, but for models that return probabilities instead of text.

    Early signalLanguage modelsRust+0 stars today, ≈ 134 by evening

  2. 6.4
    NVIDIA/Model-Optimizer

    NVIDIA library for compressing and accelerating models: quantization, pruning, distillation, NAS and speculative decoding with export to TensorRT-LLM, vLLM, SGLang. For ML engineers preparing models for deployment.

    BreakoutLanguage modelsPython+0 stars today, ≈ 255 by evening

  3. 5.4
    nokia-applied-research/AnyJev

    Turns any LLM into a decision model with typed questions (choice, yes/no, score) that return real probabilities read from the next-token distribution, without training. L0 removes position and label-prior bias; L1 adds…

    Early signalLanguage modelsPython+0 stars today, ≈ 152 by evening

  4. 4.7
    Badtheorylabs/interference-search

    A search method that reasons over explicit states instead of a language model's linear transcript: many branches expand at once, duplicates merge, dead ends are dropped by a trained judge, and survivors advance…

    Early signalLanguage modelsPython+0 stars today, ≈ 81 by evening

  5. 4.4
    Liuziyu77/Valen

    A multimodal decision model built on a Qwen3.5-0.8B/2B backbone: it takes text, images and video with an instruction and returns probabilities over supplied candidates without generating answer tokens. The repo…

    Early signalLanguage modelsPython+0 stars today, ≈ 59 by evening

  6. 4.3
    mizorewww/laya-mlx

    Native MLX runtime for Laya typed decision models: returns probabilities for choices, scores or truth values without text generation, running locally on Apple Silicon.

    CoolingLanguage modelsPython+0 stars today, ≈ 71 by evening