Seismograph

What’s gaining stars on GitHub right now

yifanzhang-pro/KLPO

Official Project Page for KL-Regularized Policy Optimization for Agentic Reinforcement Learning (KLPO)

Language modelsPython
5.5 Early signal Magnitude out of 10. Star growth looks organic. Data as of September 21, 2026.
Open on Seismograph Open on GitHub

About the project

KLPO is a critic-free, single-rollout reinforcement learning method for training agentic language models, using token regression and Monte Carlo KL estimation instead of value models or response groups. It provides the loss implementation, CPU tests, and a native Molt training integration.

Useful for

README summarized by DeepSeek V4.1 Flash. Details may be inaccurate.

Why it’s trending

Stars per day

050100September 13, 2026September 21, 2026

Bars are daily stars, the line is the usual pace. Red marks spike days.

Numbers

Total stars
97
Stars in a day
58
Forks
14
Issues and pull requests
0
Watchers
0
Language
Python
License
Apache-2.0
Created
September 19, 2026
Last push
September 21, 2026

Star trust

Growth looks organic: forks and discussion are in line with active projects, and stars arrive unevenly, the way people give them.

These are heuristics, not a verdict: we judge by the repository’s behavior, not by a list of stargazers.

Spotted in

Share

README badge

Paste it into your README — the badge shows the current magnitude and links to this page.

Seismograph: 5.5
[![Seismograph](https://gitnova.dev/badge/yifanzhang-pro/KLPO.svg?lang=en)](https://gitnova.dev/en/r/yifanzhang-pro/KLPO)

Similar projects

  1. 8.3
    mizorewww/laya-mlx

    Native MLX runtime for Laya typed decision models: returns probabilities for choices, scores or truth values without text generation, running locally on Apple Silicon.

    BreakoutLanguage modelsPython+919 stars in a day

  2. 7.2
    jaredpalmer/kev

    kev is a LoRA adapter with a small readout head on top of Qwen2.5-0.5B that answers many typed questions about a document in a single forward pass, returning calibrated probabilities instead of text.

    Early signalLanguage modelsPython+366 stars in a day

  3. 7.2
    bespokelabsai/nimble

    Nimble is a model and training recipe for fast typed decisions over text: given a flat schema of enum and boolean fields, it picks an answer and returns probabilities for each option. It targets routing, condition…

    Early signalLanguage modelsPython+330 stars in a day

  4. 7.0
    FareedKhan-dev/train-llm-from-scratch

    Educational project: a from-scratch PyTorch Transformer plus a full LLM training pipeline — from raw text through SFT, reward modeling, PPO/DPO/GRPO to chat, without transformers, trl or peft.

    BreakoutLanguage modelsPython+281 stars in a day

  5. 6.9
    mizorewww/laya-coreml

    Local port of the Laya model to Apple Core ML and Neural Engine: returns typed decisions (choice, score, yes/no) without token generation, with speed and energy benchmarks.

    Early signalLanguage modelsPython+240 stars in a day

  6. 6.8
    TheoLeeCJ/SemIf

    An open reproduction of the Jev-style semantic decision interface: reads typed option probabilities directly from a 4B model's logits without generating text. Runs on a single RTX 3090.

    PeakingLanguage modelsPython+322 stars in a day