yifanzhang-pro/KLPO
Official Project Page for KL-Regularized Policy Optimization for Agentic Reinforcement Learning (KLPO)
About the project
KLPO is a critic-free, single-rollout reinforcement learning method for training agentic language models, using token regression and Monte Carlo KL estimation instead of value models or response groups. It provides the loss implementation, CPU tests, and a native Molt training integration.
Useful for
- Train an agentic LLM with KLPO loss using the provided Molt launcher
- Run the CPU toy example to verify all eight route/estimator combinations
- Compare MC-KL, TopK-KL, Binary KL and Full KL estimators on a small policy
README summarized by DeepSeek V4.1 Flash. Details may be inaccurate.
Why it’s trending
- 11 stars so far today, about 58 expected by the end of the day.
- The spike has held for 2 days in a row — not a one-off blip.
- The repository is 2 days old and already has 97 stars.
- Top new repositories this week: #179.
- About 51 forks a day — people are taking the code.
- Recent forks include notable developers: @OctopusTakopi (215 followers).
Stars per day
Bars are daily stars, the line is the usual pace. Red marks spike days.
Numbers
- Total stars
- 97
- Stars in a day
- 58
- Forks
- 14
- Issues and pull requests
- 0
- Watchers
- 0
- Language
- Python
- License
- Apache-2.0
- Created
- September 19, 2026
- Last push
- September 21, 2026
Star trust
Growth looks organic: forks and discussion are in line with active projects, and stars arrive unevenly, the way people give them.
These are heuristics, not a verdict: we judge by the repository’s behavior, not by a list of stargazers.
Spotted in
- September 21, 2026Top new repositories this week: #179
Similar projects
-
8.3
mizorewww/laya-mlx
Native MLX runtime for Laya typed decision models: returns probabilities for choices, scores or truth values without text generation, running locally on Apple Silicon.
-
7.2
jaredpalmer/kev
kev is a LoRA adapter with a small readout head on top of Qwen2.5-0.5B that answers many typed questions about a document in a single forward pass, returning calibrated probabilities instead of text.
-
7.2
bespokelabsai/nimble
Nimble is a model and training recipe for fast typed decisions over text: given a flat schema of enum and boolean fields, it picks an answer and returns probabilities for each option. It targets routing, condition…
-
7.0
FareedKhan-dev/train-llm-from-scratch
Educational project: a from-scratch PyTorch Transformer plus a full LLM training pipeline — from raw text through SFT, reward modeling, PPO/DPO/GRPO to chat, without transformers, trl or peft.
-
6.9
mizorewww/laya-coreml
Local port of the Laya model to Apple Core ML and Neural Engine: returns typed decisions (choice, score, yes/no) without token generation, with speed and energy benchmarks.
-
6.8
TheoLeeCJ/SemIf
An open reproduction of the Jev-style semantic decision interface: reads typed option probabilities directly from a 4B model's logits without generating text. Runs on a single RTX 3090.