# NanmiCoder/jev-arena > A web tool for comparing two LLMs on the same labeling task: speed, cost and accuracy on CSV/Excel comment data, with run recording and offline reports. - Magnitude: 0.8 out of 10 — Cooling - Stars: 97 total · +0 stars today, ≈ 2 by evening - Star trust: star growth looks organic - Category: Language models · Language: JavaScript · License: MIT · Created: 2026-09-19 · Last push: 2026-09-20 - GitHub: https://github.com/NanmiCoder/jev-arena · Homepage: https://nanmicoder.github.io/jev-arena/ · Page: https://gitnova.dev/en/r/NanmiCoder/jev-arena ## Useful for - Compare Jev and DeepSeek speed and cost on your own CSV of comments - Check two models' labeling accuracy on relevance, sentiment and intent - Generate an offline HTML report from a finished run without API calls ## Why it’s here - 0 stars so far today, about 2 expected by the end of the day. - The repository is 5 days old and already has 97 stars. With less than two weeks of history, there's no usual pace to compare the spike against yet. - Top new repositories this week: #172. ## Star trust Star growth looks organic. Star-trust labels are heuristics based on the repository’s behavior, not a check of every stargazer. ## Numbers - Forks: 9 - Issues and pull requests: 1 - Watchers: 0 - Average over the last week: 16 per day - Usual pace: too little history (under two weeks) - Stars in the last hour (measured): -1 ## Stars per day, last 12 days (oldest → newest, today is partial) 2026-09-13 … 2026-09-24: 0, 0, 0, 0, 0, 0, 10, 33, 36, 16, 2, 0 ## Spotted in now - Top new repositories this week: #172 ## Similar by description 1. **zhulinchng/jevper** — 2.3 · Steady · Language models · Python · +0 stars today, ≈ 8 by evening A classification wrapper over OpenAI-compatible clients: instead of prose it returns typed questions (noul, choice, score) with probabilities and confidence. Works with any client exposing responses.create or chat.completions.create,… Full card: https://gitnova.dev/en/r/zhulinchng/jevper.md 2. **fstandhartinger/jevbench** — 2.4 · Early signal · Language models · Python · +0 stars today, ≈ 17 by evening A benchmark for Jev-class decision models: the model receives state and a rubric, returns a typed answer with probabilities. It scores accuracy, calibration, speed, and cost. Full card: https://gitnova.dev/en/r/fstandhartinger/jevbench.md --- Magnitude (0–10) measures how fast and how unusually interest in a repository is growing right now. It is not a quality score. Days are UTC. “So far today” is a fact; “expected by the end of the day” is a forecast. Summaries and use cases are written by an LLM (DeepSeek V4.1 Flash) from the README and may be inaccurate: verify specific claims (benchmarks, speed, hardware) in the repository itself. Data as of 2026-09-24 02:50 UTC, updated every 30 minutes.