ArtificialAnalysis/aa-agentperf-local
Benchmark local LLM serving by replaying real agent trajectories
About the project
A tool from Artificial Analysis that measures how fast a local LLM server serves an AI agent by replaying recorded agent conversations and reporting throughput and latency, without judging output quality.
Useful for
- Benchmark throughput and latency of your own llama.cpp, vLLM, SGLang or LM Studio server on recorded agent trajectories
- Compare speed across models and hardware via managed-run with pinned recipes and SHA-256 verification
- Check server setup with the quick synthetic aa-mini-v1 replay at 8192 context
README summarized by DeepSeek V4.1 Flash. Details may be inaccurate.
Why it’s trending
- 0 stars so far today, about 7 expected by the end of the day.
- The spike has held for 3 days in a row — not a one-off blip.
- The repository is 6 days old and already has 70 stars. With less than two weeks of history, there's no usual pace to compare the spike against yet.
- Top new repositories this week: #160.
- Hacker News: “Gemini 4 Argon (High): Intelligence, Performance and Price Analysis” — 111 points, 1 day ago.
Stars per day
Bars are daily stars, the line is the usual pace. Red marks spike days.
Numbers
- Total stars
- 70
- Today
- 0 · ≈ 7 by evening
- Forks
- 1
- Issues and pull requests
- 33
- Watchers
- 0
- Language
- Python
- License
- Apache-2.0
- Latest release
- v0.3.3 · October 1, 2026
- Created
- September 26, 2026
- Last push
- October 1, 2026
Star trust
Growth looks organic: forks and discussion are in line with active projects, and stars arrive unevenly, the way people give them.
These are heuristics, not a verdict: we judge by the repository’s behavior, not by a list of stargazers.
Hacker News discussions
- Gemini 4 Argon (High): Intelligence, Performance and Price Analysis111 points, September 30, 2026
- GPT-6.1 Sol replaces GPT-6 Sol after just 7 days, with near-Astra intelligence80 points, September 30, 2026
- Sonnet 5.5 scores just behind Opus 5.5 on Artificial Analysis Intelligence Index11 points, September 28, 2026
- Gemini 4 Argon: Google is back among the top three labs in intelligence achieved6 points, October 1, 2026
Spotted in
- October 2, 2026Top new repositories this week: #160
More in this category
-
8.6
Niko1221/Strata
A C++ local inference engine that runs the 125B MoE model Qwen3.8-Flash-Next on a regular PC with one NVIDIA GPU (12-24 GB) and 64 GB RAM, exposing an OpenAI/Anthropic-compatible API on localhost.
-
6.4
MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks-TensorFold
Scripts to serve GLM-5.3-Flash on two NVIDIA DGX Sparks via TensorFold with an OpenAI-compatible API, 1M-token context, and image/video input.
-
5.8
ninjahawk/livenerf
A benchmark for tracking whether a frontier model quietly degrades after release: it runs a frozen question panel daily and statistically measures drift in accuracy and token counts against the launch-week baseline.
-
5.7
VectifyAI/PageIndex
A vectorless RAG engine that builds a hierarchical tree index of a document and retrieves by LLM reasoning over that tree, the way a human navigates a report. Aimed at long professional PDFs such as financial, legal…
-
5.3
magnitudedev/magnitude
Open source inference engine for consumer hardware that profiles your machine, recommends the best local models, then downloads, tunes, and runs them, with one-click connection to AI agents.
-
4.8
ashhart/TensorFold
LLM inference server for Apple Silicon (MLX) and NVIDIA GPUs with an OpenAI-compatible API and exact speculative decoding. Supports Nemotron, Qwen3.8, GLM-5.3 and Gemma 4 families with per-family kernels and draft…