# lkarlslund/laya.cpp > Native C++ inference for Laya models with CUDA, Vulkan, and Apple Core ML backends, without Python. Supports tokenization, inference, JSON output, and an HTTP server with automatic request batching. - Magnitude: 2.1 out of 10 — Early signal - Stars: 82 total · +0 stars today, ≈ 12 by evening - Star trust: star growth looks organic - Category: Language models · Language: C++ · License: MIT · Created: 2026-09-20 · Last push: 2026-09-23 - GitHub: https://github.com/lkarlslund/laya.cpp · Page: https://gitnova.dev/en/r/lkarlslund/laya.cpp ## Useful for - Run local Laya inference on GPU without Python for processing JSON requests - Deploy an HTTP server with automatic batching to integrate Laya into existing systems - Build the project for CUDA, Vulkan, or Core ML on a specific platform ## Why it’s here - 0 stars so far today, about 12 expected by the end of the day. - The spike has held for 3 days in a row — not a one-off blip. - The repository is 4 days old and already has 82 stars. With less than two weeks of history, there's no usual pace to compare the spike against yet. ## Star trust Star growth looks organic. Star-trust labels are heuristics based on the repository’s behavior, not a check of every stargazer. ## Numbers - Forks: 8 - Issues and pull requests: 14 - Watchers: 2 - Average over the last week: 19 per day - Usual pace: too little history (under two weeks) - Stars in the last hour (measured): 1 - Latest release: r0002 (2026-09-23) ## Stars per day, last 5 days (oldest → newest, today is partial) 2026-09-20 … 2026-09-24: 16, 24, 26, 16, 0 ## Similar by description 1. **navaneethkrishnansuresh/Inference-Engineering** — 0.4 · Steady · Language models · Jupyter Notebook · +0 stars today, ≈ 1 by evening A hands-on course on LLM inference engineering covering what happens when a model serves a request: tokenization, prefill, decode, KV cache, scheduling and batching. Aimed at developers who want to understand the runtime side of model… Full card: https://gitnova.dev/en/r/navaneethkrishnansuresh/Inference-Engineering.md 2. **0xShug0/audio.cpp** — 2.0 · Steady · Generative media · C++ · +0 stars today, ≈ 23 by evening A native C++ inference engine for audio models built on ggml: TTS, speech recognition, VAD, voice cloning, music generation. Runs without Python on CPU, CUDA, ROCm, Vulkan and Metal. Full card: https://gitnova.dev/en/r/0xShug0/audio.cpp.md 3. **deepseek-ai/deepseek-recipe** — 0.9 · Steady · Language models · Rust · +0 stars today, ≈ 1 by evening A collection of Rust libraries and Python bindings that convert API requests in different formats (Messages, Chat Completions, Responses) into the Conversation format, encode them into prompts for DeepSeek V4/V4.1 models, and convert… Full card: https://gitnova.dev/en/r/deepseek-ai/deepseek-recipe.md 4. **paradigma-inc/limite-violetto** — 1.5 · Cooling · Language models · Python · +0 stars today, ≈ 6 by evening Repository provides a vLLM plugin for serving Limite 1B — Violetto, a model trained to solve difficult mathematical problems one at a time. Weights and tokenizer are hosted on Hugging Face. Full card: https://gitnova.dev/en/r/paradigma-inc/limite-violetto.md 5. **bojieli/ai-infra-book** — 2.6 · Cooling · Learning & lists · Python · +0 stars today, ≈ 88 by evening Open book "Understanding AI Infra" by Bojie Li: quantitative derivation of LLM inference and training system design from hardware constraints and model architecture. Includes full text, PDF, calculation tools and experiments. Full card: https://gitnova.dev/en/r/bojieli/ai-infra-book.md --- Magnitude (0–10) measures how fast and how unusually interest in a repository is growing right now. It is not a quality score. Days are UTC. “So far today” is a fact; “expected by the end of the day” is a forecast. Summaries and use cases are written by an LLM (DeepSeek V4.1 Flash) from the README and may be inaccurate: verify specific claims (benchmarks, speed, hardware) in the repository itself. Data as of 2026-09-24 03:33 UTC, updated every 30 minutes.