0xShug0/audio.cpp
An all-in-one, pure C++ inference engine for audio models, powered by ggml. Supports TTS, STT, VAD, voice conversion, music generation, and more, with highly optimized performance. No Python dependency.
About the project
A native C++ inference engine for audio models built on ggml: TTS, speech recognition, VAD, voice cloning, music generation. Runs without Python on CPU, CUDA, ROCm, Vulkan and Metal.
Useful for
- Run a local TTS server without Python dependencies on your own hardware
- Compare several GGUF ASR model variants side by side in the Arena UI
- Deploy speech recognition on an edge device via the CPU backend
README summarized by DeepSeek V4.1 Flash. Details may be inaccurate.
Why it’s trending
- 32 stars today.
- GitHub Trending C++ today: #4, +42 stars.
Stars per day
Bars are daily stars, the line is the usual pace. Red marks spike days.
Numbers
- Total stars
- 2,677
- Stars in a day
- 32
- Forks
- 304
- Issues and pull requests
- 526
- Watchers
- 27
- Language
- C++
- Latest release
- v0.7.4 · September 13, 2026
- Created
- June 23, 2026
- Last push
- September 13, 2026
Star trust
Growth looks organic: forks and discussion are in line with active projects, and stars arrive unevenly, the way people give them.
These are heuristics, not a verdict: we judge by the repository’s behavior, not by a list of stargazers.
Spotted in
- September 14, 2026GitHub Trending C++ today: #4, +40 stars
- September 13, 2026GitHub Trending C++ today: #1, +42 stars
Similar projects
-
7.3
multimodal-art-projection/YuE
YuE2 is a music generation model: from lyrics and a style prompt it builds an editable melody-and-chord plan, then renders a full song with vocals and accompaniment. It supports zero-shot covers and agentic composition…
-
5.9
debpalash/VoiceStudio
A fully local, open-source ElevenLabs alternative for voice cloning, voice design, video dubbing, dictation, transcription and audiobook creation on your own hardware, with no account or API key. Supports 16 TTS and 11…
-
5.9
eternityspring/reelbench-skills
A set of Claude Code and Codex skills for video work: video-shots breaks a finished clip into a shot-by-shot analysis table (duration, shot size, camera moves), video-sync renders video with a side panel of shot info…
-
4.6
Colafornia/short-video-generator-AI
Open-source tool that turns long YouTube videos into vertical short clips: it detects highlights and adds subtitles, translation and voiceover. A free alternative to paid SaaS like OpusClip for content creators.
-
4.5
k2-fsa/OmniVoice
Massively multilingual zero-shot TTS model supporting 600+ languages with voice cloning and voice attribute control. Built on a diffusion language model-style architecture with fast inference.
-
4.5
govin-ai/dunhuang-aura-skill
An AI Skill for generating and editing commercial visuals in the Dunhuang mineral aesthetic, usable in Codex, Claude Code and other SKILL.md-compatible agent environments.