OpenBMB/VoxCPM
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
About the project
VoxCPM2 is an open-source 2B-parameter tokenizer-free TTS system that synthesizes speech in 30 languages, designs voices from text descriptions, and clones voices from short reference clips with 48kHz output.
Useful for
- Synthesize speech from text in any of 30 languages via the Python API or CLI
- Clone a speaker's voice from a short reference clip to dub content
- Deploy an OpenAI-compatible TTS server on vLLM-Omni for multi-tenant workloads
README summarized by DeepSeek V4.1 Flash. Details may be inaccurate.
Why it’s trending
- 35 stars so far today, about 49 expected by the end of the day.
- About 9 forks a day — people are taking the code.
Stars per day
Bars are daily stars, the line is the usual pace. Red marks spike days.
Numbers
- Total stars
- 37,770
- Stars in a day
- 49
- Forks
- 4,289
- Issues and pull requests
- 397
- Watchers
- 168
- Language
- Python
- License
- Apache-2.0
- Latest release
- 2.0.3 · May 11, 2026
- Created
- September 16, 2025
- Last push
- September 2, 2026
Star trust
Growth looks organic: forks and discussion are in line with active projects, and stars arrive unevenly, the way people give them.
These are heuristics, not a verdict: we judge by the repository’s behavior, not by a list of stargazers.
Spotted in
- September 17, 2026GitHub Trending Python today: #18, +99 stars
- September 16, 2026GitHub Trending Python today: #18, +98 stars
- September 15, 2026GitHub Trending today: #14, +216 stars; GitHub Trending Python today: #7, +216 stars
- September 14, 2026GitHub Trending today: #14, +204 stars; GitHub Trending Python today: #7, +204 stars
Similar projects
-
7.9
mcncarl/jianying-headless
Local automation tool for Jianying Professional on macOS: builds native drafts from a JSON plan, edits copies of existing multi-track projects, and exports MP4 via the native engine on demand. Requires installed…
-
6.2
wide-trace/open-higgsfield
Open-source self-hosted studio for image and video generation across 32 models (8 image, 24 video) in one interface with a single prompt bar, gallery and per-model settings.
-
5.0
kuhnhomeuk-cell/procedural-film
An agent skill that turns a topic into a 30-second vertical film: all visuals are drawn on canvas in vanilla JavaScript and sound is synthesised with Web Audio, with no media assets. Output is a single HTML player plus…
-
4.9
XGEN-Labs/XGEN-JING
XGEN-JING is an egocentric interactive experience model built on MiniMax-H3 that generates first-person video and audio for navigation, object interaction, and dialogue from actions, reference images, and observation…
-
4.8
jamiepine/voicebox
A local-first AI voice studio: clone voices, generate speech with 7 TTS engines, and dictate into any app via a global hotkey. An open-source alternative to ElevenLabs and WisprFlow running entirely on your machine.
-
4.5
inikolax/remiqora
Local GPU-powered music generation studio unifying ACE-Step 1.5 and YuE2-3B in one Vue interface with a multitrack DAW, stem separation, MIDI transcription, and LoRA fine-tuning.