microsoft/VibeVoice
Open-Source Frontier Voice AI
About the project
Microsoft's family of open-source voice AI models: ASR for speech recognition (up to 60 minutes of audio in a single pass with diarization and timestamps) and TTS for long multi-speaker dialogue synthesis. Streaming and edge variants included.
Useful for
- Transcribe a one-hour meeting recording with speaker labels and timestamps
- Set up streaming speech recognition with custom hotwords for domain terms
- Generate a multi-speaker podcast or dialogue up to 90 minutes long
README summarized by DeepSeek V4.1 Flash. Details may be inaccurate.
Why it’s trending
- 6 stars so far today, about 16 expected by the end of the day.
- Over the last two days the pace is 1.7× that of the previous week and a half.
- GitHub Trending Python today: #13, +36 stars.
Stars per day
Bars are daily stars, the line is the usual pace. Red marks spike days.
Numbers
- Total stars
- 54,607
- Today
- 6 · ≈ 16 by evening
- Forks
- 6,147
- Issues and pull requests
- 430
- Watchers
- 269
- Language
- Python
- License
- MIT
- Created
- August 25, 2025
- Last push
- September 3, 2026
Star trust
Growth looks organic: forks and discussion are in line with active projects, and stars arrive unevenly, the way people give them.
These are heuristics, not a verdict: we judge by the repository’s behavior, not by a list of stargazers.
Spotted in
- October 2, 2026GitHub Trending Python today: #13, +36 stars
Similar by description
-
2.5
OpenWhispr/openwhispr
Cross-platform desktop voice-to-text dictation app that transcribes speech locally (Whisper, NVIDIA Parakeet) or in the cloud and pastes text into any app. It also transcribes meetings, manages notes, and connects to…
-
1.5
jankeesvw/omarchy-meeting-recorder
Records meetings on Omarchy: mic and system audio as separate tracks, transcribed locally with whisper.cpp, with speaker labels, chapters and a player.
-
1.5
k2-fsa/sherpa-onnx
A collection of offline speech processing tools built on next-gen Kaldi and onnxruntime: speech recognition and synthesis, speaker diarization and identification, VAD, speech enhancement, source separation, and keyword…
-
1.3
nobodywho-ooo/nobodywho
A Rust inference engine for running LLMs locally with bindings for Python, Kotlin, Swift, Flutter, React Native and Godot, plus speech-to-text and text-to-speech. It runs GGUF chat models offline on desktop and mobile…
-
2.5
Zackriya-Solutions/meetily
A local AI meeting assistant that records and transcribes audio in real time using Whisper/Parakeet and generates summaries via Ollama or other LLMs, without sending data to the cloud.
-
1.6
0xShug0/audio.cpp
A native C++ inference engine for audio models built on ggml: TTS, speech recognition, VAD, voice cloning, music generation. Runs without Python on CPU, CUDA, ROCm, Vulkan and Metal.