leemysw/yovoice
Open-source voice creation for macOS and Windows. Local TTS, voice cloning, and emotion control — no cloud APIs or per-character fees.
About the project
A local macOS and Windows tool that turns text into speech with voice cloning and emotion control, without cloud APIs or per-character fees. Supports several TTS models (IndexTTS, VoxCPM2, OmniVoice, Qwen3-TTS) via audio.cpp.
Useful for
- Generate speech locally from text with voice cloning from a reference clip
- Tune emotion and delivery for narration or voiceover without cloud services
- Have an AI agent use the CLI to voice a text file with a chosen voice
README summarized by DeepSeek V4.1 Flash. Details may be inaccurate.
Why it’s trending
- 0 stars so far today, about 13 expected by the end of the day.
- Over the last two days the pace is 2.4× that of the previous week and a half.
- The spike has held for 3 days in a row — not a one-off blip.
- The repository is 4 days old and already has 58 stars.
Stars per day
Bars are daily stars, the line is the usual pace. Red marks spike days.
Numbers
- Total stars
- 58
- Stars in a day
- 13
- Forks
- 8
- Issues and pull requests
- 3
- Watchers
- 1
- Language
- TypeScript
- License
- Apache-2.0
- Latest release
- v0.1.2 · September 18, 2026
- Created
- September 15, 2026
- Last push
- September 18, 2026
Star trust
Growth looks organic: forks and discussion are in line with active projects, and stars arrive unevenly, the way people give them.
These are heuristics, not a verdict: we judge by the repository’s behavior, not by a list of stargazers.
Spotted in
- September 19, 2026Top new repositories this week: #199
Similar projects
-
7.6
mcncarl/jianying-headless
Local automation tool for Jianying Professional on macOS: builds native drafts from a JSON plan, edits copies of existing multi-track projects, and exports MP4 via the native engine on demand. Requires installed…
-
6.3
wide-trace/open-higgsfield
Open-source self-hosted studio for image and video generation across 32 models (8 image, 24 video) in one interface with a single prompt bar, gallery and per-model settings.
-
5.3
XGEN-Labs/XGEN-JING
XGEN-JING is an egocentric interactive experience model built on MiniMax-H3 that generates first-person video and audio for navigation, object interaction, and dialogue from actions, reference images, and observation…
-
4.4
inikolax/remiqora
Local GPU-powered music generation studio unifying ACE-Step 1.5 and YuE2-3B in one Vue interface with a multitrack DAW, stem separation, MIDI transcription, and LoRA fine-tuning.
-
4.2
kuhnhomeuk-cell/procedural-film
An agent skill that turns a topic into a 30-second vertical film: all visuals are drawn on canvas in vanilla JavaScript and sound is synthesised with Web Audio, with no media assets. Output is a single HTML player plus…
-
4.1
jtydhr88/music-composition-skills
A set of 29 agent skills for Claude Code and OpenAI Codex that help compose and arrange popular music: from a brief the agent builds an ARR-SPEC with key, tempo, section map, harmony and energy curve. That spec is then…