Tencent-Hunyuan/AuK
AuK: An Open-Source Foundational Model for Speech Generation and Editing
About the project
AuK is an open-source 1.5B foundation model for speech generation and editing via natural-language instructions: zero-shot TTS, content and acoustic editing, emotion, enhancement and source separation.
Useful for
- Synthesize speech in a reference voice from target text
- Change emotion or timbre of a recording while keeping content and voice
- Extract a target speaker or vocals from a mixed recording
README summarized by DeepSeek V4.1 Flash. Details may be inaccurate.
Why it’s trending
- 20 stars so far today, about 24 expected by the end of the day.
- Interest is fading: the two-day pace is 31% of the previous week and a half.
- The repository is 30 days old and already has 1,111 stars.
Stars per day
Bars are daily stars, the line is the usual pace. Red marks spike days.
Numbers
- Total stars
- 1,111
- Stars in a day
- 24
- Forks
- 81
- Issues and pull requests
- 20
- Watchers
- 6
- Language
- Python
- Created
- August 19, 2026
- Last push
- September 16, 2026
Star trust
Growth looks organic: forks and discussion are in line with active projects, and stars arrive unevenly, the way people give them.
These are heuristics, not a verdict: we judge by the repository’s behavior, not by a list of stargazers.
Spotted in
- September 17, 2026Top new repositories this month: #73
- September 16, 2026Top new repositories this month: #81
- September 15, 2026Top new repositories this month: #89
- September 14, 2026Top new repositories this month: #97
Similar projects
-
7.9
mcncarl/jianying-headless
Local automation tool for Jianying Professional on macOS: builds native drafts from a JSON plan, edits copies of existing multi-track projects, and exports MP4 via the native engine on demand. Requires installed…
-
6.1
wide-trace/open-higgsfield
Open-source self-hosted studio for image and video generation across 32 models (8 image, 24 video) in one interface with a single prompt bar, gallery and per-model settings.
-
4.9
kuhnhomeuk-cell/procedural-film
An agent skill that turns a topic into a 30-second vertical film: all visuals are drawn on canvas in vanilla JavaScript and sound is synthesised with Web Audio, with no media assets. Output is a single HTML player plus…
-
4.9
XGEN-Labs/XGEN-JING
XGEN-JING is an egocentric interactive experience model built on MiniMax-H3 that generates first-person video and audio for navigation, object interaction, and dialogue from actions, reference images, and observation…
-
4.8
jamiepine/voicebox
A local-first AI voice studio: clone voices, generate speech with 7 TTS engines, and dictate into any app via a global hotkey. An open-source alternative to ElevenLabs and WisprFlow running entirely on your machine.
-
4.4
inikolax/remiqora
Local GPU-powered music generation studio unifying ACE-Step 1.5 and YuE2-3B in one Vue interface with a multitrack DAW, stem separation, MIDI transcription, and LoRA fine-tuning.