Source Code Hub
Audio editing and speech generation just got a complete upgrade. Tencent has just released AuK, an open-source foundational model for speech generation and editing that truly stands out.
This model is an all-in-one audio powerhouse. It offers zero-shot voice cloning, allowing users to replicate a voice from a short audio sample and generate new speech in that voice. It also features instruct-based TTS, where speech can be generated from text with specific vocal design instructions.
What makes AuK particularly impressive is its comprehensive editing capabilities. It can perform speech content editing, letting users seamlessly add, delete, or replace words in existing spoken audio. It extends this to lyric content editing, maintaining the original melody while altering sung words, which is quite remarkable for music production.
Beyond content, AuK delves into paralinguistic editing, enabling changes in emotion, vocal timbre (e.g., male to female voice), and even stripping away heavy accents. Users can also inject non-verbal sounds like sneezes or laughs, or convert normal speech directly into a whisper.
For audio quality, AuK includes full speech enhancement, denoising, and multi-speaker separation, effectively cleaning up and clarifying audio. It also supports vocal extraction from music and general quality improvement.
Looking at the benchmarks, AuK is a strong performer, beating many state-of-the-art models in various speech generation and editing tasks. Its word error rate is impressively low, and its speaker similarity scores are high. However, it's worth noting that features like de-accenting, while a massive improvement, might not always achieve 100% perfection, and some non-verbal sound insertions are decent but not entirely flawless.
AuK is available in two versions: a 1.5 billion parameter base model for high-quality generation, and AuK-Flash, a distilled model for faster inference with near-teacher quality in just four steps. Both models should be easily runnable locally on systems with under 12GB of VRAM.
This model represents a significant leap forward in audio manipulation, offering a suite of tools that previously required multiple different applications, all integrated into one coherent system.
You can explore the AuK model and its capabilities on GitHub and Hugging Face:
https://github.com/TencentARC/AuK
What are your thoughts on this comprehensive audio editing and generation model? How do you see yourself using such a tool?
For more cutting-edge AI news and developments, make sure to subscribe to t.me/iaosai.
1 · 118 ·