Web appOpen in Telegram

PostAudio editing and speech generation just got a complete upgrade. Tencent has just released AuK, an open-source foundatio…

10 September 2026
S
Source Code Hub
Audio editing and speech generation just got a complete upgrade. Tencent has just released AuK, an open-source foundational model for speech generation and editing that truly stands out. This model is an all-in-one audio powerhouse. It offers zero-shot voice cloning, allowing users to replicate a voice from a short audio sample and generate new speech in that voice. It also features instruct-based TTS, where speech can be generated from text with specific vocal design instructions. What makes AuK particularly impressive is its comprehensive editing capabilities. It can perform speech content editing, letting users seamlessly add, delete, or replace words in existing spoken audio. It extends this to lyric content editing, maintaining the original melody while altering sung words, which is quite remarkable for music production. Beyond content, AuK delves into paralinguistic editing, enabling changes in emotion, vocal timbre (e.g., male to female voice), and even stripping away heavy accents. Users can also inject non-verbal sounds like sneezes or laughs, or convert normal speech directly into a whisper. For audio quality, AuK includes full speech enhancement, denoising, and multi-speaker separation, effectively cleaning up and clarifying audio. It also supports vocal extraction from music and general quality improvement. Looking at the benchmarks, AuK is a strong performer, beating many state-of-the-art models in various speech generation and editing tasks. Its word error rate is impressively low, and its speaker similarity scores are high. However, it's worth noting that features like de-accenting, while a massive improvement, might not always achieve 100% perfection, and some non-verbal sound insertions are decent but not entirely flawless. AuK is available in two versions: a 1.5 billion parameter base model for high-quality generation, and AuK-Flash, a distilled model for faster inference with near-teacher quality in just four steps. Both models should be easily runnable locally on systems with under 12GB of VRAM. This model represents a significant leap forward in audio manipulation, offering a suite of tools that previously required multiple different applications, all integrated into one coherent system. You can explore the AuK model and its capabilities on GitHub and Hugging Face: https://github.com/TencentARC/AuK What are your thoughts on this comprehensive audio editing and generation model? How do you see yourself using such a tool? For more cutting-edge AI news and developments, make sure to subscribe to t.me/iaosai.
1 · 118 ·

Nearby in the feed

SSource Code HubClass of '09 Access — v1.0.1 is out A quick follow-up patch for the Class of '09 Access game mod. The biggest issue from the initial v1.0.0 release was that sevSSource Code HubHouse Access 1.1.13 — Released The compatibility issue affecting the Steam version of House Party has been fixed. The mod no longer loses all functionality when
this message
SSource Code HubPlease help me click this invite link for Freebuff https://freebuff.com/get-started?ref=ref-55b32564-1521-4ad4-80e8-a75648129a1b&referrer=swarup+baralSSource Code HubHere is a new tts model, extremely lightweight. and one of if not the smallest neural tts models yet. This new model, called sanoTTS by ampixa, is truly somethi
SSource Code HubSource Code Hub@code_tg_web · channel · Tech
293subscribers116average post reach
Venue feed Open in Telegram

An open public feed from the search index ChatCrawler — “Google for public Telegram”; refreshed as the venue is crawled. Times are UTC.

Public content only, official Telegram API. About · FAQ · What we do not do · Remove a page · Catalog · Search · How we count