腾讯混元开源语音模型 AuK,统一语音生成与编辑并推出 4 步推理的 AuK-Flash

Tencent Hy · @TencentHunyuan · X·2026-09-10 18:33·33分钟前
AI 导读

腾讯混元发布开源语音基础模型 AuK,支持统一语音生成与编辑,可通过自然语言指令加参考音频实现零样本 TTS、指令控制生成、内容与音色/风格/情感编辑、去口音、语速音调控制、增强去噪及多说话人和音乐分离。

Tencent Hy@TencentHunyuan
58AI 编辑部评分,满分 100

腾讯混元开源语音模型 AuK,统一语音生成与编辑并推出 4 步推理的 AuK-Flash

2026-09-10 18:33· 33分钟前
AI 导读

腾讯混元发布开源语音基础模型 AuK,支持统一语音生成与编辑,可通过自然语言指令加参考音频实现零样本 TTS、指令控制生成、内容与音色/风格/情感编辑、去口音、语速音调控制、增强去噪及多说话人和音乐分离。

🚀 AuK is officially here. Nano banana🍌 for audio

An open-source foundation model for unified speech generation and editing. Natural-language instructions + reference audio. One interface. Zero-shot TTS. Instruction-controlled generation. Content editing. Whisper-conversion. De-accent. Timbre/style/emotion edit. Speed/Pitch control. Enhancement, denoising, multi-speaker and music separation. Also releasing AuK-Flash: 4-step inference. ~4.5× faster under matched conditions. Code, weights, and demo are live. Try it and share your feedback.

🤗 Paper & upvote: https://huggingface.co/papers/2609.08936 ⭐ GitHub & star: https://github.com/Tencent-Hunyuan/AuK