Tencent's AuK Bundles Voice Cloning and Speech Editing
The new open-weights model handles zero-shot TTS alongside enhancement and separation, aiming to be a broad speech toolkit rather than a single-purpose voice engine.

Tencent has released AuK, a zero-shot text-to-speech system that arrives with an unusually broad remit. Rather than focusing on a single task, the model spans voice cloning, speech editing, audio enhancement and source separation, according to its Hugging Face repository.
Zero-shot cloning is the headline capability: it lets the model reproduce a target voice from a short reference sample without task-specific fine-tuning. Paired with speech editing—modifying existing recordings rather than synthesizing from scratch—AuK positions itself closer to a production audio toolkit than a standalone TTS demo.
Why it matters
Most open speech models specialize, forcing developers to stitch together separate systems for synthesis, cleanup and separation. Consolidating those functions in one release could simplify pipelines for anyone building voice interfaces, dubbing tools or podcast production workflows.
A few practical caveats remain:
- AuK ships under an "other" license, so teams should read the terms before commercial use.
- Parameter count, supported languages and sample rate are not specified in the release record.
- Voice cloning capabilities raise the usual concerns about consent and misuse.
As with any open cloning model, the value will depend on audio quality and how permissive the license turns out to be. The weights and documentation are available now on Hugging Face.
Sources
- Visit
tencent/AuK
Hugging Face
More in Text → Speech
Breeze-TTS-2 Brings Open Voice Cloning to English
BreezeBlue's second-generation text-to-speech model pairs voice cloning with controllable direction, all under an open release on Hugging Face.

Audio8 debuts a compact 0.1B preview TTS model
The lightweight text-to-speech model brings zero-shot voice cloning to a footprint small enough to run almost anywhere.
NVIDIA opens Magpie TTS for multilingual voice agents
The company releases open weights for a low-latency, multilingual text-to-speech model aimed at real-time conversational systems.
0 comments
No comments yet. Be the first to weigh in.