Released 2026-09-09 · speech generation editing
Deployment evidence
Tencent Hunyuan · 1.5B family model total · Not separately reported active · Not applicable
Best-fit scenarios
Instruction-driven speech generation, voice cloning, editing, enhancement and source separation where a self-hosted unified audio workflow is more important than a managed API.
Not recommended
- Treating vendor-published quality or speed results as independent or ChinaAPI-reproduced evidence
- Sizing hardware from the 6.76 GB AuK repository alone while omitting the external Qwen2.5-Omni-3B encoder, runtime state and activations
- Exposing voice-cloning or speech-editing features without consent, impersonation-abuse and content-handling controls
- Promising streaming output, high concurrency or an SLA before workload-specific benchmarks
Zero-shot and described-voice TTS
Vendor-statedTencent documents zero-shot TTS from reference audio and instruction-based TTS from a voice description through the same natural-language interface.
Capability sourceSpeech and lyric editing
Vendor-statedThe official task set covers replacing, inserting or removing spoken content, plus rewriting lyrics while preserving the recording's melody and voice.
Capability sourceEnhancement and source separation
Vendor-statedThe official recipes cover denoising, dereverberation, speech separation, singing-voice extraction and target-speaker extraction from natural-language instructions.
Capability sourceLower-latency research with AuK-Flash
Public evidenceThe technical report states that the distilled four-step AuK-Flash achieves a 4.5x wall-clock speedup over AuK under matched conditions. This is a vendor-published experiment, not a ChinaAPI reproduction.
Capability sourceOfficial launch / minimum
Evidence COfficial single gpu launch path not hardware minimum
Tencent documents launching either AuK base or AuK-Flash on cuda:0. It does not name the accelerator or memory floor, and the 6.76 GB model repository excludes the separately downloaded Qwen2.5-Omni-3B encoder plus runtime headroom, so this proves a single-GPU code path rather than a minimum hardware configuration. Primary recipe
200-person team
Evidence EBenchmark required
Start with AuK-Flash and benchmark real clip durations, reference-audio use, request mix and latency. SGLang-Omni documents default maximum batches of 8 for conditioning, 16 for DiT sampling and 4 for equal-length VAE decode, but publishes no capacity result for this 200-seat profile; an operational service also needs a second failure-domain worker.
Commercial API
Evidence BServing recipe benchmark and safety review required
SGLang-Omni provides /v1/audio/speech and /generate routes plus dynamic batching, but caps target duration at 30 seconds by default and does not implement incremental audio streaming. A public service still needs measured capacity, isolation, admission control, storage policy, consent and impersonation-abuse safeguards, and redundant workers. Primary recipe
Commercial-use check
MIT
MaaS / hosted service: No model-specific MaaS restriction was identified in the published MIT license.
Attribution: Include Tencent's copyright notice and the MIT permission notice in all copies or substantial portions; the software and weights are provided without warranty.
Known limitations and open questions
- No ChinaAPI reproduction of the downloadable checkpoints, quality claims or 4.5x speed result
- Neither Tencent nor SGLang-Omni publishes an accelerator-memory minimum or throughput result for the defined team and commercial profiles
- The downloadable AuK repository is 6.76 GB, but Qwen2.5-Omni-3B is an additional required encoder download and runtime memory is not included
- SGLang-Omni defaults to a 30-second duration cap and does not provide incremental audio streaming
- Prompt enhancement and optional transcription can introduce separate LLM, ASR, credential and data-handling dependencies
