Open-weight deployment dossier · verified 2026-09-12

AuK / AuK-Flash: self-hosting hardware, scenarios and commercial license

Instruction-driven speech generation, voice cloning, editing, enhancement and source separation where a self-hosted unified audio workflow is more important than a managed API.

1.5B family model
total parameters
Not applicable
context window
MIT
published license

Released 2026-09-09 · speech generation editing

Deployment evidence

Tencent Hunyuan · 1.5B family model total · Not separately reported active · Not applicable

PrecisionEach official AuK or AuK-Flash repository totals 6.76 GB: a 6.12 GB BF16 DiT checkpoint plus a shared 637 MB VAE; the external Qwen2.5-Omni-3B encoder is a separate downloadServing pathsOfficial PyTorch CLI and Python API, ComfyUI, SGLang-Omni

Best-fit scenarios

Instruction-driven speech generation, voice cloning, editing, enhancement and source separation where a self-hosted unified audio workflow is more important than a managed API.

Not recommended
  • Treating vendor-published quality or speed results as independent or ChinaAPI-reproduced evidence
  • Sizing hardware from the 6.76 GB AuK repository alone while omitting the external Qwen2.5-Omni-3B encoder, runtime state and activations
  • Exposing voice-cloning or speech-editing features without consent, impersonation-abuse and content-handling controls
  • Promising streaming output, high concurrency or an SLA before workload-specific benchmarks

Zero-shot and described-voice TTS

Vendor-stated

Tencent documents zero-shot TTS from reference audio and instruction-based TTS from a voice description through the same natural-language interface.

Capability source

Speech and lyric editing

Vendor-stated

The official task set covers replacing, inserting or removing spoken content, plus rewriting lyrics while preserving the recording's melody and voice.

Capability source

Enhancement and source separation

Vendor-stated

The official recipes cover denoising, dereverberation, speech separation, singing-voice extraction and target-speaker extraction from natural-language instructions.

Capability source

Lower-latency research with AuK-Flash

Public evidence

The technical report states that the distilled four-step AuK-Flash achieves a 4.5x wall-clock speedup over AuK under matched conditions. This is a vendor-published experiment, not a ChinaAPI reproduction.

Capability source

Official launch / minimum

Evidence C

Official single gpu launch path not hardware minimum

Tencent documents launching either AuK base or AuK-Flash on cuda:0. It does not name the accelerator or memory floor, and the 6.76 GB model repository excludes the separately downloaded Qwen2.5-Omni-3B encoder plus runtime headroom, so this proves a single-GPU code path rather than a minimum hardware configuration. Primary recipe

200-person team

Evidence E

Benchmark required

Start with AuK-Flash and benchmark real clip durations, reference-audio use, request mix and latency. SGLang-Omni documents default maximum batches of 8 for conditioning, 16 for DiT sampling and 4 for equal-length VAE decode, but publishes no capacity result for this 200-seat profile; an operational service also needs a second failure-domain worker.

Commercial API

Evidence B

Serving recipe benchmark and safety review required

SGLang-Omni provides /v1/audio/speech and /generate routes plus dynamic batching, but caps target duration at 30 seconds by default and does not implement incremental audio streaming. A public service still needs measured capacity, isolation, admission control, storage policy, consent and impersonation-abuse safeguards, and redundant workers. Primary recipe

Commercial-use check

MIT

MaaS / hosted service: No model-specific MaaS restriction was identified in the published MIT license.

Attribution: Include Tencent's copyright notice and the MIT permission notice in all copies or substantial portions; the software and weights are provided without warranty.

Read the primary license text

Known limitations and open questions
  • No ChinaAPI reproduction of the downloadable checkpoints, quality claims or 4.5x speed result
  • Neither Tencent nor SGLang-Omni publishes an accelerator-memory minimum or throughput result for the defined team and commercial profiles
  • The downloadable AuK repository is 6.76 GB, but Qwen2.5-Omni-3B is an additional required encoder download and runtime memory is not included
  • SGLang-Omni defaults to a 30-second duration cap and does not provide incremental audio streaming
  • Prompt enhancement and optional transcription can introduce separate LLM, ASR, credential and data-handling dependencies

Evidence boundaries

A launch shape is not a production SLA.

The minimum tier records the smallest official or inference-framework configuration we found. The 200-person and commercial API tiers still require measurements against real prompt length, output length, concurrency, latency and redundancy targets.

Read the complete index methodology

  1. AChinaAPI reproduced
  2. Binference-framework official validated recipe
  3. Cmodel-vendor documented configuration
  4. Dthird-party reproduction
  5. Ecapacity estimate only

Compare before deploying

Review every verified model or compare hosted access.

The directory keeps model selection separate from the evidence and capacity details on this page.