Released 2026-07-06 · general agent
Deployment evidence
Tencent Hunyuan · 295B backbone plus 3.8B MTP layer total · 21B backbone per token active · 256K tokens
Best-fit scenarios
Large private coding, productivity and tool-using agents that can support an official eight-accelerator tensor-parallel worker.
Not recommended
- Planning from the 21B active count instead of loading the full checkpoint
- Treating an eight-GPU launch command as a production capacity result
Coding and productivity agents
Vendor-statedTencent positions Hy3 for coding, productivity, reasoning and agent tasks. Published capability results are vendor evaluations, not ChinaAPI reproductions.
Capability sourceLong-context tool workflows
Vendor-statedThe official configuration supports 256K context and documents tool and reasoning parsers for vLLM and SGLang.
Capability sourcePrivate OpenAI-compatible endpoint
ChinaAPI inferenceThe official serving commands make an operator-controlled endpoint plausible, but capacity still depends on prompt length and concurrency.
Official launch / minimum
Evidence COfficial vendor floor
The vendor recommends eight H20-3e GPUs or accelerators with more memory for the documented vLLM and SGLang TP8 path. This proves an official launch shape, not an SLA. Primary recipe
200-person team
Evidence EBenchmark required
Benchmark one TP8 worker against the 200-seat profile, then add an independent worker if maintenance and failure recovery must preserve service.
Commercial API
Evidence COfficial framework paths
Official vLLM and SGLang OpenAI-compatible paths exist; a public service still needs redundant TP8 workers, admission control and pinned parsers. Primary recipe
Commercial-use check
Apache-2.0
MaaS / hosted service: No model-specific MaaS restriction identified in Apache-2.0.
Attribution: Provide the license and required notices, preserve attribution notices, and mark modified files; trademark rights are not granted.
Known limitations and open questions
- No ChinaAPI hardware or quality reproduction
- The official eight-GPU shape does not state concurrency or latency
- Full-context KV-cache demand must be measured separately
