Released 2026-08-28 · general agent
Deployment evidence
Tencent Hy Team · 770B backbone + 10B MTP total · 49B backbone + 0.7B MTP active · 1M tokens
Best-fit scenarios
Frontier-scale coding, office, game-development and scientific agents where a large multi-GPU deployment is justified.
Not recommended
- Single-workstation deployment
- Treating the vendor's internal blind evaluation as an independent benchmark
- Promising 1M-context production concurrency before KV-cache and load tests
Long-horizon software engineering
Vendor-statedTencent positions Hy4 preview for understanding, planning, debugging and verifying long-running development tasks, including front-end interaction quality.
Capability sourceOffice analysis and artifact creation
Vendor-statedThe official release describes multi-file document, spreadsheet and presentation work involving data analysis, equations and financial models.
Capability sourceGame development and scientific research
Public evidenceTencent reports evaluations and internal expert testing across playable game prototypes, AI research, molecular dynamics, condensed-matter physics and pure mathematics; these are vendor results, not ChinaAPI reproductions.
Capability sourceOfficial launch / minimum
Evidence BOfficial framework matrix
SGLang's official Hy4 matrix lists the smallest published GPU-count path as the MXFP8 checkpoint on 4x B300 or 4x GB300 at TP4 and 262K configured context. It also lists 8x B200 at TP8. These are framework recipes, not a claim that four accelerators are an absolute minimum for every engine or the full 1M context. Primary recipe
200-person team
Evidence EBenchmark required
For the defined 200-seat profile, start load testing with one official-shape FP8 worker such as 8x B200 or 4x B300. If the service is operationally important, provision a second independent worker for maintenance and failure recovery. Tencent and the serving frameworks do not publish a Hy4 throughput result for this team profile, so this is ChinaAPI capacity-planning guidance, not a vendor recommendation.
Commercial API
Evidence BOfficial serving recipe
Tencent publishes vLLM and SGLang OpenAI-compatible recipes with native MTP speculative decoding, sparse-attention support and Hy4 tool/reasoning parsers. A public service should start with at least two independently deployable workers, admission control and separate short- and long-context SLO tests; redundancy is ChinaAPI operational guidance, while the serving commands are official framework recipes. Primary recipe
Commercial-use check
Apache-2.0
MaaS / hosted service: No model-specific MaaS restriction identified in the published Apache-2.0 license.
Attribution: Provide the license and required notices, preserve applicable attribution notices, and mark modified files; trademark rights are not granted.
Known limitations and open questions
- No ChinaAPI hardware reproduction or throughput benchmark
- The official framework matrix sizes common deployments at 131K or 262K context rather than demonstrating the advertised 1M maximum under production concurrency
- No official Hy4 throughput result is published for the 200-person profile used in this dossier
- The preview may reason longer than necessary and over-verify its own work, according to the vendor
- A 1M context window can make KV-cache demand the production bottleneck even when weights fit
