Open-weight deployment dossier · verified 2026-09-12

Hy4 preview: self-hosting hardware, scenarios and commercial license

Frontier-scale coding, office, game-development and scientific agents where a large multi-GPU deployment is justified.

770B backbone + 10B MTP
total parameters
1M tokens
context window
Apache-2.0
published license

Released 2026-08-28 · general agent

Deployment evidence

Tencent Hy Team · 770B backbone + 10B MTP total · 49B backbone + 0.7B MTP active · 1M tokens

PrecisionOfficial BF16 and MXFP8 checkpoints; Hugging Face file metadata totals about 1,452.85 GiB and 757.88 GiB of safetensors respectively as reviewed on 2026-09-07Serving pathsvLLM, SGLang

Best-fit scenarios

Frontier-scale coding, office, game-development and scientific agents where a large multi-GPU deployment is justified.

Not recommended
  • Single-workstation deployment
  • Treating the vendor's internal blind evaluation as an independent benchmark
  • Promising 1M-context production concurrency before KV-cache and load tests

Long-horizon software engineering

Vendor-stated

Tencent positions Hy4 preview for understanding, planning, debugging and verifying long-running development tasks, including front-end interaction quality.

Capability source

Office analysis and artifact creation

Vendor-stated

The official release describes multi-file document, spreadsheet and presentation work involving data analysis, equations and financial models.

Capability source

Game development and scientific research

Public evidence

Tencent reports evaluations and internal expert testing across playable game prototypes, AI research, molecular dynamics, condensed-matter physics and pure mathematics; these are vendor results, not ChinaAPI reproductions.

Capability source

Official launch / minimum

Evidence B

Official framework matrix

SGLang's official Hy4 matrix lists the smallest published GPU-count path as the MXFP8 checkpoint on 4x B300 or 4x GB300 at TP4 and 262K configured context. It also lists 8x B200 at TP8. These are framework recipes, not a claim that four accelerators are an absolute minimum for every engine or the full 1M context. Primary recipe

200-person team

Evidence E

Benchmark required

For the defined 200-seat profile, start load testing with one official-shape FP8 worker such as 8x B200 or 4x B300. If the service is operationally important, provision a second independent worker for maintenance and failure recovery. Tencent and the serving frameworks do not publish a Hy4 throughput result for this team profile, so this is ChinaAPI capacity-planning guidance, not a vendor recommendation.

Commercial API

Evidence B

Official serving recipe

Tencent publishes vLLM and SGLang OpenAI-compatible recipes with native MTP speculative decoding, sparse-attention support and Hy4 tool/reasoning parsers. A public service should start with at least two independently deployable workers, admission control and separate short- and long-context SLO tests; redundancy is ChinaAPI operational guidance, while the serving commands are official framework recipes. Primary recipe

Commercial-use check

Apache-2.0

MaaS / hosted service: No model-specific MaaS restriction identified in the published Apache-2.0 license.

Attribution: Provide the license and required notices, preserve applicable attribution notices, and mark modified files; trademark rights are not granted.

Read the primary license text

Known limitations and open questions
  • No ChinaAPI hardware reproduction or throughput benchmark
  • The official framework matrix sizes common deployments at 131K or 262K context rather than demonstrating the advertised 1M maximum under production concurrency
  • No official Hy4 throughput result is published for the 200-person profile used in this dossier
  • The preview may reason longer than necessary and over-verify its own work, according to the vendor
  • A 1M context window can make KV-cache demand the production bottleneck even when weights fit

Evidence boundaries

A launch shape is not a production SLA.

The minimum tier records the smallest official or inference-framework configuration we found. The 200-person and commercial API tiers still require measurements against real prompt length, output length, concurrency, latency and redundancy targets.

Read the complete index methodology

  1. AChinaAPI reproduced
  2. Binference-framework official validated recipe
  3. Cmodel-vendor documented configuration
  4. Dthird-party reproduction
  5. Ecapacity estimate only

Compare before deploying

Review every verified model or compare hosted access.

The directory keeps model selection separate from the evidence and capacity details on this page.