# Hy4 preview: Self-Hosting Hardware, Scenarios and Commercial License

> Frontier-scale coding, office, game-development and scientific agents where a large multi-GPU deployment is justified.

- Verified: 2026-09-12
- Released: 2026-08-28
- Canonical: https://chinaapi.ai/open-model-deployment/hy4-preview/
- Back to directory: https://chinaapi.ai/open-model-deployment/
- Machine-readable dataset: https://chinaapi.ai/data/open-model-deployment.json

## Model facts

- Vendor: Tencent Hy Team
- Parameters: 770B backbone + 10B MTP total; 49B backbone + 0.7B MTP active
- Native precision: Official BF16 and MXFP8 checkpoints; Hugging Face file metadata totals about 1,452.85 GiB and 757.88 GiB of safetensors respectively as reviewed on 2026-09-07
- Context: 1M tokens
- Category: General and agentic foundation models
- Modalities: text → text
- Serving paths: vLLM, SGLang
- Official weights: Original: https://huggingface.co/tencent/Hy4-preview?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment; Official FP8: https://huggingface.co/tencent/Hy4-preview-FP8?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- Model documentation: https://github.com/Tencent-Hunyuan/Hy4-preview?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- Deployment and cost analysis: https://chinaapi.ai/insights/hy4-preview-local-deployment-value/

## Deployment evidence

- **Official launch / minimum — evidence B:** SGLang's official Hy4 matrix lists the smallest published GPU-count path as the MXFP8 checkpoint on 4x B300 or 4x GB300 at TP4 and 262K configured context. It also lists 8x B200 at TP8. These are framework recipes, not a claim that four accelerators are an absolute minimum for every engine or the full 1M context. Source: https://docs.sglang.io/cookbook/autoregressive/Tencent/Hy4-Preview?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **200-person team — evidence E:** For the defined 200-seat profile, start load testing with one official-shape FP8 worker such as 8x B200 or 4x B300. If the service is operationally important, provision a second independent worker for maintenance and failure recovery. Tencent and the serving frameworks do not publish a Hy4 throughput result for this team profile, so this is ChinaAPI capacity-planning guidance, not a vendor recommendation.
- **Commercial API — evidence B:** Tencent publishes vLLM and SGLang OpenAI-compatible recipes with native MTP speculative decoding, sparse-attention support and Hy4 tool/reasoning parsers. A public service should start with at least two independently deployable workers, admission control and separate short- and long-context SLO tests; redundancy is ChinaAPI operational guidance, while the serving commands are official framework recipes. Source: https://docs.sglang.io/cookbook/autoregressive/Tencent/Hy4-Preview?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment

## Scenario evidence

- **Long-horizon software engineering — Vendor-stated:** Tencent positions Hy4 preview for understanding, planning, debugging and verifying long-running development tasks, including front-end interaction quality. Source: https://github.com/Tencent-Hunyuan/Hy4-preview?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Office analysis and artifact creation — Vendor-stated:** The official release describes multi-file document, spreadsheet and presentation work involving data analysis, equations and financial models. Source: https://github.com/Tencent-Hunyuan/Hy4-preview?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Game development and scientific research — Public evidence:** Tencent reports evaluations and internal expert testing across playable game prototypes, AI research, molecular dynamics, condensed-matter physics and pure mathematics; these are vendor results, not ChinaAPI reproductions. Source: https://github.com/Tencent-Hunyuan/Hy4-preview?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment

## Commercial-use check

- License: Apache-2.0 — https://github.com/Tencent-Hunyuan/Hy4-preview/blob/main/LICENSE?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- MaaS / hosted service: No model-specific MaaS restriction identified in the published Apache-2.0 license.
- Attribution: Provide the license and required notices, preserve applicable attribution notices, and mark modified files; trademark rights are not granted.

## Avoid or validate first

- Single-workstation deployment
- Treating the vendor's internal blind evaluation as an independent benchmark
- Promising 1M-context production concurrency before KV-cache and load tests

## Known limitations and open questions

- No ChinaAPI hardware reproduction or throughput benchmark
- The official framework matrix sizes common deployments at 131K or 262K context rather than demonstrating the advertised 1M maximum under production concurrency
- No official Hy4 throughput result is published for the 200-person profile used in this dossier
- The preview may reason longer than necessary and over-verify its own work, according to the vendor
- A 1M context window can make KV-cache demand the production bottleneck even when weights fit

## Evidence boundary

Loading weights, completing a first forward pass and meeting a production latency/SLA are separate thresholds. The 200-person and commercial tiers require workload-specific measurement. License summaries are product research, not legal advice.

Generated by GENERATED BY scripts/gen_open_model_deployment.py.
