Open-weight deployment dossier · verified 2026-09-12

Hy3: self-hosting hardware, scenarios and commercial license

Large private coding, productivity and tool-using agents that can support an official eight-accelerator tensor-parallel worker.

295B backbone plus 3.8B MTP layer
total parameters
256K tokens
context window
Apache-2.0
published license

Released 2026-07-06 · general agent

Deployment evidence

Tencent Hunyuan · 295B backbone plus 3.8B MTP layer total · 21B backbone per token active · 256K tokens

PrecisionBF16Serving pathsTransformers, vLLM, SGLang

Best-fit scenarios

Large private coding, productivity and tool-using agents that can support an official eight-accelerator tensor-parallel worker.

Not recommended
  • Planning from the 21B active count instead of loading the full checkpoint
  • Treating an eight-GPU launch command as a production capacity result

Coding and productivity agents

Vendor-stated

Tencent positions Hy3 for coding, productivity, reasoning and agent tasks. Published capability results are vendor evaluations, not ChinaAPI reproductions.

Capability source

Long-context tool workflows

Vendor-stated

The official configuration supports 256K context and documents tool and reasoning parsers for vLLM and SGLang.

Capability source

Private OpenAI-compatible endpoint

ChinaAPI inference

The official serving commands make an operator-controlled endpoint plausible, but capacity still depends on prompt length and concurrency.

Official launch / minimum

Evidence C

Official vendor floor

The vendor recommends eight H20-3e GPUs or accelerators with more memory for the documented vLLM and SGLang TP8 path. This proves an official launch shape, not an SLA. Primary recipe

200-person team

Evidence E

Benchmark required

Benchmark one TP8 worker against the 200-seat profile, then add an independent worker if maintenance and failure recovery must preserve service.

Commercial API

Evidence C

Official framework paths

Official vLLM and SGLang OpenAI-compatible paths exist; a public service still needs redundant TP8 workers, admission control and pinned parsers. Primary recipe

Commercial-use check

Apache-2.0

MaaS / hosted service: No model-specific MaaS restriction identified in Apache-2.0.

Attribution: Provide the license and required notices, preserve attribution notices, and mark modified files; trademark rights are not granted.

Read the primary license text

Known limitations and open questions
  • No ChinaAPI hardware or quality reproduction
  • The official eight-GPU shape does not state concurrency or latency
  • Full-context KV-cache demand must be measured separately

Evidence boundaries

A launch shape is not a production SLA.

The minimum tier records the smallest official or inference-framework configuration we found. The 200-person and commercial API tiers still require measurements against real prompt length, output length, concurrency, latency and redundancy targets.

Read the complete index methodology

  1. AChinaAPI reproduced
  2. Binference-framework official validated recipe
  3. Cmodel-vendor documented configuration
  4. Dthird-party reproduction
  5. Ecapacity estimate only

Compare before deploying

Review every verified model or compare hosted access.

The directory keeps model selection separate from the evidence and capacity details on this page.