Released 2026-07-12 · general agent
Deployment evidence
Z.ai · 753B total · not independently verified in V1 active · 1M tokens
Best-fit scenarios
Long-horizon software engineering, tool use and million-token analysis on enterprise infrastructure.
Not recommended
- Latency-sensitive single-GPU serving
- Capacity planning based only on employee count
Long-horizon software engineering
Vendor-statedOfficial release materials emphasize coding-agent and tool-use workloads.
Capability sourceMillion-token repository and document analysis
Vendor-statedThe official checkpoint supports a 1M-token context window.
Capability sourcePrivate enterprise agent service
ChinaAPI inferenceA credible fit when CPU/GPU offload or TP8 infrastructure is already available.
Official launch / minimum
Evidence BHeterogeneous reference
KTransformers documents an 8-GPU CPU/GPU-offload launch shape with 96 CPU inference threads; this is a tutorial target, not a claimed absolute minimum. Primary recipe
200-person team
Evidence EBenchmark required
Start from a replicated TP8 service only after measuring the target context mix; no official 200-seat capacity claim exists.
Commercial API
Evidence EBenchmark required
Requires redundant replicas or disaggregated prefill/decode; GPU count cannot be inferred from seats alone.
Commercial-use check
MIT
MaaS / hosted service: No model-specific MaaS restriction identified in the MIT license.
Attribution: Retain the copyright and permission notice.
Known limitations and open questions
- The 8-GPU tutorial is not an official minimum claim
- No public V1 concurrency result
