Open-weight deployment dossier · verified 2026-09-12

GLM-5.2: self-hosting hardware, scenarios and commercial license

Long-horizon software engineering, tool use and million-token analysis on enterprise infrastructure.

753B
total parameters
1M tokens
context window
MIT
published license

Released 2026-07-12 · general agent

Deployment evidence

Z.ai · 753B total · not independently verified in V1 active · 1M tokens

PrecisionBF16 and official FP8 checkpointServing pathsvLLM, SGLang, xLLM, KTransformers

Best-fit scenarios

Long-horizon software engineering, tool use and million-token analysis on enterprise infrastructure.

Not recommended
  • Latency-sensitive single-GPU serving
  • Capacity planning based only on employee count

Long-horizon software engineering

Vendor-stated

Official release materials emphasize coding-agent and tool-use workloads.

Capability source

Million-token repository and document analysis

Vendor-stated

The official checkpoint supports a 1M-token context window.

Capability source

Private enterprise agent service

ChinaAPI inference

A credible fit when CPU/GPU offload or TP8 infrastructure is already available.

Official launch / minimum

Evidence B

Heterogeneous reference

KTransformers documents an 8-GPU CPU/GPU-offload launch shape with 96 CPU inference threads; this is a tutorial target, not a claimed absolute minimum. Primary recipe

200-person team

Evidence E

Benchmark required

Start from a replicated TP8 service only after measuring the target context mix; no official 200-seat capacity claim exists.

Commercial API

Evidence E

Benchmark required

Requires redundant replicas or disaggregated prefill/decode; GPU count cannot be inferred from seats alone.

Commercial-use check

MIT

MaaS / hosted service: No model-specific MaaS restriction identified in the MIT license.

Attribution: Retain the copyright and permission notice.

Read the primary license text

Known limitations and open questions
  • The 8-GPU tutorial is not an official minimum claim
  • No public V1 concurrency result

Evidence boundaries

A launch shape is not a production SLA.

The minimum tier records the smallest official or inference-framework configuration we found. The 200-person and commercial API tiers still require measurements against real prompt length, output length, concurrency, latency and redundancy targets.

Read the complete index methodology

  1. AChinaAPI reproduced
  2. Binference-framework official validated recipe
  3. Cmodel-vendor documented configuration
  4. Dthird-party reproduction
  5. Ecapacity estimate only

Compare before deploying

Review every verified model or compare hosted access.

The directory keeps model selection separate from the evidence and capacity details on this page.