Open-weight deployment dossier · verified 2026-09-12

GLM-5.3-Flash: self-hosting hardware, scenarios and commercial license

Cost-sensitive multimodal coding, document, research and tool-using agents that need a one-million-token model specification and can support either heavy host-memory offload or a multi-GPU worker.

320B
total parameters
1M tokens
context window
MIT
published license

Released 2026-08-26 · general agent

Deployment evidence

Z.ai · 320B total · 18B per token active · 1M tokens

PrecisionTwo official weight repositories: the default native FP8 checkpoint is about 306 GiB before runtime and cache overhead; the separate BF16 variant requires roughly twice the weight memory. Third-party NVFP4 and GGUF conversions are not counted as official weight variants.Serving pathsSGLang, vLLM, KTransformers, Transformers, TokenSpeed, Unsloth

Best-fit scenarios

Cost-sensitive multimodal coding, document, research and tool-using agents that need a one-million-token model specification and can support either heavy host-memory offload or a multi-GPU worker.

Not recommended
  • Treating the 18B active-parameter count as the memory footprint; all 320B parameters still have to be stored or offloaded
  • Using the single-GPU CPU-offload route as a production baseline for a 200-person team
  • Promising 1M-context concurrency before measuring both the KDA state pool and KV cache
  • Leaving maximum reasoning effort enabled for every latency-sensitive request without evaluating token cost

Coding and visual software-engineering agents

Vendor-stated

Z.ai positions GLM-5.3-Flash for coding and agent workflows that combine repository context, screenshots, visual debugging and tool calls. The quality claims are vendor results, not ChinaAPI reproductions.

Capability source

Long documents, presentations and data analysis

Vendor-stated

The official release highlights long-document synthesis, presentation generation and data-analysis work with native image and video understanding.

Capability source

Private multimodal research and meeting workflows

ChinaAPI inference

The combination of image/video input, a 1M model context and local OpenAI-compatible serving makes it a plausible fit where raw media or long records must stay inside the operator's boundary; this is a fit assessment, not a quality benchmark.

Official launch / minimum

Evidence B

Official heterogeneous minimum

KTransformers documents a single-GPU path for the official FP8 weights on NVIDIA SM89 or SM120 hardware, with at least 350 GB of available system memory, AVX-512 CPU inference and 64 CPU inference threads. Its example uses 501,025-token context. This establishes runnable heterogeneous inference, not production latency or concurrency. Primary recipe

200-person team

Evidence B

Validated worker benchmark required

For the defined 200-seat profile, start load testing with one current framework-documented worker: vLLM publishes TP4 on one 4x GB200 tray, while SGLang validates a 4x GB300 TP4/EP4 multimodal topology. Add a second independent worker only when maintenance or failure recovery is required. No official recipe maps employee count to capacity. Primary recipe

Commercial API

Evidence B

Official serving recipe

Use at least two independently deployable workers plus admission control and separate short-, image/video- and long-context SLOs. SGLang verifies encoder disaggregation on one shared 4x GB300 node for image requests and videos up to 238,080 visual tokens; its prefill/decode path remains preview-only. Final capacity still requires measured traffic and failure tests. Primary recipe

Commercial-use check

MIT

MaaS / hosted service: No model-specific MaaS restriction identified in the MIT license published with the official weights.

Attribution: Retain the copyright and permission notice in copies or substantial portions of the weights. The weights are provided without warranty.

Read the primary license text

Known limitations and open questions
  • No ChinaAPI hardware reproduction, quality benchmark or like-for-like comparison between the downloadable checkpoint and hosted route
  • The KTransformers single-GPU recipe publishes no throughput or 200-user concurrency result and requires at least 350 GB of available system memory
  • Its documented heterogeneous example uses 501,025 tokens rather than proving the full 1,048,576-token window under load
  • vLLM currently requires a Docker build before GLM-5.3-Flash integration reaches its public repository; pin the exact serving image and framework revision
  • SGLang's prefill/decode path is preview-only and has been mechanically validated with dummy weights rather than load- or accuracy-tested
  • Reasoning defaults to maximum effort unless low or high is passed explicitly, which can materially change latency and billed output
  • Vendor capability and efficiency claims have not been reproduced by ChinaAPI

Evidence boundaries

A launch shape is not a production SLA.

The minimum tier records the smallest official or inference-framework configuration we found. The 200-person and commercial API tiers still require measurements against real prompt length, output length, concurrency, latency and redundancy targets.

Read the complete index methodology

  1. AChinaAPI reproduced
  2. Binference-framework official validated recipe
  3. Cmodel-vendor documented configuration
  4. Dthird-party reproduction
  5. Ecapacity estimate only

Compare before deploying

Review every verified model or compare hosted access.

The directory keeps model selection separate from the evidence and capacity details on this page.