Open-weight deployment dossier · verified 2026-09-12

Qwen3.8-2.4T-A95B: self-hosting hardware, scenarios and commercial license

Frontier-scale text reasoning, coding and long-horizon agent workloads where a multi-node cluster and custom commercial-license review are acceptable.

2.4T
total parameters
256K tokens
context window
Qwen3.8-Max License
published license

Released 2026-08-14 · general agent

Deployment evidence

Alibaba Qwen · 2.4T total · 95B active · 256K tokens

PrecisionBF16 plus an official FP8 checkpoint; the two repositories are precision variants of one model specificationServing pathsvLLM, SGLang, TokenSpeed

Best-fit scenarios

Frontier-scale text reasoning, coding and long-horizon agent workloads where a multi-node cluster and custom commercial-license review are acceptable.

Not recommended
  • Single-workstation deployment
  • Treating the open text-only checkpoint as identical to the hosted multimodal qwen3.8-max service
  • Commercial MaaS launch before license-threshold review

Frontier coding agents

Vendor-stated

The official open model card positions the 2.4T/95B-active checkpoint for reasoning, coding and agent tasks.

Capability source

Long-horizon tool orchestration

Vendor-stated

The open release supports flexible reasoning and a native 262K context window, extendable to about 1M tokens.

Capability source

Premium private text endpoint

ChinaAPI inference

Consider only when expected capability value justifies distributed data-center infrastructure and the custom license has been cleared.

Official launch / minimum

Evidence C

Not published

The vendor documents vLLM, SGLang and TokenSpeed serving paths but does not publish an absolute minimum hardware configuration. BF16 weights alone are roughly 4.8TB and FP8 weights roughly 2.4TB before runtime and KV cache.

200-person team

Evidence E

Requires cluster sizing

A 2.4T model requires a measured multi-node design; employee count alone cannot determine accelerator count, and no official 200-seat capacity claim exists.

Commercial API

Evidence E

Benchmark and license required

Plan redundant distributed workers, disaggregated prefill/decode where supported, admission control and long-context tests. Clear the USD 50M MaaS/AI Work Assistant threshold before launch; no ChinaAPI production topology has been validated.

Commercial-use check

Qwen3.8-Max License

MaaS / hosted service: A separate license is required before commercial use when a Model-as-a-Service or AI Work Assistant business and its affiliates exceed USD 50M aggregate revenue in any consecutive 12-month period.

Attribution: Products or services above 100M monthly active users or USD 20M monthly revenue must prominently display the model name in the user interface.

Exceptions: The license excludes mere relaying to third-party hosted models from its Model-as-a-Service definition. Review the complete license text for scope and other obligations.

Read the primary license text

Known limitations and open questions
  • No ChinaAPI hardware reproduction
  • The open checkpoint is text-only while the hosted qwen3.8-max service adds vision and built-in tools
  • The official FP8 repository is a precision variant, not a fourth model size

Evidence boundaries

A launch shape is not a production SLA.

The minimum tier records the smallest official or inference-framework configuration we found. The 200-person and commercial API tiers still require measurements against real prompt length, output length, concurrency, latency and redundancy targets.

Read the complete index methodology

  1. AChinaAPI reproduced
  2. Binference-framework official validated recipe
  3. Cmodel-vendor documented configuration
  4. Dthird-party reproduction
  5. Ecapacity estimate only

Compare before deploying

Review every verified model or compare hosted access.

The directory keeps model selection separate from the evidence and capacity details on this page.