# Qwen3.8-2.4T-A95B: Self-Hosting Hardware, Scenarios and Commercial License

> Frontier-scale text reasoning, coding and long-horizon agent workloads where a multi-node cluster and custom commercial-license review are acceptable.

- Verified: 2026-09-12
- Released: 2026-08-14
- Canonical: https://chinaapi.ai/open-model-deployment/qwen3.8-2.4t-a95b/
- Back to directory: https://chinaapi.ai/open-model-deployment/
- Machine-readable dataset: https://chinaapi.ai/data/open-model-deployment.json

## Model facts

- Vendor: Alibaba Qwen
- Parameters: 2.4T total; 95B active
- Native precision: BF16 plus an official FP8 checkpoint; the two repositories are precision variants of one model specification
- Context: 256K tokens
- Category: General and agentic foundation models
- Modalities: text → text
- Serving paths: vLLM, SGLang, TokenSpeed
- Official weights: BF16 / original: https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment; Official FP8: https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B-FP8?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- Model documentation: https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment

## Deployment evidence

- **Official launch / minimum — evidence C:** The vendor documents vLLM, SGLang and TokenSpeed serving paths but does not publish an absolute minimum hardware configuration. BF16 weights alone are roughly 4.8TB and FP8 weights roughly 2.4TB before runtime and KV cache.
- **200-person team — evidence E:** A 2.4T model requires a measured multi-node design; employee count alone cannot determine accelerator count, and no official 200-seat capacity claim exists.
- **Commercial API — evidence E:** Plan redundant distributed workers, disaggregated prefill/decode where supported, admission control and long-context tests. Clear the USD 50M MaaS/AI Work Assistant threshold before launch; no ChinaAPI production topology has been validated.

## Scenario evidence

- **Frontier coding agents — Vendor-stated:** The official open model card positions the 2.4T/95B-active checkpoint for reasoning, coding and agent tasks. Source: https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Long-horizon tool orchestration — Vendor-stated:** The open release supports flexible reasoning and a native 262K context window, extendable to about 1M tokens. Source: https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Premium private text endpoint — ChinaAPI inference:** Consider only when expected capability value justifies distributed data-center infrastructure and the custom license has been cleared.

## Commercial-use check

- License: Qwen3.8-Max License — https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B/blob/main/LICENSE?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- MaaS / hosted service: A separate license is required before commercial use when a Model-as-a-Service or AI Work Assistant business and its affiliates exceed USD 50M aggregate revenue in any consecutive 12-month period.
- Attribution: Products or services above 100M monthly active users or USD 20M monthly revenue must prominently display the model name in the user interface.
- Exceptions: The license excludes mere relaying to third-party hosted models from its Model-as-a-Service definition. Review the complete license text for scope and other obligations.

## Avoid or validate first

- Single-workstation deployment
- Treating the open text-only checkpoint as identical to the hosted multimodal qwen3.8-max service
- Commercial MaaS launch before license-threshold review

## Known limitations and open questions

- No ChinaAPI hardware reproduction
- The open checkpoint is text-only while the hosted qwen3.8-max service adds vision and built-in tools
- The official FP8 repository is a precision variant, not a fourth model size

## Evidence boundary

Loading weights, completing a first forward pass and meeting a production latency/SLA are separate thresholds. The 200-person and commercial tiers require workload-specific measurement. License summaries are product research, not legal advice.

Generated by GENERATED BY scripts/gen_open_model_deployment.py.
