Released 2026-08-26 · general agent
Deployment evidence
Alibaba Qwen · 125B main + 51B N-gram embedding + 4B MTP total · 6B per token active · 256K tokens
Best-fit scenarios
Cost-sensitive coding, office and multimodal agent workloads that need strong public benchmark results but can accept preview-stage serving software or a managed production derivative.
Not recommended
- Calling the 6B active-parameter count a 6B memory footprint
- Treating a patched community NVFP4 single-workstation run as the official deployment baseline
- Launching MaaS or an independent coding or office assistant before obtaining the separate commercial license
Coding and software-engineering agents
Public evidenceQwen reports 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro and 81.0 on SWE-bench Multilingual. These are vendor-published evaluations, not ChinaAPI reproductions.
Capability sourceLong-horizon office and tool work
Public evidenceThe official card reports 73.9 on CoWorkBench, 55.7 on JobBench and 73.5 on Toolathlon Verified, with the vendor positioning the model for coding and office tasks.
Capability sourceSingle-workstation research pilot
Third-party reproducedA third-party NVFP4 recipe reported a median 43.81 output tokens/s on one 128GB DGX Spark at 8,192-token context and single concurrency. It used a 101 GiB non-Qwen-official four-bit checkpoint and custom serving work, so it is evidence of feasibility, not an official minimum or production-capacity result.
Capability sourceOfficial launch / minimum
Evidence BOfficial validated minimum
vLLM validates the official 172.78 GiB FP8 checkpoint on 2x GB300 as the minimum supported FP8 topology. The 51B N-gram table also needs at least 51GB of host memory plus runtime headroom. This is the official-framework minimum, not a promise about useful concurrency at 262K context. Primary recipe
200-person team
Evidence BValidated worker plus redundancy
For the defined 200-seat profile, start load testing with the validated 4x H100 80GB FP8 worker and CPU-offloaded N-gram table. vLLM reports about 1,430 output tokens/s at concurrency 64 on a random 1,024-input/256-output workload, but a team service still needs a second failure-domain worker if availability matters and separate tests for real multimodal and long-context traffic. Primary recipe
Commercial API
Evidence BLicense and capacity review required
Use at least two independently deployable workers; vLLM recommends 4x GB300 FP8 per full-tray worker and also validates 8x H200 TEP8, 4x H100 with N-gram CPU offload, and 4x MI355X. Final topology depends on context, image/video mix and SLOs, and commercial MaaS requires a separate Qwen license before launch. Primary recipe
Commercial-use check
Qwen Community License 1.0
MaaS / hosted service: A separate Qwen license is required before commercial use by a licensee or affiliate that conducts a Model-as-a-Service or AI Work Assistant business. This condition has no revenue threshold in the published license.
Attribution: Include the copyright and permission notice. A commercial product or service above 100M monthly active users or USD 20M monthly revenue must prominently display the model name in its user interface.
Exceptions: Internal use is exempt from the separate-license condition only when the software, outputs and model capabilities are not made available to a third party. Mere relay to a third-party hosted model is excluded from MaaS; single-purpose tools, non-coding/non-office domain assistants, and assistants that are only a feature of another primary product are excluded from the defined AI Work Assistant category.
Known limitations and open questions
- No ChinaAPI reproduction of the downloadable checkpoint and no like-for-like quality comparison with the hosted qwen3.8-flash production model
- The official vLLM recipe requires a dedicated day-zero image; SGLang documentation said support was not yet in a tagged release when reviewed
- The validated 4x H100 throughput uses random 1,024-input/256-output requests and does not establish coding-agent, vision, video or 262K-context latency
- One-million-token context requires static YaRN and may affect shorter-context quality
