Open-weight deployment dossier · verified 2026-09-12

Qwen3.8-Flash-Next: self-hosting hardware, scenarios and commercial license

Cost-sensitive coding, office and multimodal agent workloads that need strong public benchmark results but can accept preview-stage serving software or a managed production derivative.

125B main + 51B N-gram embedding + 4B MTP
total parameters
256K tokens
context window
Qwen Community License 1.0
published license

Released 2026-08-26 · general agent

Deployment evidence

Alibaba Qwen · 125B main + 51B N-gram embedding + 4B MTP total · 6B per token active · 256K tokens

PrecisionOfficial BF16 checkpoint (335.28 GiB) and official FP8 checkpoint (172.78 GiB); community and NVIDIA NVFP4 conversions are separate artifactsServing pathsTransformers, llama.cpp, MLX, vLLM, SGLang, TokenSpeed

Best-fit scenarios

Cost-sensitive coding, office and multimodal agent workloads that need strong public benchmark results but can accept preview-stage serving software or a managed production derivative.

Not recommended
  • Calling the 6B active-parameter count a 6B memory footprint
  • Treating a patched community NVFP4 single-workstation run as the official deployment baseline
  • Launching MaaS or an independent coding or office assistant before obtaining the separate commercial license

Coding and software-engineering agents

Public evidence

Qwen reports 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro and 81.0 on SWE-bench Multilingual. These are vendor-published evaluations, not ChinaAPI reproductions.

Capability source

Long-horizon office and tool work

Public evidence

The official card reports 73.9 on CoWorkBench, 55.7 on JobBench and 73.5 on Toolathlon Verified, with the vendor positioning the model for coding and office tasks.

Capability source

Single-workstation research pilot

Third-party reproduced

A third-party NVFP4 recipe reported a median 43.81 output tokens/s on one 128GB DGX Spark at 8,192-token context and single concurrency. It used a 101 GiB non-Qwen-official four-bit checkpoint and custom serving work, so it is evidence of feasibility, not an official minimum or production-capacity result.

Capability source

Official launch / minimum

Evidence B

Official validated minimum

vLLM validates the official 172.78 GiB FP8 checkpoint on 2x GB300 as the minimum supported FP8 topology. The 51B N-gram table also needs at least 51GB of host memory plus runtime headroom. This is the official-framework minimum, not a promise about useful concurrency at 262K context. Primary recipe

200-person team

Evidence B

Validated worker plus redundancy

For the defined 200-seat profile, start load testing with the validated 4x H100 80GB FP8 worker and CPU-offloaded N-gram table. vLLM reports about 1,430 output tokens/s at concurrency 64 on a random 1,024-input/256-output workload, but a team service still needs a second failure-domain worker if availability matters and separate tests for real multimodal and long-context traffic. Primary recipe

Commercial API

Evidence B

License and capacity review required

Use at least two independently deployable workers; vLLM recommends 4x GB300 FP8 per full-tray worker and also validates 8x H200 TEP8, 4x H100 with N-gram CPU offload, and 4x MI355X. Final topology depends on context, image/video mix and SLOs, and commercial MaaS requires a separate Qwen license before launch. Primary recipe

Commercial-use check

Qwen Community License 1.0

MaaS / hosted service: A separate Qwen license is required before commercial use by a licensee or affiliate that conducts a Model-as-a-Service or AI Work Assistant business. This condition has no revenue threshold in the published license.

Attribution: Include the copyright and permission notice. A commercial product or service above 100M monthly active users or USD 20M monthly revenue must prominently display the model name in its user interface.

Exceptions: Internal use is exempt from the separate-license condition only when the software, outputs and model capabilities are not made available to a third party. Mere relay to a third-party hosted model is excluded from MaaS; single-purpose tools, non-coding/non-office domain assistants, and assistants that are only a feature of another primary product are excluded from the defined AI Work Assistant category.

Read the primary license text

Known limitations and open questions
  • No ChinaAPI reproduction of the downloadable checkpoint and no like-for-like quality comparison with the hosted qwen3.8-flash production model
  • The official vLLM recipe requires a dedicated day-zero image; SGLang documentation said support was not yet in a tagged release when reviewed
  • The validated 4x H100 throughput uses random 1,024-input/256-output requests and does not establish coding-agent, vision, video or 262K-context latency
  • One-million-token context requires static YaRN and may affect shorter-context quality

Evidence boundaries

A launch shape is not a production SLA.

The minimum tier records the smallest official or inference-framework configuration we found. The 200-person and commercial API tiers still require measurements against real prompt length, output length, concurrency, latency and redundancy targets.

Read the complete index methodology

  1. AChinaAPI reproduced
  2. Binference-framework official validated recipe
  3. Cmodel-vendor documented configuration
  4. Dthird-party reproduction
  5. Ecapacity estimate only

Compare before deploying

Review every verified model or compare hosted access.

The directory keeps model selection separate from the evidence and capacity details on this page.