Released 2026-08-14 · general agent
Deployment evidence
Alibaba Qwen · 27B total · 27B dense active · 256K tokens
Best-fit scenarios
A comparatively deployable native vision-language model for private coding, computer-use, document and general agent workloads.
Not recommended
- Assuming a 27B weight-memory estimate includes KV cache and vision processing
- Promising full 262K context or production concurrency from a single workstation without load tests
Coding and repository agents
Public evidenceThe official model card reports Terminal Bench 2.1, SWE-bench Pro, DeepSWE, NL2Repo and QwenSWEBench results; these are vendor-reported, not ChinaAPI reproductions.
Capability sourceComputer-use and mobile agents
Public evidenceThe official card publishes OSWorld, WebArena and AndroidWorld evaluations for agent interaction with visual interfaces.
Capability sourcePrivate multimodal assistant
ChinaAPI inferenceThe dense 27B shape and native image/video inputs make it the more practical of the two Qwen3.8 open releases for controlled enterprise deployment.
Official launch / minimum
Evidence EMemory estimate not official
BF16 weights alone are about 54GB before runtime and KV cache. The vendor publishes BF16 and FP8 checkpoints but no absolute minimum GPU configuration; community quantizations are not treated as an official minimum.
200-person team
Evidence EBenchmark required
Start capacity testing with two redundant 48–80GB-class workers or an equivalent tensor-parallel design, then size from the real vision, context and output mix; this is a planning estimate, not an official recommendation.
Commercial API
Evidence COfficial serving paths
Official vLLM and SGLang launch paths are published. A public service still needs redundant replicas, admission control and separate short-context, vision, 262K-context and tool-call benchmarks. Primary recipe
Commercial-use check
Apache-2.0
MaaS / hosted service: No model-specific MaaS restriction identified in Apache-2.0.
Attribution: Provide the license and required notices, preserve attribution notices, and mark modified files; trademark rights are not granted.
Known limitations and open questions
- No ChinaAPI hardware reproduction
- The official FP8 repository is a precision variant, not a third Qwen3.8 model size
- Native 262K and extended 1M context materially increase memory demand
