Open-weight deployment dossier · verified 2026-09-12

Qwen3.8-27B: self-hosting hardware, scenarios and commercial license

A comparatively deployable native vision-language model for private coding, computer-use, document and general agent workloads.

27B
total parameters
256K tokens
context window
Apache-2.0
published license

Released 2026-08-14 · general agent

Deployment evidence

Alibaba Qwen · 27B total · 27B dense active · 256K tokens

PrecisionBF16 plus an official FP8 checkpoint; the two repositories are precision variants of one model specificationServing pathsTransformers, vLLM, SGLang

Best-fit scenarios

A comparatively deployable native vision-language model for private coding, computer-use, document and general agent workloads.

Not recommended
  • Assuming a 27B weight-memory estimate includes KV cache and vision processing
  • Promising full 262K context or production concurrency from a single workstation without load tests

Coding and repository agents

Public evidence

The official model card reports Terminal Bench 2.1, SWE-bench Pro, DeepSWE, NL2Repo and QwenSWEBench results; these are vendor-reported, not ChinaAPI reproductions.

Capability source

Computer-use and mobile agents

Public evidence

The official card publishes OSWorld, WebArena and AndroidWorld evaluations for agent interaction with visual interfaces.

Capability source

Private multimodal assistant

ChinaAPI inference

The dense 27B shape and native image/video inputs make it the more practical of the two Qwen3.8 open releases for controlled enterprise deployment.

Official launch / minimum

Evidence E

Memory estimate not official

BF16 weights alone are about 54GB before runtime and KV cache. The vendor publishes BF16 and FP8 checkpoints but no absolute minimum GPU configuration; community quantizations are not treated as an official minimum.

200-person team

Evidence E

Benchmark required

Start capacity testing with two redundant 48–80GB-class workers or an equivalent tensor-parallel design, then size from the real vision, context and output mix; this is a planning estimate, not an official recommendation.

Commercial API

Evidence C

Official serving paths

Official vLLM and SGLang launch paths are published. A public service still needs redundant replicas, admission control and separate short-context, vision, 262K-context and tool-call benchmarks. Primary recipe

Commercial-use check

Apache-2.0

MaaS / hosted service: No model-specific MaaS restriction identified in Apache-2.0.

Attribution: Provide the license and required notices, preserve attribution notices, and mark modified files; trademark rights are not granted.

Read the primary license text

Known limitations and open questions
  • No ChinaAPI hardware reproduction
  • The official FP8 repository is a precision variant, not a third Qwen3.8 model size
  • Native 262K and extended 1M context materially increase memory demand

Evidence boundaries

A launch shape is not a production SLA.

The minimum tier records the smallest official or inference-framework configuration we found. The 200-person and commercial API tiers still require measurements against real prompt length, output length, concurrency, latency and redundancy targets.

Read the complete index methodology

  1. AChinaAPI reproduced
  2. Binference-framework official validated recipe
  3. Cmodel-vendor documented configuration
  4. Dthird-party reproduction
  5. Ecapacity estimate only

Compare before deploying

Review every verified model or compare hosted access.

The directory keeps model selection separate from the evidence and capacity details on this page.