Open-weight deployment dossier · verified 2026-09-12

Qwen3.5-9B: self-hosting hardware, scenarios and commercial license

The local and edge-friendly baseline for private multimodal assistants, extraction and moderate-volume APIs.

9B
total parameters
256K tokens
context window
Apache-2.0
published license

Released 2026-02-16 · general agent

Deployment evidence

Alibaba Qwen · 9B total · 9B active · 256K tokens

PrecisionBF16; multiple quantized formats availableServing pathsTransformers, vLLM, SGLang, KTransformers, llama.cpp

Best-fit scenarios

The local and edge-friendly baseline for private multimodal assistants, extraction and moderate-volume APIs.

Not recommended
  • Assuming 1M extended context fits a 24GB card
  • Highest-complexity long-horizon agents without task-specific evaluation

Local multimodal assistant

ChinaAPI inference

The 9B dense shape is the most accessible formal model in this ledger.

Visual extraction and classification

Vendor-stated

The official model is a unified vision-language foundation model.

Capability source

Cost-sensitive private API

ChinaAPI inference

A better first benchmark target than frontier-scale MoE models when concurrency and budget dominate.

Official launch / minimum

Evidence E

Memory estimate not official

A 24GB-class GPU is a reasonable BF16 short-context planning target, but the vendor does not label this an official minimum.

200-person team

Evidence E

Best low cost candidate

Benchmark one or two 24–48GB workers before considering larger models; exact replicas depend on output length.

Commercial API

Evidence C

Mainstream framework support

Official Transformers, vLLM and SGLang examples are published; add redundant replicas for availability. Primary recipe

Commercial-use check

Apache-2.0

MaaS / hosted service: No model-specific MaaS restriction identified in Apache-2.0.

Attribution: Provide the license and notices, preserve attribution notices, and mark modified files.

Read the primary license text

Known limitations and open questions
  • 24GB guidance is an estimate, not ChinaAPI reproduction
  • Native 262K and extended 1M contexts materially increase KV-cache demand

Evidence boundaries

A launch shape is not a production SLA.

The minimum tier records the smallest official or inference-framework configuration we found. The 200-person and commercial API tiers still require measurements against real prompt length, output length, concurrency, latency and redundancy targets.

Read the complete index methodology

  1. AChinaAPI reproduced
  2. Binference-framework official validated recipe
  3. Cmodel-vendor documented configuration
  4. Dthird-party reproduction
  5. Ecapacity estimate only

Compare before deploying

Review every verified model or compare hosted access.

The directory keeps model selection separate from the evidence and capacity details on this page.