Released 2026-02-16 · general agent
Deployment evidence
Alibaba Qwen · 9B total · 9B active · 256K tokens
Best-fit scenarios
The local and edge-friendly baseline for private multimodal assistants, extraction and moderate-volume APIs.
Not recommended
- Assuming 1M extended context fits a 24GB card
- Highest-complexity long-horizon agents without task-specific evaluation
Local multimodal assistant
ChinaAPI inferenceThe 9B dense shape is the most accessible formal model in this ledger.
Visual extraction and classification
Vendor-statedThe official model is a unified vision-language foundation model.
Capability sourceCost-sensitive private API
ChinaAPI inferenceA better first benchmark target than frontier-scale MoE models when concurrency and budget dominate.
Official launch / minimum
Evidence EMemory estimate not official
A 24GB-class GPU is a reasonable BF16 short-context planning target, but the vendor does not label this an official minimum.
200-person team
Evidence EBest low cost candidate
Benchmark one or two 24–48GB workers before considering larger models; exact replicas depend on output length.
Commercial API
Evidence CMainstream framework support
Official Transformers, vLLM and SGLang examples are published; add redundant replicas for availability. Primary recipe
Commercial-use check
Apache-2.0
MaaS / hosted service: No model-specific MaaS restriction identified in Apache-2.0.
Attribution: Provide the license and notices, preserve attribution notices, and mark modified files.
Known limitations and open questions
- 24GB guidance is an estimate, not ChinaAPI reproduction
- Native 262K and extended 1M contexts materially increase KV-cache demand
