Open-weight deployment dossier · verified 2026-09-12

Qwen3.6-35B-A3B: self-hosting hardware, scenarios and commercial license

A comparatively deployable private multimodal agent baseline with strong ecosystem coverage.

35B
total parameters
256K tokens
context window
Apache-2.0
published license

Released 2026-04-16 · general agent

Deployment evidence

Alibaba Qwen · 35B total · 3B active · 256K tokens

PrecisionBF16; official and community quantized formats availableServing pathsTransformers, vLLM, SGLang, llama.cpp, MLX

Best-fit scenarios

A comparatively deployable private multimodal agent baseline with strong ecosystem coverage.

Not recommended
  • Assuming TP4 equals four production replicas
  • Unvalidated parser upgrades in a critical tool loop

Private coding and office assistant

ChinaAPI inference

The 35B/3B-active shape and broad serving support make it a practical baseline for controlled workloads.

Visual document and UI understanding

Vendor-stated

Qwen3.6 is released as a native multimodal model family.

Capability source

Tool-calling agent service

Vendor-stated

Official serving examples include tool parsing and long-context operation.

Capability source

Official launch / minimum

Evidence C

Official launch shape

Official serving examples use TP4 and 262K context. Smaller quantized short-context shapes exist, but V1 does not call them the official minimum.

200-person team

Evidence E

Best v1 benchmark candidate

Use two measured serving replicas as the initial HA design candidate; exact GPUs remain pending load tests.

Commercial API

Evidence E

Benchmark required

Scale through replicated workers after measuring prefill-heavy and decode-heavy traffic separately.

Commercial-use check

Apache-2.0

MaaS / hosted service: No model-specific MaaS restriction identified in Apache-2.0.

Attribution: Provide the license and required notices, preserve attribution notices, and mark modified files; trademark rights are not granted.

Read the primary license text

Known limitations and open questions
  • Official TP4 example is a launch shape, not a capacity guarantee
  • Tool-call parser issues have been reported for related Qwen3.5 configurations

Evidence boundaries

A launch shape is not a production SLA.

The minimum tier records the smallest official or inference-framework configuration we found. The 200-person and commercial API tiers still require measurements against real prompt length, output length, concurrency, latency and redundancy targets.

Read the complete index methodology

  1. AChinaAPI reproduced
  2. Binference-framework official validated recipe
  3. Cmodel-vendor documented configuration
  4. Dthird-party reproduction
  5. Ecapacity estimate only

Compare before deploying

Review every verified model or compare hosted access.

The directory keeps model selection separate from the evidence and capacity details on this page.