Official launch / minimum
One request, short context, no SLA; proves runnable, not production-ready.
Open model deployment index · verified 2026-08-03
A source-linked ledger of official weights, best-fit scenarios, launch configurations, 200-person team assumptions, public API architecture and commercial-use restrictions. We label vendor claims and estimates.

Choose by scenario
“Best fit” is not a universal quality ranking. Each recommendation is labelled as vendor-stated, backed by a named public evaluation, or inferred by ChinaAPI from the model’s modality and deployment shape.
| Model | Best-fit use | Avoid / validate first | Primary claim type |
|---|---|---|---|
| Kimi K3General and agentic foundation models | Frontier-scale long-context multimodal and coding agent workloads where a cluster deployment is acceptable. | Single-workstation deployments; Low-cost high-QPS chat without aggressive batching or distillation | Vendor-stated |
| GLM-5.2General and agentic foundation models | Long-horizon software engineering, tool use and million-token analysis on enterprise infrastructure. | Latency-sensitive single-GPU serving; Capacity planning based only on employee count | Vendor-stated |
| DeepSeek V4 ProGeneral and agentic foundation models | Maximum-quality reasoning, coding and long-context agents when an eight-accelerator worker is viable. | Budget workstation inference; Public API launch without at least one redundant worker | Public evidence |
| MiniMax M3General and agentic foundation models | Multimodal agent, coding and long-context workloads with explicit commercial-license review. | Commercial launch before license notice and attribution review; Stable-production serving without pinned framework versions | Vendor-stated |
| Qwen3.6-35B-A3BGeneral and agentic foundation models | A comparatively deployable private multimodal agent baseline with strong ecosystem coverage. | Assuming TP4 equals four production replicas; Unvalidated parser upgrades in a critical tool loop | ChinaAPI inference |
| DeepSeek V4 FlashGeneral and agentic foundation models | A smaller DeepSeek V4 worker for high-frequency coding, reasoning and agent traffic. | Treating Flash as quality-equivalent to Pro on every task; Launching a public API without measured parser and long-context behavior | Public evidence |
| Step 3.7 FlashGeneral and agentic foundation models | High-frequency multimodal agents, financial-document parsing, verified search loops and concurrent coding. | Using the 128GB local path as a production throughput claim; Assuming vendor benchmark throughput transfers to long-context prefill | Vendor-stated |
| LongCat 2.0General and agentic foundation models | Long-horizon coding, search, repository edits and tool-driven agents on GPU or NPU clusters. | Workstation deployment; Quoting in-house benchmark results as ChinaAPI reproduction | Vendor-stated |
| MiMo V2.5General and agentic foundation models | Native omnimodal understanding, long-context reasoning and agentic workflows across text, image, video and audio. | Small single-GPU deployment; Using stale config or tokenizer files from the initial release | Vendor-stated |
| Qwen3.5-9BGeneral and agentic foundation models | The local and edge-friendly baseline for private multimodal assistants, extraction and moderate-volume APIs. | Assuming 1M extended context fits a 24GB card; Highest-complexity long-horizon agents without task-specific evaluation | ChinaAPI inference |
| Qwen-AgentWorld-35B-A3BDigital-world and environment models | A language world model for simulating MCP, Search, Terminal, SWE, Android, Web and OS agent environments. | General-purpose chat replacement; Treating simulated success as proof of real-environment reliability | Vendor-stated |
| Xiaomi-Robotics-0Physical-world vision-language-action models | Real-time robotic manipulation research and post-training across supported embodiments and simulation suites. | Direct deployment on an unvalidated physical robot; Safety-critical control without independent interlocks and task-specific validation | Vendor-stated |
| NAVAJoint audio-video generation | Synchronized audio-video generation with multi-speaker timbre control, camera prompting and image continuation. | Unconsented face or voice cloning; Low-latency interactive video generation | Vendor-stated |
| MiMo-V2.5-ASRAutomatic speech recognition | Chinese and English transcription across dialects, code-switching, meetings, songs, noise and knowledge-heavy speech. | Capacity promises before real-time-factor and batch testing; Assuming speaker overlap performance replaces diarization requirements | Vendor-stated |
Evidence ledger
“Minimum” means the smallest official or framework-documented launch shape we found—not a production recommendation. A 200-seat estimate is workload-dependent, while a public API also needs redundancy and rollout capacity.
| Model | Commercial terms | Official launch / minimum | Commercial API reference |
|---|---|---|---|
| Kimi K3Moonshot AI · 2.8T | Commercial use: conditionalKimi K3 License | C Vendor documents supported engines but does not publish a smallest runnable GPU configuration. | B NVIDIA Dynamo publishes full-1M profiles using 8x GB300 or 16x GB200 per aggregated worker; disaggregated profiles use more GPUs. |
| GLM-5.2Z.ai · 753B | Commercial use: permittedMIT | B KTransformers documents an 8-GPU CPU/GPU-offload launch shape with 96 CPU inference threads; this is a tutorial target, not a claimed absolute minimum. | E Requires redundant replicas or disaggregated prefill/decode; GPU count cannot be inferred from seats alone. |
| DeepSeek V4 ProDeepSeek · 862B reported by model repository | Commercial use: permittedMIT | B vLLM publishes 8x B300, 8x H200, 8x MI355X and GB200 profiles; none is labelled the absolute minimum. | B 8x H200 is documented with context capped at 800K to preserve KV headroom; production redundancy requires additional workers. |
| MiniMax M3MiniMax · approximately 428B | Commercial use: conditionalMiniMax Community License | D A 4x RTX PRO 6000 NVFP4 recipe is under review upstream; V1 does not promote an unmerged recipe to validated minimum. | B Aggregated and disaggregated vLLM recipes are still landing; pinning a nightly build may be required. |
| Qwen3.6-35B-A3BAlibaba Qwen · 35B | Commercial use: permittedApache-2.0 | C Official serving examples use TP4 and 262K context. Smaller quantized short-context shapes exist, but V1 does not call them the official minimum. | E Scale through replicated workers after measuring prefill-heavy and decode-heavy traffic separately. |
| DeepSeek V4 FlashDeepSeek · 284B | Commercial use: permittedMIT | C Official local instructions and engine integrations are published, but no absolute minimum GPU count is claimed. | B vLLM and SGLang support are available; production still requires redundant workers and pinned encoding logic. |
| Step 3.7 FlashStepFun · 198B | Commercial use: permittedApache-2.0 | C The official GGUF path specifies about 120GB minimum unified memory/VRAM and recommends 128GB. | C Official examples publish TP4 NVFP4 and TP8 FP8/BF16 serving shapes; replicas are still required for HA. |
| LongCat 2.0Meituan · 1.6T | Commercial use: permittedMIT | C GPU and NPU paths are documented, but the vendor does not publish an absolute minimum hardware shape. | C Official GPU and NPU serving paths exist; topology, redundancy and throughput remain operator-specific. |
| MiMo V2.5Xiaomi MiMo · 310B | Commercial use: permittedMIT | C Transformers can load the checkpoint, but the vendor does not state an absolute minimum GPU configuration. | C The official card shows an FP8 SGLang DP2×TP8 configuration at 262K context. |
| Qwen3.5-9BAlibaba Qwen · 9B | Commercial use: permittedApache-2.0 | E A 24GB-class GPU is a reasonable BF16 short-context planning target, but the vendor does not label this an official minimum. | C Official Transformers, vLLM and SGLang examples are published; add redundant replicas for availability. |
| Qwen-AgentWorld-35B-A3BAlibaba Qwen · 35B | Commercial use: permittedApache-2.0 | C Official SGLang and vLLM examples use tensor parallel size 4. | E Expose behind a task-specific simulator contract; do not market it as a normal chat-completions quality substitute. |
| Xiaomi-Robotics-0Xiaomi Robotics · 4.7B | Commercial use: permittedApache-2.0 | C The vendor says BF16 inference is optimized for consumer GPUs but publishes no exact minimum VRAM. | E Prefer an on-robot or near-edge safety architecture; a remote multi-tenant API is not the default production shape. |
| NAVABaidu ERNIE Team · 6.3B backbone | Commercial use: conditionalApache-2.0 | C The model card supports single-GPU inference but does not publish exact minimum VRAM. | C The official Ulysses SP8 path reports roughly one minute for a 720p synchronized clip; HA needs additional workers. |
| MiMo-V2.5-ASRXiaomi MiMo · not published | Commercial use: permittedApache-2.0 | C The official local Gradio and Python path requires CUDA 12+, but exact minimum VRAM is not published. | E The release provides local inference code, not a validated multi-tenant serving recipe; build queueing, batching and observability. |
Workload contract
One request, short context, no SLA; proves runnable, not production-ready.
200 seats, 80-120 DAU, 20-30 peak online, 8-16 concurrent generations, 8K-32K ordinary prompts and occasional 128K agent work.
Public multi-tenant API, redundancy, rolling updates, rate limits and at least 99.9% availability target. Final GPU count requires measured traffic and latency targets.
Model dossiers
Every model entry keeps “documented,” “validated” and “estimated” separate. Commercial summaries are a first-pass product review, not legal advice.
Model category
10 verified deployable models
Released 2026-07-27 · general agent
Moonshot AI · 2.8T total · 104B active · 1M tokens
Best-fit scenarios
Positioned for repository-scale coding and long-horizon agent work.
Capability sourceOfficial materials expose native multimodal inputs and a 1M-token context window.
Capability sourceBest considered when model capability justifies data-center-class full-context workers.
Not published
Vendor documents supported engines but does not publish a smallest runnable GPU configuration.
Requires cluster sizing
Not responsibly specifiable before a concurrency and context benchmark; workstation deployment is not credible.
Validated reference
NVIDIA Dynamo publishes full-1M profiles using 8x GB300 or 16x GB200 per aggregated worker; disaggregated profiles use more GPUs. Primary recipe
Commercial-use check
MaaS / hosted service: Separate agreement required when a MaaS operator and affiliates exceed USD 20M aggregate revenue over any consecutive 12 months.
Attribution: Display Kimi K3 prominently above 100M MAU or USD 20M monthly revenue.
Exceptions: The cited requirements do not apply to internal use or use through Moonshot official products or certified inference partners.
Released 2026-07-12 · general agent
Z.ai · 753B total · not independently verified in V1 active · 1M tokens
Best-fit scenarios
Official release materials emphasize coding-agent and tool-use workloads.
Capability sourceThe official checkpoint supports a 1M-token context window.
Capability sourceA credible fit when CPU/GPU offload or TP8 infrastructure is already available.
Heterogeneous reference
KTransformers documents an 8-GPU CPU/GPU-offload launch shape with 96 CPU inference threads; this is a tutorial target, not a claimed absolute minimum. Primary recipe
Benchmark required
Start from a replicated TP8 service only after measuring the target context mix; no official 200-seat capacity claim exists.
Benchmark required
Requires redundant replicas or disaggregated prefill/decode; GPU count cannot be inferred from seats alone.
Commercial-use check
MaaS / hosted service: No model-specific MaaS restriction identified in the MIT license.
Attribution: Retain the copyright and permission notice.
Released 2026-06-22 · general agent
DeepSeek · 862B reported by model repository total · not independently verified in V1 active · 1M tokens
Best-fit scenarios
The official model card publishes reasoning, coding and agent benchmark results.
Capability sourceThe release is designed around efficient million-token context intelligence.
Capability sourceUse when quality has more value than minimum infrastructure cost.
Validated reference not minimum
vLLM publishes 8x B300, 8x H200, 8x MI355X and GB200 profiles; none is labelled the absolute minimum. Primary recipe
Benchmark required
An 8-GPU worker may be a capacity building block, but replicas depend on concurrency and output-token demand.
Validated reference
8x H200 is documented with context capped at 800K to preserve KV headroom; production redundancy requires additional workers. Primary recipe
Commercial-use check
MaaS / hosted service: No model-specific MaaS restriction identified in the MIT license.
Attribution: Retain the copyright and permission notice.
Released 2026-06-01 · general agent
MiniMax · approximately 428B total · approximately 23B active · 1M tokens
Best-fit scenarios
Official materials position M3 for multimodal perception and agent tasks.
Capability sourceThe official release reports coding and agent evaluations.
Capability sourcePotentially suitable after the revenue, notice and attribution clauses are cleared.
Candidate recipe under review
A 4x RTX PRO 6000 NVFP4 recipe is under review upstream; V1 does not promote an unmerged recipe to validated minimum.
Benchmark required
A quantized multi-GPU worker is plausible, but the 200-seat recommendation needs measured TTFT, throughput and context mix.
Framework support maturing
Aggregated and disaggregated vLLM recipes are still landing; pinning a nightly build may be required.
Commercial-use check
MaaS / hosted service: Commercial API and hosted use are Commercial Use. Above USD 20M yearly revenue obtain prior written authorization; otherwise send the required one-time notice.
Attribution: Prominently display Built with MiniMax M3 for commercial use.
Prohibited uses: The license includes specified unlawful, military and harmful-use restrictions.
Released 2026-04-16 · general agent
Alibaba Qwen · 35B total · 3B active · 256K tokens
Best-fit scenarios
The 35B/3B-active shape and broad serving support make it a practical baseline for controlled workloads.
Qwen3.6 is released as a native multimodal model family.
Capability sourceOfficial serving examples include tool parsing and long-context operation.
Capability sourceOfficial launch shape
Official serving examples use TP4 and 262K context. Smaller quantized short-context shapes exist, but V1 does not call them the official minimum.
Best v1 benchmark candidate
Use two measured serving replicas as the initial HA design candidate; exact GPUs remain pending load tests.
Benchmark required
Scale through replicated workers after measuring prefill-heavy and decode-heavy traffic separately.
Commercial-use check
MaaS / hosted service: No model-specific MaaS restriction identified in Apache-2.0.
Attribution: Provide the license and required notices, preserve attribution notices, and mark modified files; trademark rights are not granted.
Released 2026-06-22 · general agent
DeepSeek · 284B total · 13B active · 1M tokens
Best-fit scenarios
The official card reports coding and agent results and compares Flash reasoning modes.
Capability sourceThe model supports a 1M-token context with a 284B/13B-active MoE shape.
Capability sourcePrefer over V4 Pro when infrastructure cost and concurrency matter more than maximum quality.
Vendor run path not minimum
Official local instructions and engine integrations are published, but no absolute minimum GPU count is claimed. Primary recipe
Benchmark required
Use one measured tensor-parallel worker as the capacity baseline; add replicas only after concurrency tests.
Framework supported
vLLM and SGLang support are available; production still requires redundant workers and pinned encoding logic. Primary recipe
Commercial-use check
MaaS / hosted service: No model-specific MaaS restriction identified in the MIT license.
Attribution: Retain the copyright and permission notice.
Released 2026-05-28 · general agent
StepFun · 198B total · approximately 11B active · 256K tokens
Best-fit scenarios
Official materials highlight charts, GUIs, wireframes and structured-code extraction.
Capability sourceThe vendor publishes ClawEval, Toolathlon and tool-use evaluations.
Capability sourceThe model is explicitly engineered for high-frequency production agent workloads.
Capability sourceOfficial local minimum
The official GGUF path specifies about 120GB minimum unified memory/VRAM and recommends 128GB. Primary recipe
Production candidate
Start with one TP4 NVFP4 or TP8 FP8 worker and benchmark the actual image and context mix.
Official launch shapes
Official examples publish TP4 NVFP4 and TP8 FP8/BF16 serving shapes; replicas are still required for HA. Primary recipe
Commercial-use check
MaaS / hosted service: No model-specific MaaS restriction identified in Apache-2.0.
Attribution: Provide the license and notices, preserve attribution notices, and mark modified files.
Released 2026-07-02 · general agent
Meituan · 1.6T total · approximately 48B active · 1M tokens
Best-fit scenarios
Official materials emphasize repository edits and integrations with coding-agent harnesses.
Capability sourceThe vendor publishes coding, BrowseComp, RWSearch and agent evaluations.
Capability sourceAn official SGLang-FluentLLM NPU serving path is linked.
Capability sourceNot published
GPU and NPU paths are documented, but the vendor does not publish an absolute minimum hardware shape.
Cluster sizing required
A 1.6T model needs a measured cluster design; employee count alone is not a capacity input.
Official serving paths
Official GPU and NPU serving paths exist; topology, redundancy and throughput remain operator-specific. Primary recipe
Commercial-use check
MaaS / hosted service: No model-specific MaaS restriction identified in the MIT license.
Attribution: Retain the copyright and permission notice.
Released 2026-04-27 · general agent
Xiaomi MiMo · 310B total · 15B active · 1M tokens
Best-fit scenarios
The checkpoint natively accepts text, image, video and audio.
Capability sourceThe model supports up to 1M context and dedicated visual and audio encoders.
Capability sourceThe official card publishes multimodal, coding, agent and long-context evaluations.
Capability sourceNot published
Transformers can load the checkpoint, but the vendor does not state an absolute minimum GPU configuration.
Benchmark required
Use the official distributed recipe as a starting worker and size replicas from the real modality mix.
Official distributed recipe
The official card shows an FP8 SGLang DP2×TP8 configuration at 262K context. Primary recipe
Commercial-use check
MaaS / hosted service: No model-specific MaaS restriction identified in the MIT license.
Attribution: Retain the copyright and permission notice.
Released 2026-02-16 · general agent
Alibaba Qwen · 9B total · 9B active · 256K tokens
Best-fit scenarios
The 9B dense shape is the most accessible formal model in this ledger.
The official model is a unified vision-language foundation model.
Capability sourceA better first benchmark target than frontier-scale MoE models when concurrency and budget dominate.
Memory estimate not official
A 24GB-class GPU is a reasonable BF16 short-context planning target, but the vendor does not label this an official minimum.
Best low cost candidate
Benchmark one or two 24–48GB workers before considering larger models; exact replicas depend on output length.
Mainstream framework support
Official Transformers, vLLM and SGLang examples are published; add redundant replicas for availability. Primary recipe
Commercial-use check
MaaS / hosted service: No model-specific MaaS restriction identified in Apache-2.0.
Attribution: Provide the license and notices, preserve attribution notices, and mark modified files.
Model category
1 verified deployable model
Released 2026-06-24 · world model
Alibaba Qwen · 35B total · 3B active · 256K tokens
Best-fit scenarios
The model predicts environment transitions across seven unified domains.
Capability sourceOfficial materials highlight controllable simulation and fictional-world construction.
Capability sourceUseful as a simulator component, not as a drop-in customer chatbot.
Official tp4 launch
Official SGLang and vLLM examples use tensor parallel size 4. Primary recipe
Not a seat based service
Size by simulation jobs and trajectory length, not employee seats.
Specialized service
Expose behind a task-specific simulator contract; do not market it as a normal chat-completions quality substitute.
Commercial-use check
MaaS / hosted service: No model-specific MaaS restriction identified in Apache-2.0.
Attribution: Provide the license and notices, preserve attribution notices, and mark modified files.
Model category
1 verified deployable model
Released 2026-02-11 · robotics vla
Xiaomi Robotics · 4.7B total · 4.7B active · Not applicable
Best-fit scenarios
The official deployment guide targets robotic manipulation with asynchronous real-time execution.
Capability sourceOfficial fine-tuned checkpoints and evaluation guides are released for all three suites.
Capability sourcePost-training code is available for adapting the model to new data.
Capability sourceConsumer gpu claim no vram
The vendor says BF16 inference is optimized for consumer GPUs but publishes no exact minimum VRAM. Primary recipe
Robot fleet profile required
Size by robots, camera rate and control latency; a 200-seat office profile is not applicable.
Edge control not public api
Prefer an on-robot or near-edge safety architecture; a remote multi-tenant API is not the default production shape.
Commercial-use check
MaaS / hosted service: No model-specific MaaS restriction identified in Apache-2.0.
Attribution: Provide the license and notices, preserve attribution notices, and mark modified files.
Model category
1 verified deployable model
Released 2026-05-28 · audio video generation
Baidu ERNIE Team · 6.3B backbone total · 6.3B active · Not applicable
Best-fit scenarios
NAVA jointly generates video, scene audio and speech rather than aligning separate outputs after generation.
Capability sourceThe official checkpoint supports up to two reference voices bound to speech spans.
Capability sourceThe vendor reports VerseBench synchronization and quality results plus an 8-GPU fast path.
Capability sourceSingle gpu supported no vram
The model card supports single-GPU inference but does not publish exact minimum VRAM. Primary recipe
Media queue required
Size a queued render farm by jobs per hour, resolution and duration; office-seat assumptions do not apply.
Official 8gpu reference
The official Ulysses SP8 path reports roughly one minute for a 720p synchronized clip; HA needs additional workers. Primary recipe
Commercial-use check
MaaS / hosted service: The model card states Apache-2.0, but bundled LTX audio-VAE artifacts carry an additional community license that must be reviewed for the shipped stack.
Attribution: Preserve Apache notices and the notices/licenses for bundled upstream components.
Prohibited uses: The model card prohibits depicting real persons without consent, including face or voice likeness reproduction.
Model category
1 verified deployable model
Released 2026-04-23 · speech recognition
Xiaomi MiMo · not published total · not published active · Not applicable
Best-fit scenarios
Official support includes Wu, Cantonese, Hokkien, Sichuanese and Chinese-English code switching.
Capability sourceThe release targets overlapping speakers, heavy noise and far-field capture.
Capability sourceThe vendor reports evaluations across dialects, lyrics and complex English scenarios.
Capability sourceNot published
The official local Gradio and Python path requires CUDA 12+, but exact minimum VRAM is not published. Primary recipe
Audio workload profile required
Size by audio hours, peak simultaneous streams and latency rather than office seats.
Custom service required
The release provides local inference code, not a validated multi-tenant serving recipe; build queueing, batching and observability.
Commercial-use check
MaaS / hosted service: No model-specific MaaS restriction identified in Apache-2.0.
Attribution: Provide the license and notices, preserve attribution notices, and mark modified files.
Research watchlist
These projects have visible open-model value, but unclear commercial rights, partial artifacts or an unconfirmed directly deployable checkpoint keep them outside the verified count.
License unclear
Official weights and inference code exist, but the repository does not clearly name a standard model license; commercial rights need clarification.
Official projectPartial release
Official code and selected weights are available as a research preview, but the complete deployment surface is not yet equivalent to the formal ledger.
Official projectWeights not confirmed
The official repository is public, but a complete, directly deployable released checkpoint was not sufficiently confirmed at this cutoff.
Official projectMethodology
Release window: 2025-08-03 through 2026-08-03. Included models have downloadable official weights and at least one documented executable self-hosting path. API-only announcements, unreleased weights, community-only ports and research checkpoints without a serving path are excluded.
Important: loading weights, completing the first forward pass and meeting a production latency/SLA are three separate thresholds. Context length also changes KV-cache demand, so seat count alone cannot produce a reliable GPU bill.
Quick answers
Short answers for buyers, infrastructure teams and AI agents. The evidence ledger above is authoritative when a summary and a model-specific row differ.
As of 2026-08-03, this verified index includes Kimi K3, GLM-5.2, DeepSeek V4 Pro, MiniMax M3, Qwen3.6-35B-A3B, DeepSeek V4 Flash, Step 3.7 Flash, LongCat 2.0, MiMo V2.5, Qwen3.5-9B, Qwen-AgentWorld-35B-A3B, Xiaomi-Robotics-0, NAVA, MiMo-V2.5-ASR. Each has downloadable official weights and at least one documented executable serving path.
There is no single trustworthy minimum. We report the smallest official or inference-framework launch shape found and label its evidence source. A launch shape is not a production capacity guarantee.
GPU count depends on DAU, peak concurrency, prompt and output length, context mix, latency targets and redundancy. Our comparison assumes 80–120 DAU, 20–30 peak online and 8–16 concurrent generations, then marks where load testing is still required.
MIT and Apache-2.0 releases generally permit commercial use with notice obligations, while community licenses and bundled dependencies can add revenue, authorization, attribution or use restrictions. Review the complete license stack before launch.
Start with the workload: general agents, environment simulation, physical-world robot control, joint audio-video generation and ASR are not interchangeable. The scenario guide labels vendor claims, public evidence and ChinaAPI deployment inferences separately.
Prefer an API?
Use the same workload assumptions to compare self-hosting against an OpenAI-compatible managed endpoint.