# Chinese Open Models: Deployment Hardware and Commercial Licenses

> Source-linked deployment evidence for downloadable Chinese open-weight foundation models. Estimates are labelled as estimates.

- Verified: 2026-08-03
- Release window: 2025-08-03 to 2026-08-03
- Canonical: https://chinaapi.ai/open-model-deployment/
- JSON: https://chinaapi.ai/data/open-model-deployment.json
- Terminology: this index uses the common search term ‘open model,’ but records the actual license. Downloadable weights do not automatically imply OSI-style open source.

## Workload profiles

- Official launch / minimum: One request, short context, no SLA; proves runnable, not production-ready.
- 200-person team: 200 seats, 80-120 DAU, 20-30 peak online, 8-16 concurrent generations, 8K-32K ordinary prompts and occasional 128K agent work.
- Commercial API: Public multi-tenant API, redundancy, rolling updates, rate limits and at least 99.9% availability target. Final GPU count requires measured traffic and latency targets.

## Model evidence ledger

### Kimi K3

- Vendor: Moonshot AI
- Released: 2026-07-27
- Parameters: 2.8T total; 104B active
- Context: 1M tokens
- Category: General and agentic foundation models
- Modalities: text, image, video → text
- Best-fit summary: Frontier-scale long-context multimodal and coding agent workloads where a cluster deployment is acceptable.
- Official weights: https://huggingface.co/moonshotai/Kimi-K3?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- Serving paths: vLLM, SGLang, TokenSpeed
- License: Kimi K3 License — https://github.com/MoonshotAI/Kimi-K3/blob/main/LICENSE?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- MaaS / hosted service: Separate agreement required when a MaaS operator and affiliates exceed USD 20M aggregate revenue over any consecutive 12 months.
- Attribution: Display Kimi K3 prominently above 100M MAU or USD 20M monthly revenue.

- **Official launch / minimum — evidence C:** Vendor documents supported engines but does not publish a smallest runnable GPU configuration.
- **200-person team — evidence E:** Not responsibly specifiable before a concurrency and context benchmark; workstation deployment is not credible.
- **Commercial API — evidence B:** NVIDIA Dynamo publishes full-1M profiles using 8x GB300 or 16x GB200 per aggregated worker; disaggregated profiles use more GPUs. Source: https://github.com/ai-dynamo/dynamo/blob/main/recipes/kimi-k3/README.md?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment

#### Scenario evidence

- **Long-context coding agents — Vendor-stated:** Positioned for repository-scale coding and long-horizon agent work. Source: https://github.com/MoonshotAI/Kimi-K3?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Multimodal document and video analysis — Vendor-stated:** Official materials expose native multimodal inputs and a 1M-token context window. Source: https://github.com/MoonshotAI/Kimi-K3?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **High-value enterprise agent workflows — ChinaAPI inference:** Best considered when model capability justifies data-center-class full-context workers.

Not recommended: Single-workstation deployments; Low-cost high-QPS chat without aggressive batching or distillation

Limitations: No ChinaAPI hardware reproduction; Full-context reference configurations are data-center cluster class

### GLM-5.2

- Vendor: Z.ai
- Released: 2026-07-12
- Parameters: 753B total; not independently verified in V1 active
- Context: 1M tokens
- Category: General and agentic foundation models
- Modalities: text → text
- Best-fit summary: Long-horizon software engineering, tool use and million-token analysis on enterprise infrastructure.
- Official weights: https://huggingface.co/zai-org/GLM-5.2?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- Serving paths: vLLM, SGLang, xLLM, KTransformers
- License: MIT — https://huggingface.co/zai-org/GLM-5.2/blob/main/LICENSE?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- MaaS / hosted service: No model-specific MaaS restriction identified in the MIT license.
- Attribution: Retain the copyright and permission notice.

- **Official launch / minimum — evidence B:** KTransformers documents an 8-GPU CPU/GPU-offload launch shape with 96 CPU inference threads; this is a tutorial target, not a claimed absolute minimum. Source: https://github.com/kvcache-ai/ktransformers/blob/main/doc/en/kt-kernel/GLM-5.2-Tutorial.md?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **200-person team — evidence E:** Start from a replicated TP8 service only after measuring the target context mix; no official 200-seat capacity claim exists.
- **Commercial API — evidence E:** Requires redundant replicas or disaggregated prefill/decode; GPU count cannot be inferred from seats alone.

#### Scenario evidence

- **Long-horizon software engineering — Vendor-stated:** Official release materials emphasize coding-agent and tool-use workloads. Source: https://github.com/zai-org/GLM-5?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Million-token repository and document analysis — Vendor-stated:** The official checkpoint supports a 1M-token context window. Source: https://huggingface.co/zai-org/GLM-5.2?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Private enterprise agent service — ChinaAPI inference:** A credible fit when CPU/GPU offload or TP8 infrastructure is already available.

Not recommended: Latency-sensitive single-GPU serving; Capacity planning based only on employee count

Limitations: The 8-GPU tutorial is not an official minimum claim; No public V1 concurrency result

### DeepSeek V4 Pro

- Vendor: DeepSeek
- Released: 2026-06-22
- Parameters: 862B reported by model repository total; not independently verified in V1 active
- Context: 1M tokens
- Category: General and agentic foundation models
- Modalities: text → text
- Best-fit summary: Maximum-quality reasoning, coding and long-context agents when an eight-accelerator worker is viable.
- Official weights: https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- Serving paths: vLLM, SGLang
- License: MIT — https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/LICENSE?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- MaaS / hosted service: No model-specific MaaS restriction identified in the MIT license.
- Attribution: Retain the copyright and permission notice.

- **Official launch / minimum — evidence B:** vLLM publishes 8x B300, 8x H200, 8x MI355X and GB200 profiles; none is labelled the absolute minimum. Source: https://github.com/vllm-project/recipes/blob/main/models/deepseek-ai/DeepSeek-V4-Pro.yaml?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **200-person team — evidence E:** An 8-GPU worker may be a capacity building block, but replicas depend on concurrency and output-token demand.
- **Commercial API — evidence B:** 8x H200 is documented with context capped at 800K to preserve KV headroom; production redundancy requires additional workers. Source: https://github.com/vllm-project/recipes/blob/main/models/deepseek-ai/DeepSeek-V4-Pro.yaml?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment

#### Scenario evidence

- **Complex reasoning and coding — Public evidence:** The official model card publishes reasoning, coding and agent benchmark results. Source: https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Million-token retrieval and analysis — Vendor-stated:** The release is designed around efficient million-token context intelligence. Source: https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Premium private agent endpoint — ChinaAPI inference:** Use when quality has more value than minimum infrastructure cost.

Not recommended: Budget workstation inference; Public API launch without at least one redundant worker

Limitations: Some SM120 workstation kernels have unresolved reports; Loading weights is not proof that the first forward pass succeeds

### MiniMax M3

- Vendor: MiniMax
- Released: 2026-06-01
- Parameters: approximately 428B total; approximately 23B active
- Context: 1M tokens
- Category: General and agentic foundation models
- Modalities: text, image, video → text
- Best-fit summary: Multimodal agent, coding and long-context workloads with explicit commercial-license review.
- Official weights: https://huggingface.co/MiniMaxAI/MiniMax-M3?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- Serving paths: Transformers, vLLM nightly, SGLang
- License: MiniMax Community License — https://huggingface.co/MiniMaxAI/MiniMax-M3/blob/main/LICENSE?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- MaaS / hosted service: Commercial API and hosted use are Commercial Use. Above USD 20M yearly revenue obtain prior written authorization; otherwise send the required one-time notice.
- Attribution: Prominently display Built with MiniMax M3 for commercial use.

- **Official launch / minimum — evidence D:** A 4x RTX PRO 6000 NVFP4 recipe is under review upstream; V1 does not promote an unmerged recipe to validated minimum.
- **200-person team — evidence E:** A quantized multi-GPU worker is plausible, but the 200-seat recommendation needs measured TTFT, throughput and context mix.
- **Commercial API — evidence B:** Aggregated and disaggregated vLLM recipes are still landing; pinning a nightly build may be required.

#### Scenario evidence

- **Multimodal agent workflows — Vendor-stated:** Official materials position M3 for multimodal perception and agent tasks. Source: https://github.com/MiniMax-AI/MiniMax-M3?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Coding and tool orchestration — Public evidence:** The official release reports coding and agent evaluations. Source: https://huggingface.co/MiniMaxAI/MiniMax-M3?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Commercial embedded assistant — ChinaAPI inference:** Potentially suitable after the revenue, notice and attribution clauses are cleared.

Not recommended: Commercial launch before license notice and attribution review; Stable-production serving without pinned framework versions

Limitations: Stable vLLM release support was not complete at the V1 cutoff; Do not equate a pending recipe with successful ChinaAPI reproduction

### Qwen3.6-35B-A3B

- Vendor: Alibaba Qwen
- Released: 2026-04-16
- Parameters: 35B total; 3B active
- Context: 256K tokens
- Category: General and agentic foundation models
- Modalities: text, image, video → text
- Best-fit summary: A comparatively deployable private multimodal agent baseline with strong ecosystem coverage.
- Official weights: https://huggingface.co/Qwen/Qwen3.6-35B-A3B?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- Serving paths: Transformers, vLLM, SGLang, llama.cpp, MLX
- License: Apache-2.0 — https://huggingface.co/Qwen/Qwen3.6-35B-A3B/blob/main/LICENSE?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- MaaS / hosted service: No model-specific MaaS restriction identified in Apache-2.0.
- Attribution: Provide the license and required notices, preserve attribution notices, and mark modified files; trademark rights are not granted.

- **Official launch / minimum — evidence C:** Official serving examples use TP4 and 262K context. Smaller quantized short-context shapes exist, but V1 does not call them the official minimum.
- **200-person team — evidence E:** Use two measured serving replicas as the initial HA design candidate; exact GPUs remain pending load tests.
- **Commercial API — evidence E:** Scale through replicated workers after measuring prefill-heavy and decode-heavy traffic separately.

#### Scenario evidence

- **Private coding and office assistant — ChinaAPI inference:** The 35B/3B-active shape and broad serving support make it a practical baseline for controlled workloads.
- **Visual document and UI understanding — Vendor-stated:** Qwen3.6 is released as a native multimodal model family. Source: https://github.com/QwenLM/Qwen3.6?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Tool-calling agent service — Vendor-stated:** Official serving examples include tool parsing and long-context operation. Source: https://huggingface.co/Qwen/Qwen3.6-35B-A3B?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment

Not recommended: Assuming TP4 equals four production replicas; Unvalidated parser upgrades in a critical tool loop

Limitations: Official TP4 example is a launch shape, not a capacity guarantee; Tool-call parser issues have been reported for related Qwen3.5 configurations

### DeepSeek V4 Flash

- Vendor: DeepSeek
- Released: 2026-06-22
- Parameters: 284B total; 13B active
- Context: 1M tokens
- Category: General and agentic foundation models
- Modalities: text → text
- Best-fit summary: A smaller DeepSeek V4 worker for high-frequency coding, reasoning and agent traffic.
- Official weights: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- Serving paths: Transformers, vLLM, SGLang
- License: MIT — https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash/blob/main/LICENSE?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- MaaS / hosted service: No model-specific MaaS restriction identified in the MIT license.
- Attribution: Retain the copyright and permission notice.

- **Official launch / minimum — evidence C:** Official local instructions and engine integrations are published, but no absolute minimum GPU count is claimed. Source: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **200-person team — evidence E:** Use one measured tensor-parallel worker as the capacity baseline; add replicas only after concurrency tests.
- **Commercial API — evidence B:** vLLM and SGLang support are available; production still requires redundant workers and pinned encoding logic. Source: https://github.com/vllm-project/vllm-project.github.io/blob/main/_posts/2026-04-24-deepseek-v4.md?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment

#### Scenario evidence

- **High-frequency coding agents — Public evidence:** The official card reports coding and agent results and compares Flash reasoning modes. Source: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Million-token analysis — Vendor-stated:** The model supports a 1M-token context with a 284B/13B-active MoE shape. Source: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Throughput-sensitive private endpoint — ChinaAPI inference:** Prefer over V4 Pro when infrastructure cost and concurrency matter more than maximum quality.

Not recommended: Treating Flash as quality-equivalent to Pro on every task; Launching a public API without measured parser and long-context behavior

Limitations: No ChinaAPI hardware reproduction; The release uses a dedicated encoding implementation rather than a Jinja chat template

### Step 3.7 Flash

- Vendor: StepFun
- Released: 2026-05-28
- Parameters: 198B total; approximately 11B active
- Context: 256K tokens
- Category: General and agentic foundation models
- Modalities: text, image → text
- Best-fit summary: High-frequency multimodal agents, financial-document parsing, verified search loops and concurrent coding.
- Official weights: https://huggingface.co/stepfun-ai/Step-3.7-Flash?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- Serving paths: vLLM, SGLang, Transformers, llama.cpp
- License: Apache-2.0 — https://github.com/stepfun-ai/Step-3.7-Flash/blob/main/LICENSE?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- MaaS / hosted service: No model-specific MaaS restriction identified in Apache-2.0.
- Attribution: Provide the license and notices, preserve attribution notices, and mark modified files.

- **Official launch / minimum — evidence C:** The official GGUF path specifies about 120GB minimum unified memory/VRAM and recommends 128GB. Source: https://github.com/stepfun-ai/Step-3.7-Flash?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **200-person team — evidence E:** Start with one TP4 NVFP4 or TP8 FP8 worker and benchmark the actual image and context mix.
- **Commercial API — evidence C:** Official examples publish TP4 NVFP4 and TP8 FP8/BF16 serving shapes; replicas are still required for HA. Source: https://github.com/stepfun-ai/Step-3.7-Flash?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment

#### Scenario evidence

- **Visual documents and UI-to-code — Vendor-stated:** Official materials highlight charts, GUIs, wireframes and structured-code extraction. Source: https://github.com/stepfun-ai/Step-3.7-Flash?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Search and tool orchestration — Public evidence:** The vendor publishes ClawEval, Toolathlon and tool-use evaluations. Source: https://github.com/stepfun-ai/Step-3.7-Flash?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Concurrent coding agents — Vendor-stated:** The model is explicitly engineered for high-frequency production agent workloads. Source: https://github.com/stepfun-ai/Step-3.7-Flash?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment

Not recommended: Using the 128GB local path as a production throughput claim; Assuming vendor benchmark throughput transfers to long-context prefill

Limitations: Published 400 tok/s is a vendor benchmark, not a universal SLA; NVFP4 support depends on recent engine and GPU paths

### LongCat 2.0

- Vendor: Meituan
- Released: 2026-07-02
- Parameters: 1.6T total; approximately 48B active
- Context: 1M tokens
- Category: General and agentic foundation models
- Modalities: text → text
- Best-fit summary: Long-horizon coding, search, repository edits and tool-driven agents on GPU or NPU clusters.
- Official weights: https://huggingface.co/meituan-longcat/LongCat-2.0?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- Serving paths: SGLang, SGLang-FluentLLM
- License: MIT — https://github.com/meituan-longcat/LongCat-2.0/blob/main/LICENSE?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- MaaS / hosted service: No model-specific MaaS restriction identified in the MIT license.
- Attribution: Retain the copyright and permission notice.

- **Official launch / minimum — evidence C:** GPU and NPU paths are documented, but the vendor does not publish an absolute minimum hardware shape.
- **200-person team — evidence E:** A 1.6T model needs a measured cluster design; employee count alone is not a capacity input.
- **Commercial API — evidence C:** Official GPU and NPU serving paths exist; topology, redundancy and throughput remain operator-specific. Source: https://github.com/meituan-longcat/LongCat-2.0?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment

#### Scenario evidence

- **Repository-scale coding — Vendor-stated:** Official materials emphasize repository edits and integrations with coding-agent harnesses. Source: https://github.com/meituan-longcat/LongCat-2.0?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Search and general agents — Public evidence:** The vendor publishes coding, BrowseComp, RWSearch and agent evaluations. Source: https://github.com/meituan-longcat/LongCat-2.0?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **NPU-based sovereign deployment — Vendor-stated:** An official SGLang-FluentLLM NPU serving path is linked. Source: https://github.com/meituan-longcat/LongCat-2.0?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment

Not recommended: Workstation deployment; Quoting in-house benchmark results as ChinaAPI reproduction

Limitations: No official minimum GPU count; Most published evaluation values are vendor-measured

### MiMo V2.5

- Vendor: Xiaomi MiMo
- Released: 2026-04-27
- Parameters: 310B total; 15B active
- Context: 1M tokens
- Category: General and agentic foundation models
- Modalities: text, image, video, audio → text
- Best-fit summary: Native omnimodal understanding, long-context reasoning and agentic workflows across text, image, video and audio.
- Official weights: https://huggingface.co/XiaomiMiMo/MiMo-V2.5?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- Serving paths: Transformers, SGLang, vLLM
- License: MIT — https://huggingface.co/XiaomiMiMo/MiMo-V2.5/blob/main/LICENSE?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- MaaS / hosted service: No model-specific MaaS restriction identified in the MIT license.
- Attribution: Retain the copyright and permission notice.

- **Official launch / minimum — evidence C:** Transformers can load the checkpoint, but the vendor does not state an absolute minimum GPU configuration.
- **200-person team — evidence E:** Use the official distributed recipe as a starting worker and size replicas from the real modality mix.
- **Commercial API — evidence C:** The official card shows an FP8 SGLang DP2×TP8 configuration at 262K context. Source: https://huggingface.co/XiaomiMiMo/MiMo-V2.5?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment

#### Scenario evidence

- **Omnimodal research and support agents — Vendor-stated:** The checkpoint natively accepts text, image, video and audio. Source: https://huggingface.co/XiaomiMiMo/MiMo-V2.5?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Long video, audio and document analysis — Vendor-stated:** The model supports up to 1M context and dedicated visual and audio encoders. Source: https://huggingface.co/XiaomiMiMo/MiMo-V2.5?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Multimodal tool-using agents — Public evidence:** The official card publishes multimodal, coding, agent and long-context evaluations. Source: https://huggingface.co/XiaomiMiMo/MiMo-V2.5?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment

Not recommended: Small single-GPU deployment; Using stale config or tokenizer files from the initial release

Limitations: Official deployment example is not a minimum; Audio and video traffic need separate encoder-capacity measurements

### Qwen3.5-9B

- Vendor: Alibaba Qwen
- Released: 2026-02-16
- Parameters: 9B total; 9B active
- Context: 256K tokens
- Category: General and agentic foundation models
- Modalities: text, image, video → text
- Best-fit summary: The local and edge-friendly baseline for private multimodal assistants, extraction and moderate-volume APIs.
- Official weights: https://huggingface.co/Qwen/Qwen3.5-9B?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- Serving paths: Transformers, vLLM, SGLang, KTransformers, llama.cpp
- License: Apache-2.0 — https://huggingface.co/Qwen/Qwen3.5-9B/blob/main/LICENSE?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- MaaS / hosted service: No model-specific MaaS restriction identified in Apache-2.0.
- Attribution: Provide the license and notices, preserve attribution notices, and mark modified files.

- **Official launch / minimum — evidence E:** A 24GB-class GPU is a reasonable BF16 short-context planning target, but the vendor does not label this an official minimum.
- **200-person team — evidence E:** Benchmark one or two 24–48GB workers before considering larger models; exact replicas depend on output length.
- **Commercial API — evidence C:** Official Transformers, vLLM and SGLang examples are published; add redundant replicas for availability. Source: https://huggingface.co/Qwen/Qwen3.5-9B?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment

#### Scenario evidence

- **Local multimodal assistant — ChinaAPI inference:** The 9B dense shape is the most accessible formal model in this ledger.
- **Visual extraction and classification — Vendor-stated:** The official model is a unified vision-language foundation model. Source: https://huggingface.co/Qwen/Qwen3.5-9B?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Cost-sensitive private API — ChinaAPI inference:** A better first benchmark target than frontier-scale MoE models when concurrency and budget dominate.

Not recommended: Assuming 1M extended context fits a 24GB card; Highest-complexity long-horizon agents without task-specific evaluation

Limitations: 24GB guidance is an estimate, not ChinaAPI reproduction; Native 262K and extended 1M contexts materially increase KV-cache demand

### Qwen-AgentWorld-35B-A3B

- Vendor: Alibaba Qwen
- Released: 2026-06-24
- Parameters: 35B total; 3B active
- Context: 256K tokens
- Category: Digital-world and environment models
- Modalities: text, environment state, action history → predicted environment state
- Best-fit summary: A language world model for simulating MCP, Search, Terminal, SWE, Android, Web and OS agent environments.
- Official weights: https://huggingface.co/Qwen/Qwen-AgentWorld-35B-A3B?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- Serving paths: SGLang, vLLM
- License: Apache-2.0 — https://github.com/QwenLM/Qwen-AgentWorld/blob/main/LICENSE?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- MaaS / hosted service: No model-specific MaaS restriction identified in Apache-2.0.
- Attribution: Provide the license and notices, preserve attribution notices, and mark modified files.

- **Official launch / minimum — evidence C:** Official SGLang and vLLM examples use tensor parallel size 4. Source: https://github.com/QwenLM/Qwen-AgentWorld?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **200-person team — evidence E:** Size by simulation jobs and trajectory length, not employee seats.
- **Commercial API — evidence E:** Expose behind a task-specific simulator contract; do not market it as a normal chat-completions quality substitute.

#### Scenario evidence

- **Agent environment simulation — Vendor-stated:** The model predicts environment transitions across seven unified domains. Source: https://github.com/QwenLM/Qwen-AgentWorld?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Synthetic trajectories and perturbation tests — Vendor-stated:** Official materials highlight controllable simulation and fictional-world construction. Source: https://github.com/QwenLM/Qwen-AgentWorld?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Agent regression evaluation — ChinaAPI inference:** Useful as a simulator component, not as a drop-in customer chatbot.

Not recommended: General-purpose chat replacement; Treating simulated success as proof of real-environment reliability

Limitations: World-model outputs are simulations; The 35B release covers seven named domains, not arbitrary physical environments

### Xiaomi-Robotics-0

- Vendor: Xiaomi Robotics
- Released: 2026-02-11
- Parameters: 4.7B total; 4.7B active
- Context: Not applicable
- Category: Physical-world vision-language-action models
- Modalities: robot camera, language instruction, proprioception → robot action
- Best-fit summary: Real-time robotic manipulation research and post-training across supported embodiments and simulation suites.
- Official weights: https://huggingface.co/XiaomiRobotics/Xiaomi-Robotics-0?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- Serving paths: Transformers, PyTorch
- License: Apache-2.0 — https://github.com/XiaomiRobotics/Xiaomi-Robotics-0/blob/main/LICENSE?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- MaaS / hosted service: No model-specific MaaS restriction identified in Apache-2.0.
- Attribution: Provide the license and notices, preserve attribution notices, and mark modified files.

- **Official launch / minimum — evidence C:** The vendor says BF16 inference is optimized for consumer GPUs but publishes no exact minimum VRAM. Source: https://github.com/XiaomiRobotics/Xiaomi-Robotics-0?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **200-person team — evidence E:** Size by robots, camera rate and control latency; a 200-seat office profile is not applicable.
- **Commercial API — evidence E:** Prefer an on-robot or near-edge safety architecture; a remote multi-tenant API is not the default production shape.

#### Scenario evidence

- **Robot manipulation — Vendor-stated:** The official deployment guide targets robotic manipulation with asynchronous real-time execution. Source: https://github.com/XiaomiRobotics/Xiaomi-Robotics-0?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **LIBERO, CALVIN and SimplerEnv evaluation — Public evidence:** Official fine-tuned checkpoints and evaluation guides are released for all three suites. Source: https://github.com/XiaomiRobotics/Xiaomi-Robotics-0?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Embodiment-specific post-training — Vendor-stated:** Post-training code is available for adapting the model to new data. Source: https://github.com/XiaomiRobotics/Xiaomi-Robotics-0?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment

Not recommended: Direct deployment on an unvalidated physical robot; Safety-critical control without independent interlocks and task-specific validation

Limitations: Simulation benchmarks do not prove physical-world safety; Exact VRAM and end-to-end control latency are not published

### NAVA

- Vendor: Baidu ERNIE Team
- Released: 2026-05-28
- Parameters: 6.3B backbone total; 6.3B active
- Context: Not applicable
- Category: Joint audio-video generation
- Modalities: text, image, reference voice → video, stereo audio, speech
- Best-fit summary: Synchronized audio-video generation with multi-speaker timbre control, camera prompting and image continuation.
- Official weights: https://huggingface.co/baidu/NAVA?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- Serving paths: PyTorch, Ulysses sequence parallel
- License: Apache-2.0 — https://huggingface.co/baidu/NAVA/blob/main/LICENSE?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- MaaS / hosted service: The model card states Apache-2.0, but bundled LTX audio-VAE artifacts carry an additional community license that must be reviewed for the shipped stack.
- Attribution: Preserve Apache notices and the notices/licenses for bundled upstream components.

- **Official launch / minimum — evidence C:** The model card supports single-GPU inference but does not publish exact minimum VRAM. Source: https://huggingface.co/baidu/NAVA?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **200-person team — evidence E:** Size a queued render farm by jobs per hour, resolution and duration; office-seat assumptions do not apply.
- **Commercial API — evidence C:** The official Ulysses SP8 path reports roughly one minute for a 720p synchronized clip; HA needs additional workers. Source: https://huggingface.co/baidu/NAVA?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment

#### Scenario evidence

- **Synchronized short-form audio-video — Vendor-stated:** NAVA jointly generates video, scene audio and speech rather than aligning separate outputs after generation. Source: https://huggingface.co/baidu/NAVA?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Multi-speaker and reference-timbre scenes — Vendor-stated:** The official checkpoint supports up to two reference voices bound to speech spans. Source: https://huggingface.co/baidu/NAVA?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **720p creative generation — Public evidence:** The vendor reports VerseBench synchronization and quality results plus an 8-GPU fast path. Source: https://huggingface.co/baidu/NAVA?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment

Not recommended: Unconsented face or voice cloning; Low-latency interactive video generation

Limitations: Default clips are about 6–10 seconds; The full dependency stack includes component-specific notices beyond the headline Apache license

### MiMo-V2.5-ASR

- Vendor: Xiaomi MiMo
- Released: 2026-04-23
- Parameters: not published total; not published active
- Context: Not applicable
- Category: Automatic speech recognition
- Modalities: audio → punctuated text
- Best-fit summary: Chinese and English transcription across dialects, code-switching, meetings, songs, noise and knowledge-heavy speech.
- Official weights: https://huggingface.co/XiaomiMiMo/MiMo-V2.5-ASR?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- Serving paths: PyTorch, Gradio
- License: Apache-2.0 — https://github.com/XiaomiMiMo/MiMo-V2.5-ASR/blob/main/LICENSE?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- MaaS / hosted service: No model-specific MaaS restriction identified in Apache-2.0.
- Attribution: Provide the license and notices, preserve attribution notices, and mark modified files.

- **Official launch / minimum — evidence C:** The official local Gradio and Python path requires CUDA 12+, but exact minimum VRAM is not published. Source: https://github.com/XiaomiMiMo/MiMo-V2.5-ASR?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **200-person team — evidence E:** Size by audio hours, peak simultaneous streams and latency rather than office seats.
- **Commercial API — evidence E:** The release provides local inference code, not a validated multi-tenant serving recipe; build queueing, batching and observability.

#### Scenario evidence

- **Chinese dialect and code-switch transcription — Vendor-stated:** Official support includes Wu, Cantonese, Hokkien, Sichuanese and Chinese-English code switching. Source: https://github.com/XiaomiMiMo/MiMo-V2.5-ASR?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Meetings and noisy far-field audio — Vendor-stated:** The release targets overlapping speakers, heavy noise and far-field capture. Source: https://github.com/XiaomiMiMo/MiMo-V2.5-ASR?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Knowledge-heavy and lyric transcription — Public evidence:** The vendor reports evaluations across dialects, lyrics and complex English scenarios. Source: https://github.com/XiaomiMiMo/MiMo-V2.5-ASR?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment

Not recommended: Capacity promises before real-time-factor and batch testing; Assuming speaker overlap performance replaces diarization requirements

Limitations: Model parameter count and minimum VRAM are not published; The public release does not provide an official high-throughput serving benchmark

## Research watchlist

- **ByteDance Lance — License unclear:** Official weights and inference code exist, but the repository does not clearly name a standard model license; commercial rights need clarification. Source: https://github.com/bytedance/Lance?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **GigaWorld-1 — Partial release:** Official code and selected weights are available as a research preview, but the complete deployment surface is not yet equivalent to the formal ledger. Source: https://github.com/open-gigaai/giga-world-1?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Qwen-VLA — Weights not confirmed:** The official repository is public, but a complete, directly deployable released checkpoint was not sufficiently confirmed at this cutoff. Source: https://github.com/QwenLM/Qwen-VLA?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment

## Evidence scale

- **A:** ChinaAPI reproduced
- **B:** inference-framework official validated recipe
- **C:** model-vendor documented configuration
- **D:** third-party reproduction
- **E:** capacity estimate only

## Methodology and limitations

Inclusion rule: Official weights are downloadable and at least one executable self-hosting path is documented.

Loading weights, completing a first forward pass, and meeting a production latency/SLA are separate thresholds. Context length changes KV-cache demand, so seats alone cannot determine GPU count. Scenario claims distinguish vendor positioning, public evidence and ChinaAPI deployment inference. License summaries are product research, not legal advice; read the linked primary licenses and obtain counsel before commercial launch.

## FAQ

### Which Chinese open models released in the last year can be self-hosted?

As of 2026-08-03, this verified index includes Kimi K3, GLM-5.2, DeepSeek V4 Pro, MiniMax M3, Qwen3.6-35B-A3B, DeepSeek V4 Flash, Step 3.7 Flash, LongCat 2.0, MiMo V2.5, Qwen3.5-9B, Qwen-AgentWorld-35B-A3B, Xiaomi-Robotics-0, NAVA, MiMo-V2.5-ASR. Each has downloadable official weights and at least one documented executable serving path.

### What is the minimum hardware required?

There is no single trustworthy minimum. The index reports the smallest documented launch shape and its evidence level. A launch shape is not a production capacity guarantee.

### What hardware does a 200-person team need?

GPU count depends on DAU, peak concurrency, token and context mix, latency targets and redundancy. The defined comparison profile is 80-120 DAU, 20-30 peak online and 8-16 concurrent generations for conversational models; robotics and media models use workload-specific profiles.

### Can these models be used commercially or offered as an API?

MIT and Apache-2.0 releases generally permit commercial use with notice obligations, while community licenses and bundled dependencies can add conditions. Review the complete license stack before launch.

### Which model is best for my use case?

Start with the workload category and read the claim label. General agents, environment simulators, robotics VLA, audio-video generators and ASR models are not interchangeable.

Generated by GENERATED BY scripts/gen_open_model_deployment.py.
