Open model deployment index · verified 2026-09-12

What does it actually take to run China’s open frontier models?

Compare the full verified directory here, then open one focused dossier for primary sources, launch configurations, 200-person team assumptions, API architecture and commercial-use restrictions.

31
verified deployable models
7
model categories
3
deployment tiers
Abstract data-center hardware capacity blocks in ChinaAPI blue, violet and cyan
Capacity shape, not a benchmark result.

Model directory

Find the right dossier without scrolling through every model.

Search names, vendors, serving frameworks and use cases, or narrow the directory by category and commercial-use status. Each result opens a dedicated evidence page.

31 of 31 dossiers shown

ModelBest fitFootprintLaunch evidenceLicense
DeepSeek V4.1 FlashDeepSeek · released 2026-09-10 Frontier-scale private multimodal coding, document and long-context agent workloads where a 500-plus-GB checkpoint and dedicated serving build are acceptable. 552B backbone + 196B Engram memory1M tokens · text, image BOfficial capacity floor Commercial use: permittedMIT View dossier
GLM-5.3Z.ai · released 2026-08-28 Frontier text-only coding, cybersecurity, tool-use and long-horizon agent workloads that can support an eight-accelerator FP8 worker or a multi-node BF16 deployment. 753B checkpoint; official serving guides describe an approximately 743B-744B architecture1M tokens · text BOfficial single node fp8 Commercial use: conditionalGLM-5.3 License View dossier
AuK / AuK-FlashTencent Hunyuan · released 2026-09-09 Instruction-driven speech generation, voice cloning, editing, enhancement and source separation where a self-hosted unified audio workflow is more important than a managed API. 1.5B family modelNot applicable · text, audio COfficial single gpu launch path not hardware minimum Commercial use: permittedMIT View dossier
Qwen3.8-Flash-NextAlibaba Qwen · released 2026-08-26 Cost-sensitive coding, office and multimodal agent workloads that need strong public benchmark results but can accept preview-stage serving software or a managed production derivative. 125B main + 51B N-gram embedding + 4B MTP256K tokens · text, image, video BOfficial validated minimum Commercial use: conditionalQwen Community License 1.0 View dossier
Hy4 previewTencent Hy Team · released 2026-08-28 Frontier-scale coding, office, game-development and scientific agents where a large multi-GPU deployment is justified. 770B backbone + 10B MTP1M tokens · text BOfficial framework matrix Commercial use: permittedApache-2.0 View dossier
Qwen3.8-27BAlibaba Qwen · released 2026-08-14 A comparatively deployable native vision-language model for private coding, computer-use, document and general agent workloads. 27B256K tokens · text, image, video EMemory estimate not official Commercial use: permittedApache-2.0 View dossier
Qwen3.8-2.4T-A95BAlibaba Qwen · released 2026-08-14 Frontier-scale text reasoning, coding and long-horizon agent workloads where a multi-node cluster and custom commercial-license review are acceptable. 2.4T256K tokens · text CNot published Commercial use: conditionalQwen3.8-Max License View dossier
Kimi K3Moonshot AI · released 2026-07-27 Frontier-scale long-context multimodal and coding agent workloads where a cluster deployment is acceptable. 2.8T1M tokens · text, image, video CNot published Commercial use: conditionalKimi K3 License View dossier
GLM-5.3-FlashZ.ai · released 2026-08-26 Cost-sensitive multimodal coding, document, research and tool-using agents that need a one-million-token model specification and can support either heavy host-memory offload or a multi-GPU worker. 320B1M tokens · text, image, video BOfficial heterogeneous minimum Commercial use: permittedMIT View dossier
GLM-5.2Z.ai · released 2026-07-12 Long-horizon software engineering, tool use and million-token analysis on enterprise infrastructure. 753B1M tokens · text BHeterogeneous reference Commercial use: permittedMIT View dossier
DeepSeek V4 ProDeepSeek · released 2026-06-22 Maximum-quality reasoning, coding and long-context agents when an eight-accelerator worker is viable. 862B reported by model repository1M tokens · text BValidated reference not minimum Commercial use: permittedMIT View dossier
MiniMax M3MiniMax · released 2026-06-01 Multimodal agent, coding and long-context workloads with explicit commercial-license review. approximately 428B1M tokens · text, image, video DCandidate recipe under review Commercial use: conditionalMiniMax Community License View dossier
Qwen3.6-35B-A3BAlibaba Qwen · released 2026-04-16 A comparatively deployable private multimodal agent baseline with strong ecosystem coverage. 35B256K tokens · text, image, video COfficial launch shape Commercial use: permittedApache-2.0 View dossier
DeepSeek V4 Flash 0731DeepSeek · released 2026-07-31 The superseding DeepSeek V4 Flash checkpoint for high-frequency coding, reasoning and agent traffic, with an attached DSpark speculative draft module. 304B checkpoint metadata (284B target model plus attached DSpark draft module)1M tokens · text BOfficial framework floor Commercial use: permittedMIT View dossier
Step 3.7 FlashStepFun · released 2026-05-28 High-frequency multimodal agents, financial-document parsing, verified search loops and concurrent coding. 198B256K tokens · text, image COfficial local minimum Commercial use: permittedApache-2.0 View dossier
LongCat 2.0Meituan · released 2026-07-02 Long-horizon coding, search, repository edits and tool-driven agents on GPU or NPU clusters. 1.6T1M tokens · text CNot published Commercial use: permittedMIT View dossier
MiMo V2.5Xiaomi MiMo · released 2026-04-27 Native omnimodal understanding, long-context reasoning and agentic workflows across text, image, video and audio. 310B1M tokens · text, image, video, audio CNot published Commercial use: permittedMIT View dossier
Qwen3.5-9BAlibaba Qwen · released 2026-02-16 The local and edge-friendly baseline for private multimodal assistants, extraction and moderate-volume APIs. 9B256K tokens · text, image, video EMemory estimate not official Commercial use: permittedApache-2.0 View dossier
Qwen-AgentWorld-35B-A3BAlibaba Qwen · released 2026-06-24 A language world model for simulating MCP, Search, Terminal, SWE, Android, Web and OS agent environments. 35B256K tokens · text, environment state, action history COfficial tp4 launch Commercial use: permittedApache-2.0 View dossier
Xiaomi-Robotics-0Xiaomi Robotics · released 2026-02-11 Real-time robotic manipulation research and post-training across supported embodiments and simulation suites. 4.7BNot applicable · robot camera, language instruction, proprioception CConsumer gpu claim no vram Commercial use: permittedApache-2.0 View dossier
NAVABaidu ERNIE Team · released 2026-05-28 Synchronized audio-video generation with multi-speaker timbre control, camera prompting and image continuation. 6.3B backboneNot applicable · text, image, reference voice CSingle gpu supported no vram Commercial use: conditionalApache-2.0 View dossier
MiMo-V2.5-ASRXiaomi MiMo · released 2026-04-23 Chinese and English transcription across dialects, code-switching, meetings, songs, noise and knowledge-heavy speech. not publishedNot applicable · audio CNot published Commercial use: permittedApache-2.0 View dossier
Hy3Tencent Hunyuan · released 2026-07-06 Large private coding, productivity and tool-using agents that can support an official eight-accelerator tensor-parallel worker. 295B backbone plus 3.8B MTP layer256K tokens · text COfficial vendor floor Commercial use: permittedApache-2.0 View dossier
LongCat-Flash-Lite-SparseMeituan · released 2026-07-31 Sparse long-context coding, search and tool agents with an unusually small active path and an official single-H20 serving example. 69B1M tokens · text COfficial single accelerator path Commercial use: permittedMIT View dossier
Ling-3.0-flashAnt Group InclusionAI · released 2026-08-02 Hybrid-linear sparse model for coding, research and long-horizon agents with official four- and eight-GPU serving shapes. 124B256K tokens · text COfficial vendor floor Commercial use: permittedMIT View dossier
Ling-3.0-tinyAnt Group InclusionAI · released 2026-08-10 Compact sparse model for local coding and agent assistants on Apple silicon, DGX Spark and conventional GPU servers. 7.9B256K tokens · text COfficial local paths Commercial use: permittedMIT View dossier
MiniCPM5-2BOpenBMB · released 2026-09-08 On-device and resource-constrained assistant baseline with official server, desktop and Apple-silicon runtime options. 2.52B (1.98B non-embedding)128K tokens · text COfficial local paths no hardware floor Commercial use: permittedApache-2.0 View dossier
Spark X2.5XHToken · released 2026-09-01 Compact million-token multilingual model family for local agents, coding assistants and edge or desktop deployment. 4B primary; 1.7B official variant1M tokens · text COfficial single device path no vram floor Commercial use: permittedApache-2.0 View dossier
AREX-TurboBAAI · released 2026-07-23 Compact deep-research model for long-horizon information seeking, evidence aggregation and multi-constraint verification. 4B dense256K tokens · text COfficial paths no hardware floor Commercial use: permittedApache-2.0 View dossier
LLaDA2.2-flashAnt Group InclusionAI · released 2026-07-16 Diffusion-language-model research checkpoint for agent reasoning, tool use and iterative editing, with an executable Transformers path but no vendor-confirmed production server yet. 100B non-embedding; 103B repository metadata128K tokens · text CTransformers path no hardware floor Commercial use: permittedApache-2.0 View dossier
Unlimited-OCRBaidu · released 2026-06-22 Compact document-vision model for one-shot long-document OCR, multi-page PDF parsing and structured layout extraction. 3B32K tokens · image, PDF, text prompt COfficial single gpu path no vram floor Commercial use: permittedMIT View dossier

Workload contract

The same GPU count does not answer three different questions.

01

Official launch / minimum

One request, short context, no SLA; proves runnable, not production-ready.

02

200-person team

200 seats, 80-120 DAU, 20-30 peak online, 8-16 concurrent generations, 8K-32K ordinary prompts and occasional 128K agent work.

03

Commercial API

Public multi-tenant API, redundancy, rolling updates, rate limits and at least 99.9% availability target. Final GPU count requires measured traffic and latency targets.

Research watchlist

Interesting releases that do not yet clear the formal inclusion bar.

These projects have visible open-model value, but unclear commercial rights, partial artifacts or an unconfirmed directly deployable checkpoint keep them outside the verified count.

License unclear

ByteDance Lance

Official weights and inference code exist, but the repository does not clearly name a standard model license; commercial rights need clarification.

Official project

Partial release

GigaWorld-1

Official code and selected weights are available as a research preview, but the complete deployment surface is not yet equivalent to the formal ledger.

Official project

Weights not confirmed

Qwen-VLA

The official repository is public, but a complete, directly deployable released checkpoint was not sufficiently confirmed at this cutoff.

Official project

Methodology

How evidence earns a label.

Release window: 2025-09-12 through 2026-09-12. Included models have downloadable official weights and at least one documented executable self-hosting path. API-only announcements, unreleased weights, community-only ports and research checkpoints without a serving path are excluded.

Important: loading weights, completing the first forward pass and meeting a production latency/SLA are three separate thresholds. Context length also changes KV-cache demand, so seat count alone cannot produce a reliable GPU bill.

  1. AChinaAPI reproduced
  2. Binference-framework official validated recipe
  3. Cmodel-vendor documented configuration
  4. Dthird-party reproduction
  5. Ecapacity estimate only

Quick answers

Chinese open-model deployment FAQ.

Short answers for buyers, infrastructure teams and AI agents. Each model dossier is authoritative when this directory summary is not detailed enough.

Which Chinese open models released in the last year can be self-hosted?

As of 2026-09-12, this verified index includes DeepSeek V4.1 Flash, GLM-5.3, AuK / AuK-Flash, Qwen3.8-Flash-Next, Hy4 preview, Qwen3.8-27B, Qwen3.8-2.4T-A95B, Kimi K3, GLM-5.3-Flash, GLM-5.2, DeepSeek V4 Pro, MiniMax M3, Qwen3.6-35B-A3B, DeepSeek V4 Flash 0731, Step 3.7 Flash, LongCat 2.0, MiMo V2.5, Qwen3.5-9B, Qwen-AgentWorld-35B-A3B, Xiaomi-Robotics-0, NAVA, MiMo-V2.5-ASR, Hy3, LongCat-Flash-Lite-Sparse, Ling-3.0-flash, Ling-3.0-tiny, MiniCPM5-2B, Spark X2.5, AREX-Turbo, LLaDA2.2-flash, Unlimited-OCR. Each has downloadable official weights and at least one documented executable serving path.

What is the minimum hardware required to deploy these models?

There is no single trustworthy minimum. We report the smallest official or inference-framework launch shape found and label its evidence source. A launch shape is not a production capacity guarantee.

What hardware does a 200-person team need?

GPU count depends on DAU, peak concurrency, prompt and output length, context mix, latency targets and redundancy. Our comparison assumes 80–120 DAU, 20–30 peak online and 8–16 concurrent generations, then marks where load testing is still required.

Can these models be used commercially or offered as an API?

MIT and Apache-2.0 releases generally permit commercial use with notice obligations, while community licenses and bundled dependencies can add revenue, authorization, attribution or use restrictions. Review the complete license stack before launch.

Which model is best for my use case?

Start with the workload: general agents, environment simulation, physical-world robot control, joint audio-video generation and ASR are not interchangeable. The scenario guide labels vendor claims, public evidence and ChinaAPI deployment inferences separately.

Prefer an API?

Compare hosted model access before buying a cluster.

Use the same workload assumptions to compare self-hosting against an OpenAI-compatible managed endpoint.

View live API pricing