Open model deployment index · verified 2026-08-03

What does it actually take to run China’s open frontier models?

A source-linked ledger of official weights, best-fit scenarios, launch configurations, 200-person team assumptions, public API architecture and commercial-use restrictions. We label vendor claims and estimates.

14
verified deployable models
5
model categories
3
deployment tiers
Abstract data-center hardware capacity blocks in ChinaAPI blue, violet and cyan
Capacity shape, not a benchmark result.

Choose by scenario

Start with the job, then check the hardware and license.

“Best fit” is not a universal quality ranking. Each recommendation is labelled as vendor-stated, backed by a named public evaluation, or inferred by ChinaAPI from the model’s modality and deployment shape.

ModelBest-fit useAvoid / validate firstPrimary claim type
Kimi K3General and agentic foundation models Frontier-scale long-context multimodal and coding agent workloads where a cluster deployment is acceptable. Single-workstation deployments; Low-cost high-QPS chat without aggressive batching or distillation Vendor-stated
GLM-5.2General and agentic foundation models Long-horizon software engineering, tool use and million-token analysis on enterprise infrastructure. Latency-sensitive single-GPU serving; Capacity planning based only on employee count Vendor-stated
DeepSeek V4 ProGeneral and agentic foundation models Maximum-quality reasoning, coding and long-context agents when an eight-accelerator worker is viable. Budget workstation inference; Public API launch without at least one redundant worker Public evidence
MiniMax M3General and agentic foundation models Multimodal agent, coding and long-context workloads with explicit commercial-license review. Commercial launch before license notice and attribution review; Stable-production serving without pinned framework versions Vendor-stated
Qwen3.6-35B-A3BGeneral and agentic foundation models A comparatively deployable private multimodal agent baseline with strong ecosystem coverage. Assuming TP4 equals four production replicas; Unvalidated parser upgrades in a critical tool loop ChinaAPI inference
DeepSeek V4 FlashGeneral and agentic foundation models A smaller DeepSeek V4 worker for high-frequency coding, reasoning and agent traffic. Treating Flash as quality-equivalent to Pro on every task; Launching a public API without measured parser and long-context behavior Public evidence
Step 3.7 FlashGeneral and agentic foundation models High-frequency multimodal agents, financial-document parsing, verified search loops and concurrent coding. Using the 128GB local path as a production throughput claim; Assuming vendor benchmark throughput transfers to long-context prefill Vendor-stated
LongCat 2.0General and agentic foundation models Long-horizon coding, search, repository edits and tool-driven agents on GPU or NPU clusters. Workstation deployment; Quoting in-house benchmark results as ChinaAPI reproduction Vendor-stated
MiMo V2.5General and agentic foundation models Native omnimodal understanding, long-context reasoning and agentic workflows across text, image, video and audio. Small single-GPU deployment; Using stale config or tokenizer files from the initial release Vendor-stated
Qwen3.5-9BGeneral and agentic foundation models The local and edge-friendly baseline for private multimodal assistants, extraction and moderate-volume APIs. Assuming 1M extended context fits a 24GB card; Highest-complexity long-horizon agents without task-specific evaluation ChinaAPI inference
Qwen-AgentWorld-35B-A3BDigital-world and environment models A language world model for simulating MCP, Search, Terminal, SWE, Android, Web and OS agent environments. General-purpose chat replacement; Treating simulated success as proof of real-environment reliability Vendor-stated
Xiaomi-Robotics-0Physical-world vision-language-action models Real-time robotic manipulation research and post-training across supported embodiments and simulation suites. Direct deployment on an unvalidated physical robot; Safety-critical control without independent interlocks and task-specific validation Vendor-stated
NAVAJoint audio-video generation Synchronized audio-video generation with multi-speaker timbre control, camera prompting and image continuation. Unconsented face or voice cloning; Low-latency interactive video generation Vendor-stated
MiMo-V2.5-ASRAutomatic speech recognition Chinese and English transcription across dialects, code-switching, meetings, songs, noise and knowledge-heavy speech. Capacity promises before real-time-factor and batch testing; Assuming speaker overlap performance replaces diarization requirements Vendor-stated
Vendor-statedCapability or use case explicitly described by the model vendor.Public evidenceCapability supported by a named public benchmark or published evaluation; not reproduced by ChinaAPI.ChinaAPI inferenceChinaAPI fit assessment inferred from modality, model size, context and documented serving path; not a quality benchmark.

Evidence ledger

One table for hardware, scale and license risk.

“Minimum” means the smallest official or framework-documented launch shape we found—not a production recommendation. A 200-seat estimate is workload-dependent, while a public API also needs redundancy and rollout capacity.

ModelCommercial termsOfficial launch / minimumCommercial API reference
Kimi K3Moonshot AI · 2.8T Commercial use: conditionalKimi K3 License C Vendor documents supported engines but does not publish a smallest runnable GPU configuration. B NVIDIA Dynamo publishes full-1M profiles using 8x GB300 or 16x GB200 per aggregated worker; disaggregated profiles use more GPUs.
GLM-5.2Z.ai · 753B Commercial use: permittedMIT B KTransformers documents an 8-GPU CPU/GPU-offload launch shape with 96 CPU inference threads; this is a tutorial target, not a claimed absolute minimum. E Requires redundant replicas or disaggregated prefill/decode; GPU count cannot be inferred from seats alone.
DeepSeek V4 ProDeepSeek · 862B reported by model repository Commercial use: permittedMIT B vLLM publishes 8x B300, 8x H200, 8x MI355X and GB200 profiles; none is labelled the absolute minimum. B 8x H200 is documented with context capped at 800K to preserve KV headroom; production redundancy requires additional workers.
MiniMax M3MiniMax · approximately 428B Commercial use: conditionalMiniMax Community License D A 4x RTX PRO 6000 NVFP4 recipe is under review upstream; V1 does not promote an unmerged recipe to validated minimum. B Aggregated and disaggregated vLLM recipes are still landing; pinning a nightly build may be required.
Qwen3.6-35B-A3BAlibaba Qwen · 35B Commercial use: permittedApache-2.0 C Official serving examples use TP4 and 262K context. Smaller quantized short-context shapes exist, but V1 does not call them the official minimum. E Scale through replicated workers after measuring prefill-heavy and decode-heavy traffic separately.
DeepSeek V4 FlashDeepSeek · 284B Commercial use: permittedMIT C Official local instructions and engine integrations are published, but no absolute minimum GPU count is claimed. B vLLM and SGLang support are available; production still requires redundant workers and pinned encoding logic.
Step 3.7 FlashStepFun · 198B Commercial use: permittedApache-2.0 C The official GGUF path specifies about 120GB minimum unified memory/VRAM and recommends 128GB. C Official examples publish TP4 NVFP4 and TP8 FP8/BF16 serving shapes; replicas are still required for HA.
LongCat 2.0Meituan · 1.6T Commercial use: permittedMIT C GPU and NPU paths are documented, but the vendor does not publish an absolute minimum hardware shape. C Official GPU and NPU serving paths exist; topology, redundancy and throughput remain operator-specific.
MiMo V2.5Xiaomi MiMo · 310B Commercial use: permittedMIT C Transformers can load the checkpoint, but the vendor does not state an absolute minimum GPU configuration. C The official card shows an FP8 SGLang DP2×TP8 configuration at 262K context.
Qwen3.5-9BAlibaba Qwen · 9B Commercial use: permittedApache-2.0 E A 24GB-class GPU is a reasonable BF16 short-context planning target, but the vendor does not label this an official minimum. C Official Transformers, vLLM and SGLang examples are published; add redundant replicas for availability.
Qwen-AgentWorld-35B-A3BAlibaba Qwen · 35B Commercial use: permittedApache-2.0 C Official SGLang and vLLM examples use tensor parallel size 4. E Expose behind a task-specific simulator contract; do not market it as a normal chat-completions quality substitute.
Xiaomi-Robotics-0Xiaomi Robotics · 4.7B Commercial use: permittedApache-2.0 C The vendor says BF16 inference is optimized for consumer GPUs but publishes no exact minimum VRAM. E Prefer an on-robot or near-edge safety architecture; a remote multi-tenant API is not the default production shape.
NAVABaidu ERNIE Team · 6.3B backbone Commercial use: conditionalApache-2.0 C The model card supports single-GPU inference but does not publish exact minimum VRAM. C The official Ulysses SP8 path reports roughly one minute for a 720p synchronized clip; HA needs additional workers.
MiMo-V2.5-ASRXiaomi MiMo · not published Commercial use: permittedApache-2.0 C The official local Gradio and Python path requires CUDA 12+, but exact minimum VRAM is not published. E The release provides local inference code, not a validated multi-tenant serving recipe; build queueing, batching and observability.

Workload contract

The same GPU count does not answer three different questions.

01

Official launch / minimum

One request, short context, no SLA; proves runnable, not production-ready.

02

200-person team

200 seats, 80-120 DAU, 20-30 peak online, 8-16 concurrent generations, 8K-32K ordinary prompts and occasional 128K agent work.

03

Commercial API

Public multi-tenant API, redundancy, rolling updates, rate limits and at least 99.9% availability target. Final GPU count requires measured traffic and latency targets.

Model dossiers

Primary sources, deployment evidence and commercial-use checks.

Every model entry keeps “documented,” “validated” and “estimated” separate. Commercial summaries are a first-pass product review, not legal advice.

Model category

General and agentic foundation models

10 verified deployable models

Released 2026-07-27 · general agent

Kimi K3

Moonshot AI · 2.8T total · 104B active · 1M tokens

PrecisionMXFP4 routed experts; BF16 dense components; MXFP8 activationsServing pathsvLLM, SGLang, TokenSpeed

Best-fit scenarios

Frontier-scale long-context multimodal and coding agent workloads where a cluster deployment is acceptable.

Not recommended
  • Single-workstation deployments
  • Low-cost high-QPS chat without aggressive batching or distillation

Long-context coding agents

Vendor-stated

Positioned for repository-scale coding and long-horizon agent work.

Capability source

Multimodal document and video analysis

Vendor-stated

Official materials expose native multimodal inputs and a 1M-token context window.

Capability source

High-value enterprise agent workflows

ChinaAPI inference

Best considered when model capability justifies data-center-class full-context workers.

Official launch / minimum

Evidence C

Not published

Vendor documents supported engines but does not publish a smallest runnable GPU configuration.

200-person team

Evidence E

Requires cluster sizing

Not responsibly specifiable before a concurrency and context benchmark; workstation deployment is not credible.

Commercial API

Evidence B

Validated reference

NVIDIA Dynamo publishes full-1M profiles using 8x GB300 or 16x GB200 per aggregated worker; disaggregated profiles use more GPUs. Primary recipe

Commercial-use check

Kimi K3 License

MaaS / hosted service: Separate agreement required when a MaaS operator and affiliates exceed USD 20M aggregate revenue over any consecutive 12 months.

Attribution: Display Kimi K3 prominently above 100M MAU or USD 20M monthly revenue.

Exceptions: The cited requirements do not apply to internal use or use through Moonshot official products or certified inference partners.

Read the primary license text

Known limitations and open questions
  • No ChinaAPI hardware reproduction
  • Full-context reference configurations are data-center cluster class

Released 2026-07-12 · general agent

GLM-5.2

Z.ai · 753B total · not independently verified in V1 active · 1M tokens

PrecisionBF16 and official FP8 checkpointServing pathsvLLM, SGLang, xLLM, KTransformers

Best-fit scenarios

Long-horizon software engineering, tool use and million-token analysis on enterprise infrastructure.

Not recommended
  • Latency-sensitive single-GPU serving
  • Capacity planning based only on employee count

Long-horizon software engineering

Vendor-stated

Official release materials emphasize coding-agent and tool-use workloads.

Capability source

Million-token repository and document analysis

Vendor-stated

The official checkpoint supports a 1M-token context window.

Capability source

Private enterprise agent service

ChinaAPI inference

A credible fit when CPU/GPU offload or TP8 infrastructure is already available.

Official launch / minimum

Evidence B

Heterogeneous reference

KTransformers documents an 8-GPU CPU/GPU-offload launch shape with 96 CPU inference threads; this is a tutorial target, not a claimed absolute minimum. Primary recipe

200-person team

Evidence E

Benchmark required

Start from a replicated TP8 service only after measuring the target context mix; no official 200-seat capacity claim exists.

Commercial API

Evidence E

Benchmark required

Requires redundant replicas or disaggregated prefill/decode; GPU count cannot be inferred from seats alone.

Commercial-use check

MIT

MaaS / hosted service: No model-specific MaaS restriction identified in the MIT license.

Attribution: Retain the copyright and permission notice.

Read the primary license text

Known limitations and open questions
  • The 8-GPU tutorial is not an official minimum claim
  • No public V1 concurrency result

Released 2026-06-22 · general agent

DeepSeek V4 Pro

DeepSeek · 862B reported by model repository total · not independently verified in V1 active · 1M tokens

Precisionmixed-precision checkpoint, approximately 960 GB in vLLM recipeServing pathsvLLM, SGLang

Best-fit scenarios

Maximum-quality reasoning, coding and long-context agents when an eight-accelerator worker is viable.

Not recommended
  • Budget workstation inference
  • Public API launch without at least one redundant worker

Complex reasoning and coding

Public evidence

The official model card publishes reasoning, coding and agent benchmark results.

Capability source

Million-token retrieval and analysis

Vendor-stated

The release is designed around efficient million-token context intelligence.

Capability source

Premium private agent endpoint

ChinaAPI inference

Use when quality has more value than minimum infrastructure cost.

Official launch / minimum

Evidence B

Validated reference not minimum

vLLM publishes 8x B300, 8x H200, 8x MI355X and GB200 profiles; none is labelled the absolute minimum. Primary recipe

200-person team

Evidence E

Benchmark required

An 8-GPU worker may be a capacity building block, but replicas depend on concurrency and output-token demand.

Commercial API

Evidence B

Validated reference

8x H200 is documented with context capped at 800K to preserve KV headroom; production redundancy requires additional workers. Primary recipe

Commercial-use check

MIT

MaaS / hosted service: No model-specific MaaS restriction identified in the MIT license.

Attribution: Retain the copyright and permission notice.

Read the primary license text

Known limitations and open questions
  • Some SM120 workstation kernels have unresolved reports
  • Loading weights is not proof that the first forward pass succeeds

Released 2026-06-01 · general agent

MiniMax M3

MiniMax · approximately 428B total · approximately 23B active · 1M tokens

PrecisionBF16 plus official MXFP8; NVIDIA NVFP4 conversion availableServing pathsTransformers, vLLM nightly, SGLang

Best-fit scenarios

Multimodal agent, coding and long-context workloads with explicit commercial-license review.

Not recommended
  • Commercial launch before license notice and attribution review
  • Stable-production serving without pinned framework versions

Multimodal agent workflows

Vendor-stated

Official materials position M3 for multimodal perception and agent tasks.

Capability source

Coding and tool orchestration

Public evidence

The official release reports coding and agent evaluations.

Capability source

Commercial embedded assistant

ChinaAPI inference

Potentially suitable after the revenue, notice and attribution clauses are cleared.

Official launch / minimum

Evidence D

Candidate recipe under review

A 4x RTX PRO 6000 NVFP4 recipe is under review upstream; V1 does not promote an unmerged recipe to validated minimum.

200-person team

Evidence E

Benchmark required

A quantized multi-GPU worker is plausible, but the 200-seat recommendation needs measured TTFT, throughput and context mix.

Commercial API

Evidence B

Framework support maturing

Aggregated and disaggregated vLLM recipes are still landing; pinning a nightly build may be required.

Commercial-use check

MiniMax Community License

MaaS / hosted service: Commercial API and hosted use are Commercial Use. Above USD 20M yearly revenue obtain prior written authorization; otherwise send the required one-time notice.

Attribution: Prominently display Built with MiniMax M3 for commercial use.

Prohibited uses: The license includes specified unlawful, military and harmful-use restrictions.

Read the primary license text

Known limitations and open questions
  • Stable vLLM release support was not complete at the V1 cutoff
  • Do not equate a pending recipe with successful ChinaAPI reproduction

Released 2026-04-16 · general agent

Qwen3.6-35B-A3B

Alibaba Qwen · 35B total · 3B active · 256K tokens

PrecisionBF16; official and community quantized formats availableServing pathsTransformers, vLLM, SGLang, llama.cpp, MLX

Best-fit scenarios

A comparatively deployable private multimodal agent baseline with strong ecosystem coverage.

Not recommended
  • Assuming TP4 equals four production replicas
  • Unvalidated parser upgrades in a critical tool loop

Private coding and office assistant

ChinaAPI inference

The 35B/3B-active shape and broad serving support make it a practical baseline for controlled workloads.

Visual document and UI understanding

Vendor-stated

Qwen3.6 is released as a native multimodal model family.

Capability source

Tool-calling agent service

Vendor-stated

Official serving examples include tool parsing and long-context operation.

Capability source

Official launch / minimum

Evidence C

Official launch shape

Official serving examples use TP4 and 262K context. Smaller quantized short-context shapes exist, but V1 does not call them the official minimum.

200-person team

Evidence E

Best v1 benchmark candidate

Use two measured serving replicas as the initial HA design candidate; exact GPUs remain pending load tests.

Commercial API

Evidence E

Benchmark required

Scale through replicated workers after measuring prefill-heavy and decode-heavy traffic separately.

Commercial-use check

Apache-2.0

MaaS / hosted service: No model-specific MaaS restriction identified in Apache-2.0.

Attribution: Provide the license and required notices, preserve attribution notices, and mark modified files; trademark rights are not granted.

Read the primary license text

Known limitations and open questions
  • Official TP4 example is a launch shape, not a capacity guarantee
  • Tool-call parser issues have been reported for related Qwen3.5 configurations

Released 2026-06-22 · general agent

DeepSeek V4 Flash

DeepSeek · 284B total · 13B active · 1M tokens

PrecisionFP4 experts with FP8 dense componentsServing pathsTransformers, vLLM, SGLang

Best-fit scenarios

A smaller DeepSeek V4 worker for high-frequency coding, reasoning and agent traffic.

Not recommended
  • Treating Flash as quality-equivalent to Pro on every task
  • Launching a public API without measured parser and long-context behavior

High-frequency coding agents

Public evidence

The official card reports coding and agent results and compares Flash reasoning modes.

Capability source

Million-token analysis

Vendor-stated

The model supports a 1M-token context with a 284B/13B-active MoE shape.

Capability source

Throughput-sensitive private endpoint

ChinaAPI inference

Prefer over V4 Pro when infrastructure cost and concurrency matter more than maximum quality.

Official launch / minimum

Evidence C

Vendor run path not minimum

Official local instructions and engine integrations are published, but no absolute minimum GPU count is claimed. Primary recipe

200-person team

Evidence E

Benchmark required

Use one measured tensor-parallel worker as the capacity baseline; add replicas only after concurrency tests.

Commercial API

Evidence B

Framework supported

vLLM and SGLang support are available; production still requires redundant workers and pinned encoding logic. Primary recipe

Commercial-use check

MIT

MaaS / hosted service: No model-specific MaaS restriction identified in the MIT license.

Attribution: Retain the copyright and permission notice.

Read the primary license text

Known limitations and open questions
  • No ChinaAPI hardware reproduction
  • The release uses a dedicated encoding implementation rather than a Jinja chat template

Released 2026-05-28 · general agent

Step 3.7 Flash

StepFun · 198B total · approximately 11B active · 256K tokens

PrecisionBF16, FP8, NVFP4 and GGUF releasesServing pathsvLLM, SGLang, Transformers, llama.cpp

Best-fit scenarios

High-frequency multimodal agents, financial-document parsing, verified search loops and concurrent coding.

Not recommended
  • Using the 128GB local path as a production throughput claim
  • Assuming vendor benchmark throughput transfers to long-context prefill

Visual documents and UI-to-code

Vendor-stated

Official materials highlight charts, GUIs, wireframes and structured-code extraction.

Capability source

Search and tool orchestration

Public evidence

The vendor publishes ClawEval, Toolathlon and tool-use evaluations.

Capability source

Concurrent coding agents

Vendor-stated

The model is explicitly engineered for high-frequency production agent workloads.

Capability source

Official launch / minimum

Evidence C

Official local minimum

The official GGUF path specifies about 120GB minimum unified memory/VRAM and recommends 128GB. Primary recipe

200-person team

Evidence E

Production candidate

Start with one TP4 NVFP4 or TP8 FP8 worker and benchmark the actual image and context mix.

Commercial API

Evidence C

Official launch shapes

Official examples publish TP4 NVFP4 and TP8 FP8/BF16 serving shapes; replicas are still required for HA. Primary recipe

Commercial-use check

Apache-2.0

MaaS / hosted service: No model-specific MaaS restriction identified in Apache-2.0.

Attribution: Provide the license and notices, preserve attribution notices, and mark modified files.

Read the primary license text

Known limitations and open questions
  • Published 400 tok/s is a vendor benchmark, not a universal SLA
  • NVFP4 support depends on recent engine and GPU paths

Released 2026-07-02 · general agent

LongCat 2.0

Meituan · 1.6T total · approximately 48B active · 1M tokens

PrecisionVendor checkpoint; exact serving precision varies by recipeServing pathsSGLang, SGLang-FluentLLM

Best-fit scenarios

Long-horizon coding, search, repository edits and tool-driven agents on GPU or NPU clusters.

Not recommended
  • Workstation deployment
  • Quoting in-house benchmark results as ChinaAPI reproduction

Repository-scale coding

Vendor-stated

Official materials emphasize repository edits and integrations with coding-agent harnesses.

Capability source

Search and general agents

Public evidence

The vendor publishes coding, BrowseComp, RWSearch and agent evaluations.

Capability source

NPU-based sovereign deployment

Vendor-stated

An official SGLang-FluentLLM NPU serving path is linked.

Capability source

Official launch / minimum

Evidence C

Not published

GPU and NPU paths are documented, but the vendor does not publish an absolute minimum hardware shape.

200-person team

Evidence E

Cluster sizing required

A 1.6T model needs a measured cluster design; employee count alone is not a capacity input.

Commercial API

Evidence C

Official serving paths

Official GPU and NPU serving paths exist; topology, redundancy and throughput remain operator-specific. Primary recipe

Commercial-use check

MIT

MaaS / hosted service: No model-specific MaaS restriction identified in the MIT license.

Attribution: Retain the copyright and permission notice.

Read the primary license text

Known limitations and open questions
  • No official minimum GPU count
  • Most published evaluation values are vendor-measured

Released 2026-04-27 · general agent

MiMo V2.5

Xiaomi MiMo · 310B total · 15B active · 1M tokens

PrecisionBF16 and FP8 componentsServing pathsTransformers, SGLang, vLLM

Best-fit scenarios

Native omnimodal understanding, long-context reasoning and agentic workflows across text, image, video and audio.

Not recommended
  • Small single-GPU deployment
  • Using stale config or tokenizer files from the initial release

Omnimodal research and support agents

Vendor-stated

The checkpoint natively accepts text, image, video and audio.

Capability source

Long video, audio and document analysis

Vendor-stated

The model supports up to 1M context and dedicated visual and audio encoders.

Capability source

Multimodal tool-using agents

Public evidence

The official card publishes multimodal, coding, agent and long-context evaluations.

Capability source

Official launch / minimum

Evidence C

Not published

Transformers can load the checkpoint, but the vendor does not state an absolute minimum GPU configuration.

200-person team

Evidence E

Benchmark required

Use the official distributed recipe as a starting worker and size replicas from the real modality mix.

Commercial API

Evidence C

Official distributed recipe

The official card shows an FP8 SGLang DP2×TP8 configuration at 262K context. Primary recipe

Commercial-use check

MIT

MaaS / hosted service: No model-specific MaaS restriction identified in the MIT license.

Attribution: Retain the copyright and permission notice.

Read the primary license text

Known limitations and open questions
  • Official deployment example is not a minimum
  • Audio and video traffic need separate encoder-capacity measurements

Released 2026-02-16 · general agent

Qwen3.5-9B

Alibaba Qwen · 9B total · 9B active · 256K tokens

PrecisionBF16; multiple quantized formats availableServing pathsTransformers, vLLM, SGLang, KTransformers, llama.cpp

Best-fit scenarios

The local and edge-friendly baseline for private multimodal assistants, extraction and moderate-volume APIs.

Not recommended
  • Assuming 1M extended context fits a 24GB card
  • Highest-complexity long-horizon agents without task-specific evaluation

Local multimodal assistant

ChinaAPI inference

The 9B dense shape is the most accessible formal model in this ledger.

Visual extraction and classification

Vendor-stated

The official model is a unified vision-language foundation model.

Capability source

Cost-sensitive private API

ChinaAPI inference

A better first benchmark target than frontier-scale MoE models when concurrency and budget dominate.

Official launch / minimum

Evidence E

Memory estimate not official

A 24GB-class GPU is a reasonable BF16 short-context planning target, but the vendor does not label this an official minimum.

200-person team

Evidence E

Best low cost candidate

Benchmark one or two 24–48GB workers before considering larger models; exact replicas depend on output length.

Commercial API

Evidence C

Mainstream framework support

Official Transformers, vLLM and SGLang examples are published; add redundant replicas for availability. Primary recipe

Commercial-use check

Apache-2.0

MaaS / hosted service: No model-specific MaaS restriction identified in Apache-2.0.

Attribution: Provide the license and notices, preserve attribution notices, and mark modified files.

Read the primary license text

Known limitations and open questions
  • 24GB guidance is an estimate, not ChinaAPI reproduction
  • Native 262K and extended 1M contexts materially increase KV-cache demand

Model category

Digital-world and environment models

1 verified deployable model

Released 2026-06-24 · world model

Qwen-AgentWorld-35B-A3B

Alibaba Qwen · 35B total · 3B active · 256K tokens

PrecisionBF16Serving pathsSGLang, vLLM

Best-fit scenarios

A language world model for simulating MCP, Search, Terminal, SWE, Android, Web and OS agent environments.

Not recommended
  • General-purpose chat replacement
  • Treating simulated success as proof of real-environment reliability

Agent environment simulation

Vendor-stated

The model predicts environment transitions across seven unified domains.

Capability source

Synthetic trajectories and perturbation tests

Vendor-stated

Official materials highlight controllable simulation and fictional-world construction.

Capability source

Agent regression evaluation

ChinaAPI inference

Useful as a simulator component, not as a drop-in customer chatbot.

Official launch / minimum

Evidence C

Official tp4 launch

Official SGLang and vLLM examples use tensor parallel size 4. Primary recipe

200-person team

Evidence E

Not a seat based service

Size by simulation jobs and trajectory length, not employee seats.

Commercial API

Evidence E

Specialized service

Expose behind a task-specific simulator contract; do not market it as a normal chat-completions quality substitute.

Commercial-use check

Apache-2.0

MaaS / hosted service: No model-specific MaaS restriction identified in Apache-2.0.

Attribution: Provide the license and notices, preserve attribution notices, and mark modified files.

Read the primary license text

Known limitations and open questions
  • World-model outputs are simulations
  • The 35B release covers seven named domains, not arbitrary physical environments

Model category

Physical-world vision-language-action models

1 verified deployable model

Released 2026-02-11 · robotics vla

Xiaomi-Robotics-0

Xiaomi Robotics · 4.7B total · 4.7B active · Not applicable

PrecisionBF16 with Flash Attention 2Serving pathsTransformers, PyTorch

Best-fit scenarios

Real-time robotic manipulation research and post-training across supported embodiments and simulation suites.

Not recommended
  • Direct deployment on an unvalidated physical robot
  • Safety-critical control without independent interlocks and task-specific validation

Robot manipulation

Vendor-stated

The official deployment guide targets robotic manipulation with asynchronous real-time execution.

Capability source

LIBERO, CALVIN and SimplerEnv evaluation

Public evidence

Official fine-tuned checkpoints and evaluation guides are released for all three suites.

Capability source

Embodiment-specific post-training

Vendor-stated

Post-training code is available for adapting the model to new data.

Capability source

Official launch / minimum

Evidence C

Consumer gpu claim no vram

The vendor says BF16 inference is optimized for consumer GPUs but publishes no exact minimum VRAM. Primary recipe

200-person team

Evidence E

Robot fleet profile required

Size by robots, camera rate and control latency; a 200-seat office profile is not applicable.

Commercial API

Evidence E

Edge control not public api

Prefer an on-robot or near-edge safety architecture; a remote multi-tenant API is not the default production shape.

Commercial-use check

Apache-2.0

MaaS / hosted service: No model-specific MaaS restriction identified in Apache-2.0.

Attribution: Provide the license and notices, preserve attribution notices, and mark modified files.

Read the primary license text

Known limitations and open questions
  • Simulation benchmarks do not prove physical-world safety
  • Exact VRAM and end-to-end control latency are not published

Model category

Joint audio-video generation

1 verified deployable model

Released 2026-05-28 · audio video generation

NAVA

Baidu ERNIE Team · 6.3B backbone total · 6.3B active · Not applicable

PrecisionBF16Serving pathsPyTorch, Ulysses sequence parallel

Best-fit scenarios

Synchronized audio-video generation with multi-speaker timbre control, camera prompting and image continuation.

Not recommended
  • Unconsented face or voice cloning
  • Low-latency interactive video generation

Synchronized short-form audio-video

Vendor-stated

NAVA jointly generates video, scene audio and speech rather than aligning separate outputs after generation.

Capability source

Multi-speaker and reference-timbre scenes

Vendor-stated

The official checkpoint supports up to two reference voices bound to speech spans.

Capability source

720p creative generation

Public evidence

The vendor reports VerseBench synchronization and quality results plus an 8-GPU fast path.

Capability source

Official launch / minimum

Evidence C

Single gpu supported no vram

The model card supports single-GPU inference but does not publish exact minimum VRAM. Primary recipe

200-person team

Evidence E

Media queue required

Size a queued render farm by jobs per hour, resolution and duration; office-seat assumptions do not apply.

Commercial API

Evidence C

Official 8gpu reference

The official Ulysses SP8 path reports roughly one minute for a 720p synchronized clip; HA needs additional workers. Primary recipe

Commercial-use check

Apache-2.0

MaaS / hosted service: The model card states Apache-2.0, but bundled LTX audio-VAE artifacts carry an additional community license that must be reviewed for the shipped stack.

Attribution: Preserve Apache notices and the notices/licenses for bundled upstream components.

Prohibited uses: The model card prohibits depicting real persons without consent, including face or voice likeness reproduction.

Read the primary license text

Known limitations and open questions
  • Default clips are about 6–10 seconds
  • The full dependency stack includes component-specific notices beyond the headline Apache license

Model category

Automatic speech recognition

1 verified deployable model

Released 2026-04-23 · speech recognition

MiMo-V2.5-ASR

Xiaomi MiMo · not published total · not published active · Not applicable

PrecisionCUDA 12 path with Flash AttentionServing pathsPyTorch, Gradio

Best-fit scenarios

Chinese and English transcription across dialects, code-switching, meetings, songs, noise and knowledge-heavy speech.

Not recommended
  • Capacity promises before real-time-factor and batch testing
  • Assuming speaker overlap performance replaces diarization requirements

Chinese dialect and code-switch transcription

Vendor-stated

Official support includes Wu, Cantonese, Hokkien, Sichuanese and Chinese-English code switching.

Capability source

Meetings and noisy far-field audio

Vendor-stated

The release targets overlapping speakers, heavy noise and far-field capture.

Capability source

Knowledge-heavy and lyric transcription

Public evidence

The vendor reports evaluations across dialects, lyrics and complex English scenarios.

Capability source

Official launch / minimum

Evidence C

Not published

The official local Gradio and Python path requires CUDA 12+, but exact minimum VRAM is not published. Primary recipe

200-person team

Evidence E

Audio workload profile required

Size by audio hours, peak simultaneous streams and latency rather than office seats.

Commercial API

Evidence E

Custom service required

The release provides local inference code, not a validated multi-tenant serving recipe; build queueing, batching and observability.

Commercial-use check

Apache-2.0

MaaS / hosted service: No model-specific MaaS restriction identified in Apache-2.0.

Attribution: Provide the license and notices, preserve attribution notices, and mark modified files.

Read the primary license text

Known limitations and open questions
  • Model parameter count and minimum VRAM are not published
  • The public release does not provide an official high-throughput serving benchmark

Research watchlist

Interesting releases that do not yet clear the formal inclusion bar.

These projects have visible open-model value, but unclear commercial rights, partial artifacts or an unconfirmed directly deployable checkpoint keep them outside the verified count.

License unclear

ByteDance Lance

Official weights and inference code exist, but the repository does not clearly name a standard model license; commercial rights need clarification.

Official project

Partial release

GigaWorld-1

Official code and selected weights are available as a research preview, but the complete deployment surface is not yet equivalent to the formal ledger.

Official project

Weights not confirmed

Qwen-VLA

The official repository is public, but a complete, directly deployable released checkpoint was not sufficiently confirmed at this cutoff.

Official project

Methodology

How evidence earns a label.

Release window: 2025-08-03 through 2026-08-03. Included models have downloadable official weights and at least one documented executable self-hosting path. API-only announcements, unreleased weights, community-only ports and research checkpoints without a serving path are excluded.

Important: loading weights, completing the first forward pass and meeting a production latency/SLA are three separate thresholds. Context length also changes KV-cache demand, so seat count alone cannot produce a reliable GPU bill.

  1. AChinaAPI reproduced
  2. Binference-framework official validated recipe
  3. Cmodel-vendor documented configuration
  4. Dthird-party reproduction
  5. Ecapacity estimate only

Quick answers

Chinese open-model deployment FAQ.

Short answers for buyers, infrastructure teams and AI agents. The evidence ledger above is authoritative when a summary and a model-specific row differ.

Which Chinese open models released in the last year can be self-hosted?

As of 2026-08-03, this verified index includes Kimi K3, GLM-5.2, DeepSeek V4 Pro, MiniMax M3, Qwen3.6-35B-A3B, DeepSeek V4 Flash, Step 3.7 Flash, LongCat 2.0, MiMo V2.5, Qwen3.5-9B, Qwen-AgentWorld-35B-A3B, Xiaomi-Robotics-0, NAVA, MiMo-V2.5-ASR. Each has downloadable official weights and at least one documented executable serving path.

What is the minimum hardware required to deploy these models?

There is no single trustworthy minimum. We report the smallest official or inference-framework launch shape found and label its evidence source. A launch shape is not a production capacity guarantee.

What hardware does a 200-person team need?

GPU count depends on DAU, peak concurrency, prompt and output length, context mix, latency targets and redundancy. Our comparison assumes 80–120 DAU, 20–30 peak online and 8–16 concurrent generations, then marks where load testing is still required.

Can these models be used commercially or offered as an API?

MIT and Apache-2.0 releases generally permit commercial use with notice obligations, while community licenses and bundled dependencies can add revenue, authorization, attribution or use restrictions. Review the complete license stack before launch.

Which model is best for my use case?

Start with the workload: general agents, environment simulation, physical-world robot control, joint audio-video generation and ASR are not interchangeable. The scenario guide labels vendor claims, public evidence and ChinaAPI deployment inferences separately.

Prefer an API?

Compare hosted model access before buying a cluster.

Use the same workload assumptions to compare self-hosting against an OpenAI-compatible managed endpoint.

View live API pricing