Open-weight deployment dossier · verified 2026-09-12

MiniCPM5-2B: self-hosting hardware, scenarios and commercial license

On-device and resource-constrained assistant baseline with official server, desktop and Apple-silicon runtime options.

2.52B (1.98B non-embedding)
total parameters
128K tokens
context window
Apache-2.0
published license

Released 2026-09-08 · general agent

Deployment evidence

OpenBMB · 2.52B (1.98B non-embedding) total · 2.52B dense active · 128K tokens

PrecisionBF16; official GGUF and MLX variantsServing pathsTransformers, vLLM, SGLang, llama.cpp, Ollama, LM Studio, MLX

Best-fit scenarios

On-device and resource-constrained assistant baseline with official server, desktop and Apple-silicon runtime options.

Not recommended
  • Assuming 131K context fits every official local format at useful speed
  • Selecting a quantization without task-specific accuracy checks

On-device assistant

Vendor-stated

OpenBMB positions MiniCPM5-2B for on-device and resource-constrained use and publishes GGUF and MLX variants.

Capability source

Local coding and tool use

Vendor-stated

The card documents chat, reasoning and tool-use paths across mainstream local runtimes.

Capability source

Cross-platform private endpoint

ChinaAPI inference

The broad runtime matrix makes the model a low-cost evaluation target before committing to larger checkpoints.

Official launch / minimum

Evidence C

Official local paths no hardware floor

Official llama.cpp, Ollama, LM Studio and MLX paths are published, but the vendor does not claim an absolute minimum RAM or VRAM configuration. Primary recipe

200-person team

Evidence E

Benchmark required

Benchmark a server runtime with the real output length and context mix; individual desktop compatibility is not a 200-seat capacity result.

Commercial API

Evidence C

Mainstream framework support

Official vLLM and SGLang paths exist; add replicas, rate limits and quantization acceptance tests for public use. Primary recipe

Commercial-use check

Apache-2.0

MaaS / hosted service: No model-specific MaaS restriction identified in Apache-2.0.

Attribution: Provide the license and required notices, preserve attribution notices, and mark modified files; trademark rights are not granted.

Read the primary license text

Known limitations and open questions
  • No ChinaAPI reproduction or official absolute hardware minimum
  • Long context materially increases KV-cache demand
  • Quantized variants can change quality and throughput

Evidence boundaries

A launch shape is not a production SLA.

The minimum tier records the smallest official or inference-framework configuration we found. The 200-person and commercial API tiers still require measurements against real prompt length, output length, concurrency, latency and redundancy targets.

Read the complete index methodology

  1. AChinaAPI reproduced
  2. Binference-framework official validated recipe
  3. Cmodel-vendor documented configuration
  4. Dthird-party reproduction
  5. Ecapacity estimate only

Compare before deploying

Review every verified model or compare hosted access.

The directory keeps model selection separate from the evidence and capacity details on this page.