Released 2026-09-08 · general agent
Deployment evidence
OpenBMB · 2.52B (1.98B non-embedding) total · 2.52B dense active · 128K tokens
Best-fit scenarios
On-device and resource-constrained assistant baseline with official server, desktop and Apple-silicon runtime options.
Not recommended
- Assuming 131K context fits every official local format at useful speed
- Selecting a quantization without task-specific accuracy checks
On-device assistant
Vendor-statedOpenBMB positions MiniCPM5-2B for on-device and resource-constrained use and publishes GGUF and MLX variants.
Capability sourceLocal coding and tool use
Vendor-statedThe card documents chat, reasoning and tool-use paths across mainstream local runtimes.
Capability sourceCross-platform private endpoint
ChinaAPI inferenceThe broad runtime matrix makes the model a low-cost evaluation target before committing to larger checkpoints.
Official launch / minimum
Evidence COfficial local paths no hardware floor
Official llama.cpp, Ollama, LM Studio and MLX paths are published, but the vendor does not claim an absolute minimum RAM or VRAM configuration. Primary recipe
200-person team
Evidence EBenchmark required
Benchmark a server runtime with the real output length and context mix; individual desktop compatibility is not a 200-seat capacity result.
Commercial API
Evidence CMainstream framework support
Official vLLM and SGLang paths exist; add replicas, rate limits and quantization acceptance tests for public use. Primary recipe
Commercial-use check
Apache-2.0
MaaS / hosted service: No model-specific MaaS restriction identified in Apache-2.0.
Attribution: Provide the license and required notices, preserve attribution notices, and mark modified files; trademark rights are not granted.
Known limitations and open questions
- No ChinaAPI reproduction or official absolute hardware minimum
- Long context materially increases KV-cache demand
- Quantized variants can change quality and throughput
