Open-weight deployment dossier · verified 2026-09-12

MiMo V2.5: self-hosting hardware, scenarios and commercial license

Native omnimodal understanding, long-context reasoning and agentic workflows across text, image, video and audio.

310B
total parameters
1M tokens
context window
MIT
published license

Released 2026-04-27 · general agent

Deployment evidence

Xiaomi MiMo · 310B total · 15B active · 1M tokens

PrecisionBF16 and FP8 componentsServing pathsTransformers, SGLang, vLLM

Best-fit scenarios

Native omnimodal understanding, long-context reasoning and agentic workflows across text, image, video and audio.

Not recommended
  • Small single-GPU deployment
  • Using stale config or tokenizer files from the initial release

Omnimodal research and support agents

Vendor-stated

The checkpoint natively accepts text, image, video and audio.

Capability source

Long video, audio and document analysis

Vendor-stated

The model supports up to 1M context and dedicated visual and audio encoders.

Capability source

Multimodal tool-using agents

Public evidence

The official card publishes multimodal, coding, agent and long-context evaluations.

Capability source

Official launch / minimum

Evidence C

Not published

Transformers can load the checkpoint, but the vendor does not state an absolute minimum GPU configuration.

200-person team

Evidence E

Benchmark required

Use the official distributed recipe as a starting worker and size replicas from the real modality mix.

Commercial API

Evidence C

Official distributed recipe

The official card shows an FP8 SGLang DP2×TP8 configuration at 262K context. Primary recipe

Commercial-use check

MIT

MaaS / hosted service: No model-specific MaaS restriction identified in the MIT license.

Attribution: Retain the copyright and permission notice.

Read the primary license text

Known limitations and open questions
  • Official deployment example is not a minimum
  • Audio and video traffic need separate encoder-capacity measurements

Evidence boundaries

A launch shape is not a production SLA.

The minimum tier records the smallest official or inference-framework configuration we found. The 200-person and commercial API tiers still require measurements against real prompt length, output length, concurrency, latency and redundancy targets.

Read the complete index methodology

  1. AChinaAPI reproduced
  2. Binference-framework official validated recipe
  3. Cmodel-vendor documented configuration
  4. Dthird-party reproduction
  5. Ecapacity estimate only

Compare before deploying

Review every verified model or compare hosted access.

The directory keeps model selection separate from the evidence and capacity details on this page.