Released 2026-04-27 · general agent
Deployment evidence
Xiaomi MiMo · 310B total · 15B active · 1M tokens
Best-fit scenarios
Native omnimodal understanding, long-context reasoning and agentic workflows across text, image, video and audio.
Not recommended
- Small single-GPU deployment
- Using stale config or tokenizer files from the initial release
Omnimodal research and support agents
Vendor-statedThe checkpoint natively accepts text, image, video and audio.
Capability sourceLong video, audio and document analysis
Vendor-statedThe model supports up to 1M context and dedicated visual and audio encoders.
Capability sourceMultimodal tool-using agents
Public evidenceThe official card publishes multimodal, coding, agent and long-context evaluations.
Capability sourceOfficial launch / minimum
Evidence CNot published
Transformers can load the checkpoint, but the vendor does not state an absolute minimum GPU configuration.
200-person team
Evidence EBenchmark required
Use the official distributed recipe as a starting worker and size replicas from the real modality mix.
Commercial API
Evidence COfficial distributed recipe
The official card shows an FP8 SGLang DP2×TP8 configuration at 262K context. Primary recipe
Commercial-use check
MIT
MaaS / hosted service: No model-specific MaaS restriction identified in the MIT license.
Attribution: Retain the copyright and permission notice.
Known limitations and open questions
- Official deployment example is not a minimum
- Audio and video traffic need separate encoder-capacity measurements
