Released 2026-04-23 · speech recognition
Deployment evidence
Xiaomi MiMo · not published total · not published active · Not applicable
Best-fit scenarios
Chinese and English transcription across dialects, code-switching, meetings, songs, noise and knowledge-heavy speech.
Not recommended
- Capacity promises before real-time-factor and batch testing
- Assuming speaker overlap performance replaces diarization requirements
Chinese dialect and code-switch transcription
Vendor-statedOfficial support includes Wu, Cantonese, Hokkien, Sichuanese and Chinese-English code switching.
Capability sourceMeetings and noisy far-field audio
Vendor-statedThe release targets overlapping speakers, heavy noise and far-field capture.
Capability sourceKnowledge-heavy and lyric transcription
Public evidenceThe vendor reports evaluations across dialects, lyrics and complex English scenarios.
Capability sourceOfficial launch / minimum
Evidence CNot published
The official local Gradio and Python path requires CUDA 12+, but exact minimum VRAM is not published. Primary recipe
200-person team
Evidence EAudio workload profile required
Size by audio hours, peak simultaneous streams and latency rather than office seats.
Commercial API
Evidence ECustom service required
The release provides local inference code, not a validated multi-tenant serving recipe; build queueing, batching and observability.
Commercial-use check
Apache-2.0
MaaS / hosted service: No model-specific MaaS restriction identified in Apache-2.0.
Attribution: Provide the license and notices, preserve attribution notices, and mark modified files.
Known limitations and open questions
- Model parameter count and minimum VRAM are not published
- The public release does not provide an official high-throughput serving benchmark
