# MiMo-V2.5-ASR: Self-Hosting Hardware, Scenarios and Commercial License

> Chinese and English transcription across dialects, code-switching, meetings, songs, noise and knowledge-heavy speech.

- Verified: 2026-09-12
- Released: 2026-04-23
- Canonical: https://chinaapi.ai/open-model-deployment/mimo-v2.5-asr/
- Back to directory: https://chinaapi.ai/open-model-deployment/
- Machine-readable dataset: https://chinaapi.ai/data/open-model-deployment.json

## Model facts

- Vendor: Xiaomi MiMo
- Parameters: not published total; not published active
- Native precision: CUDA 12 path with Flash Attention
- Context: Not applicable
- Category: Automatic speech recognition
- Modalities: audio → punctuated text
- Serving paths: PyTorch, Gradio
- Official weights: https://huggingface.co/XiaomiMiMo/MiMo-V2.5-ASR?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- Model documentation: https://github.com/XiaomiMiMo/MiMo-V2.5-ASR?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment

## Deployment evidence

- **Official launch / minimum — evidence C:** The official local Gradio and Python path requires CUDA 12+, but exact minimum VRAM is not published. Source: https://github.com/XiaomiMiMo/MiMo-V2.5-ASR?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **200-person team — evidence E:** Size by audio hours, peak simultaneous streams and latency rather than office seats.
- **Commercial API — evidence E:** The release provides local inference code, not a validated multi-tenant serving recipe; build queueing, batching and observability.

## Scenario evidence

- **Chinese dialect and code-switch transcription — Vendor-stated:** Official support includes Wu, Cantonese, Hokkien, Sichuanese and Chinese-English code switching. Source: https://github.com/XiaomiMiMo/MiMo-V2.5-ASR?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Meetings and noisy far-field audio — Vendor-stated:** The release targets overlapping speakers, heavy noise and far-field capture. Source: https://github.com/XiaomiMiMo/MiMo-V2.5-ASR?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Knowledge-heavy and lyric transcription — Public evidence:** The vendor reports evaluations across dialects, lyrics and complex English scenarios. Source: https://github.com/XiaomiMiMo/MiMo-V2.5-ASR?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment

## Commercial-use check

- License: Apache-2.0 — https://github.com/XiaomiMiMo/MiMo-V2.5-ASR/blob/main/LICENSE?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- MaaS / hosted service: No model-specific MaaS restriction identified in Apache-2.0.
- Attribution: Provide the license and notices, preserve attribution notices, and mark modified files.

## Avoid or validate first

- Capacity promises before real-time-factor and batch testing
- Assuming speaker overlap performance replaces diarization requirements

## Known limitations and open questions

- Model parameter count and minimum VRAM are not published
- The public release does not provide an official high-throughput serving benchmark

## Evidence boundary

Loading weights, completing a first forward pass and meeting a production latency/SLA are separate thresholds. The 200-person and commercial tiers require workload-specific measurement. License summaries are product research, not legal advice.

Generated by GENERATED BY scripts/gen_open_model_deployment.py.
