Open-weight deployment dossier · verified 2026-09-12

MiMo-V2.5-ASR: self-hosting hardware, scenarios and commercial license

Chinese and English transcription across dialects, code-switching, meetings, songs, noise and knowledge-heavy speech.

not published
total parameters
Not applicable
context window
Apache-2.0
published license

Released 2026-04-23 · speech recognition

Deployment evidence

Xiaomi MiMo · not published total · not published active · Not applicable

PrecisionCUDA 12 path with Flash AttentionServing pathsPyTorch, Gradio

Best-fit scenarios

Chinese and English transcription across dialects, code-switching, meetings, songs, noise and knowledge-heavy speech.

Not recommended
  • Capacity promises before real-time-factor and batch testing
  • Assuming speaker overlap performance replaces diarization requirements

Chinese dialect and code-switch transcription

Vendor-stated

Official support includes Wu, Cantonese, Hokkien, Sichuanese and Chinese-English code switching.

Capability source

Meetings and noisy far-field audio

Vendor-stated

The release targets overlapping speakers, heavy noise and far-field capture.

Capability source

Knowledge-heavy and lyric transcription

Public evidence

The vendor reports evaluations across dialects, lyrics and complex English scenarios.

Capability source

Official launch / minimum

Evidence C

Not published

The official local Gradio and Python path requires CUDA 12+, but exact minimum VRAM is not published. Primary recipe

200-person team

Evidence E

Audio workload profile required

Size by audio hours, peak simultaneous streams and latency rather than office seats.

Commercial API

Evidence E

Custom service required

The release provides local inference code, not a validated multi-tenant serving recipe; build queueing, batching and observability.

Commercial-use check

Apache-2.0

MaaS / hosted service: No model-specific MaaS restriction identified in Apache-2.0.

Attribution: Provide the license and notices, preserve attribution notices, and mark modified files.

Read the primary license text

Known limitations and open questions
  • Model parameter count and minimum VRAM are not published
  • The public release does not provide an official high-throughput serving benchmark

Evidence boundaries

A launch shape is not a production SLA.

The minimum tier records the smallest official or inference-framework configuration we found. The 200-person and commercial API tiers still require measurements against real prompt length, output length, concurrency, latency and redundancy targets.

Read the complete index methodology

  1. AChinaAPI reproduced
  2. Binference-framework official validated recipe
  3. Cmodel-vendor documented configuration
  4. Dthird-party reproduction
  5. Ecapacity estimate only

Compare before deploying

Review every verified model or compare hosted access.

The directory keeps model selection separate from the evidence and capacity details on this page.