Open-weight deployment dossier · verified 2026-09-12

MiniMax M3: self-hosting hardware, scenarios and commercial license

Multimodal agent, coding and long-context workloads with explicit commercial-license review.

approximately 428B
total parameters
1M tokens
context window
MiniMax Community License
published license

Released 2026-06-01 · general agent

Deployment evidence

MiniMax · approximately 428B total · approximately 23B active · 1M tokens

PrecisionBF16 plus official MXFP8; NVIDIA NVFP4 conversion availableServing pathsTransformers, vLLM nightly, SGLang

Best-fit scenarios

Multimodal agent, coding and long-context workloads with explicit commercial-license review.

Not recommended
  • Commercial launch before license notice and attribution review
  • Stable-production serving without pinned framework versions

Multimodal agent workflows

Vendor-stated

Official materials position M3 for multimodal perception and agent tasks.

Capability source

Coding and tool orchestration

Public evidence

The official release reports coding and agent evaluations.

Capability source

Commercial embedded assistant

ChinaAPI inference

Potentially suitable after the revenue, notice and attribution clauses are cleared.

Official launch / minimum

Evidence D

Candidate recipe under review

A 4x RTX PRO 6000 NVFP4 recipe is under review upstream; V1 does not promote an unmerged recipe to validated minimum.

200-person team

Evidence E

Benchmark required

A quantized multi-GPU worker is plausible, but the 200-seat recommendation needs measured TTFT, throughput and context mix.

Commercial API

Evidence B

Framework support maturing

Aggregated and disaggregated vLLM recipes are still landing; pinning a nightly build may be required.

Commercial-use check

MiniMax Community License

MaaS / hosted service: Commercial API and hosted use are Commercial Use. Above USD 20M yearly revenue obtain prior written authorization; otherwise send the required one-time notice.

Attribution: Prominently display Built with MiniMax M3 for commercial use.

Prohibited uses: The license includes specified unlawful, military and harmful-use restrictions.

Read the primary license text

Known limitations and open questions
  • Stable vLLM release support was not complete at the V1 cutoff
  • Do not equate a pending recipe with successful ChinaAPI reproduction

Evidence boundaries

A launch shape is not a production SLA.

The minimum tier records the smallest official or inference-framework configuration we found. The 200-person and commercial API tiers still require measurements against real prompt length, output length, concurrency, latency and redundancy targets.

Read the complete index methodology

  1. AChinaAPI reproduced
  2. Binference-framework official validated recipe
  3. Cmodel-vendor documented configuration
  4. Dthird-party reproduction
  5. Ecapacity estimate only

Compare before deploying

Review every verified model or compare hosted access.

The directory keeps model selection separate from the evidence and capacity details on this page.