Open-weight deployment dossier · verified 2026-09-12

Kimi K3: self-hosting hardware, scenarios and commercial license

Frontier-scale long-context multimodal and coding agent workloads where a cluster deployment is acceptable.

2.8T
total parameters
1M tokens
context window
Kimi K3 License
published license

Released 2026-07-27 · general agent

Deployment evidence

Moonshot AI · 2.8T total · 104B active · 1M tokens

PrecisionMXFP4 routed experts; BF16 dense components; MXFP8 activationsServing pathsvLLM, SGLang, TokenSpeed

Best-fit scenarios

Frontier-scale long-context multimodal and coding agent workloads where a cluster deployment is acceptable.

Not recommended
  • Single-workstation deployments
  • Low-cost high-QPS chat without aggressive batching or distillation

Long-context coding agents

Vendor-stated

Positioned for repository-scale coding and long-horizon agent work.

Capability source

Multimodal document and video analysis

Vendor-stated

Official materials expose native multimodal inputs and a 1M-token context window.

Capability source

High-value enterprise agent workflows

ChinaAPI inference

Best considered when model capability justifies data-center-class full-context workers.

Official launch / minimum

Evidence C

Not published

Vendor documents supported engines but does not publish a smallest runnable GPU configuration.

200-person team

Evidence E

Requires cluster sizing

Not responsibly specifiable before a concurrency and context benchmark; workstation deployment is not credible.

Commercial API

Evidence B

Validated reference

NVIDIA Dynamo publishes full-1M profiles using 8x GB300 or 16x GB200 per aggregated worker; disaggregated profiles use more GPUs. Primary recipe

Commercial-use check

Kimi K3 License

MaaS / hosted service: Separate agreement required when a MaaS operator and affiliates exceed USD 20M aggregate revenue over any consecutive 12 months.

Attribution: Display Kimi K3 prominently above 100M MAU or USD 20M monthly revenue.

Exceptions: The cited requirements do not apply to internal use or use through Moonshot official products or certified inference partners.

Read the primary license text

Known limitations and open questions
  • No ChinaAPI hardware reproduction
  • Full-context reference configurations are data-center cluster class

Evidence boundaries

A launch shape is not a production SLA.

The minimum tier records the smallest official or inference-framework configuration we found. The 200-person and commercial API tiers still require measurements against real prompt length, output length, concurrency, latency and redundancy targets.

Read the complete index methodology

  1. AChinaAPI reproduced
  2. Binference-framework official validated recipe
  3. Cmodel-vendor documented configuration
  4. Dthird-party reproduction
  5. Ecapacity estimate only

Compare before deploying

Review every verified model or compare hosted access.

The directory keeps model selection separate from the evidence and capacity details on this page.