Released 2026-07-27 · general agent
Deployment evidence
Moonshot AI · 2.8T total · 104B active · 1M tokens
Best-fit scenarios
Frontier-scale long-context multimodal and coding agent workloads where a cluster deployment is acceptable.
Not recommended
- Single-workstation deployments
- Low-cost high-QPS chat without aggressive batching or distillation
Long-context coding agents
Vendor-statedPositioned for repository-scale coding and long-horizon agent work.
Capability sourceMultimodal document and video analysis
Vendor-statedOfficial materials expose native multimodal inputs and a 1M-token context window.
Capability sourceHigh-value enterprise agent workflows
ChinaAPI inferenceBest considered when model capability justifies data-center-class full-context workers.
Official launch / minimum
Evidence CNot published
Vendor documents supported engines but does not publish a smallest runnable GPU configuration.
200-person team
Evidence ERequires cluster sizing
Not responsibly specifiable before a concurrency and context benchmark; workstation deployment is not credible.
Commercial API
Evidence BValidated reference
NVIDIA Dynamo publishes full-1M profiles using 8x GB300 or 16x GB200 per aggregated worker; disaggregated profiles use more GPUs. Primary recipe
Commercial-use check
Kimi K3 License
MaaS / hosted service: Separate agreement required when a MaaS operator and affiliates exceed USD 20M aggregate revenue over any consecutive 12 months.
Attribution: Display Kimi K3 prominently above 100M MAU or USD 20M monthly revenue.
Exceptions: The cited requirements do not apply to internal use or use through Moonshot official products or certified inference partners.
Known limitations and open questions
- No ChinaAPI hardware reproduction
- Full-context reference configurations are data-center cluster class
