Released 2026-06-22 · general agent
Deployment evidence
DeepSeek · 862B reported by model repository total · not independently verified in V1 active · 1M tokens
Best-fit scenarios
Maximum-quality reasoning, coding and long-context agents when an eight-accelerator worker is viable.
Not recommended
- Budget workstation inference
- Public API launch without at least one redundant worker
Complex reasoning and coding
Public evidenceThe official model card publishes reasoning, coding and agent benchmark results.
Capability sourceMillion-token retrieval and analysis
Vendor-statedThe release is designed around efficient million-token context intelligence.
Capability sourcePremium private agent endpoint
ChinaAPI inferenceUse when quality has more value than minimum infrastructure cost.
Official launch / minimum
Evidence BValidated reference not minimum
vLLM publishes 8x B300, 8x H200, 8x MI355X and GB200 profiles; none is labelled the absolute minimum. Primary recipe
200-person team
Evidence EBenchmark required
An 8-GPU worker may be a capacity building block, but replicas depend on concurrency and output-token demand.
Commercial API
Evidence BValidated reference
8x H200 is documented with context capped at 800K to preserve KV headroom; production redundancy requires additional workers. Primary recipe
Commercial-use check
MIT
MaaS / hosted service: No model-specific MaaS restriction identified in the MIT license.
Attribution: Retain the copyright and permission notice.
Known limitations and open questions
- Some SM120 workstation kernels have unresolved reports
- Loading weights is not proof that the first forward pass succeeds
