Open-weight deployment dossier · verified 2026-09-12

DeepSeek V4 Pro: self-hosting hardware, scenarios and commercial license

Maximum-quality reasoning, coding and long-context agents when an eight-accelerator worker is viable.

862B reported by model repository
total parameters
1M tokens
context window
MIT
published license

Released 2026-06-22 · general agent

Deployment evidence

DeepSeek · 862B reported by model repository total · not independently verified in V1 active · 1M tokens

Precisionmixed-precision checkpoint, approximately 960 GB in vLLM recipeServing pathsvLLM, SGLang

Best-fit scenarios

Maximum-quality reasoning, coding and long-context agents when an eight-accelerator worker is viable.

Not recommended
  • Budget workstation inference
  • Public API launch without at least one redundant worker

Complex reasoning and coding

Public evidence

The official model card publishes reasoning, coding and agent benchmark results.

Capability source

Million-token retrieval and analysis

Vendor-stated

The release is designed around efficient million-token context intelligence.

Capability source

Premium private agent endpoint

ChinaAPI inference

Use when quality has more value than minimum infrastructure cost.

Official launch / minimum

Evidence B

Validated reference not minimum

vLLM publishes 8x B300, 8x H200, 8x MI355X and GB200 profiles; none is labelled the absolute minimum. Primary recipe

200-person team

Evidence E

Benchmark required

An 8-GPU worker may be a capacity building block, but replicas depend on concurrency and output-token demand.

Commercial API

Evidence B

Validated reference

8x H200 is documented with context capped at 800K to preserve KV headroom; production redundancy requires additional workers. Primary recipe

Commercial-use check

MIT

MaaS / hosted service: No model-specific MaaS restriction identified in the MIT license.

Attribution: Retain the copyright and permission notice.

Read the primary license text

Known limitations and open questions
  • Some SM120 workstation kernels have unresolved reports
  • Loading weights is not proof that the first forward pass succeeds

Evidence boundaries

A launch shape is not a production SLA.

The minimum tier records the smallest official or inference-framework configuration we found. The 200-person and commercial API tiers still require measurements against real prompt length, output length, concurrency, latency and redundancy targets.

Read the complete index methodology

  1. AChinaAPI reproduced
  2. Binference-framework official validated recipe
  3. Cmodel-vendor documented configuration
  4. Dthird-party reproduction
  5. Ecapacity estimate only

Compare before deploying

Review every verified model or compare hosted access.

The directory keeps model selection separate from the evidence and capacity details on this page.