# DeepSeek V4 Pro: Self-Hosting Hardware, Scenarios and Commercial License

> Maximum-quality reasoning, coding and long-context agents when an eight-accelerator worker is viable.

- Verified: 2026-09-12
- Released: 2026-06-22
- Canonical: https://chinaapi.ai/open-model-deployment/deepseek-v4-pro/
- Back to directory: https://chinaapi.ai/open-model-deployment/
- Machine-readable dataset: https://chinaapi.ai/data/open-model-deployment.json

## Model facts

- Vendor: DeepSeek
- Parameters: 862B reported by model repository total; not independently verified in V1 active
- Native precision: mixed-precision checkpoint, approximately 960 GB in vLLM recipe
- Context: 1M tokens
- Category: General and agentic foundation models
- Modalities: text → text
- Serving paths: vLLM, SGLang
- Official weights: https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- Model documentation: https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment

## Deployment evidence

- **Official launch / minimum — evidence B:** vLLM publishes 8x B300, 8x H200, 8x MI355X and GB200 profiles; none is labelled the absolute minimum. Source: https://github.com/vllm-project/recipes/blob/main/models/deepseek-ai/DeepSeek-V4-Pro.yaml?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **200-person team — evidence E:** An 8-GPU worker may be a capacity building block, but replicas depend on concurrency and output-token demand.
- **Commercial API — evidence B:** 8x H200 is documented with context capped at 800K to preserve KV headroom; production redundancy requires additional workers. Source: https://github.com/vllm-project/recipes/blob/main/models/deepseek-ai/DeepSeek-V4-Pro.yaml?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment

## Scenario evidence

- **Complex reasoning and coding — Public evidence:** The official model card publishes reasoning, coding and agent benchmark results. Source: https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Million-token retrieval and analysis — Vendor-stated:** The release is designed around efficient million-token context intelligence. Source: https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Premium private agent endpoint — ChinaAPI inference:** Use when quality has more value than minimum infrastructure cost.

## Commercial-use check

- License: MIT — https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/LICENSE?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- MaaS / hosted service: No model-specific MaaS restriction identified in the MIT license.
- Attribution: Retain the copyright and permission notice.

## Avoid or validate first

- Budget workstation inference
- Public API launch without at least one redundant worker

## Known limitations and open questions

- Some SM120 workstation kernels have unresolved reports
- Loading weights is not proof that the first forward pass succeeds

## Evidence boundary

Loading weights, completing a first forward pass and meeting a production latency/SLA are separate thresholds. The 200-person and commercial tiers require workload-specific measurement. License summaries are product research, not legal advice.

Generated by GENERATED BY scripts/gen_open_model_deployment.py.
