Released 2026-07-31 · general agent
Deployment evidence
DeepSeek · 304B checkpoint metadata (284B target model plus attached DSpark draft module) total · 13B target path; DSpark draft module attached active · 1M tokens
Best-fit scenarios
The superseding DeepSeek V4 Flash checkpoint for high-frequency coding, reasoning and agent traffic, with an attached DSpark speculative draft module.
Not recommended
- Treating Flash as quality-equivalent to Pro on every task
- Counting 13B active parameters as the checkpoint memory requirement
- Launching a public API without measured parser, speculative-decoding and long-context behavior
High-frequency coding agents
Public evidenceThe 0731 card supersedes the preview checkpoint and reports coding and agent evaluations with the attached DSpark draft module. The scores remain vendor evaluations, not ChinaAPI reproductions.
Capability sourceMillion-token analysis
Vendor-statedThe target model retains a one-million-token context and a 284B/13B-active MoE shape; the published 304B checkpoint metadata also counts the attached DSpark module.
Capability sourceThroughput-sensitive private endpoint
ChinaAPI inferenceThe official vLLM and SGLang launch paths make the superseding checkpoint a deployment candidate where concurrency matters more than the V4 Pro quality ceiling.
Official launch / minimum
Evidence BOfficial framework floor
The verified vLLM recipe serves the checkpoint on one four-GPU GB300 node at TP4. This is a framework-validated launch shape, not a throughput or concurrency result. Primary recipe
200-person team
Evidence BValidated worker benchmark required
Start with one measured TP4 GB300 worker and replay the defined 200-seat context, concurrency and reasoning mix; add an independent worker when maintenance or failure recovery is required. Primary recipe
Commercial API
Evidence BOfficial framework paths
vLLM publishes a verified four-GPU GB300 recipe and the vendor documents SGLang TP4. A commercial service still needs independent replicas, admission control, pinned encoding and speculative-decoding validation. Primary recipe
Commercial-use check
MIT
MaaS / hosted service: No model-specific MaaS restriction identified in the MIT license.
Attribution: Retain the copyright and permission notice.
Known limitations and open questions
- No ChinaAPI hardware reproduction or traffic benchmark
- The 304B checkpoint metadata includes the attached draft module and must not be compared directly with the target model's 284B architecture
- The release uses a dedicated encoding implementation rather than a Jinja chat template
- Reasoning effort, parser behavior and DSpark acceptance require workload-specific validation
