Open-weight deployment dossier · verified 2026-09-12

DeepSeek V4 Flash 0731: self-hosting hardware, scenarios and commercial license

The superseding DeepSeek V4 Flash checkpoint for high-frequency coding, reasoning and agent traffic, with an attached DSpark speculative draft module.

304B checkpoint metadata (284B target model plus attached DSpark draft module)
total parameters
1M tokens
context window
MIT
published license

Released 2026-07-31 · general agent

Deployment evidence

DeepSeek · 304B checkpoint metadata (284B target model plus attached DSpark draft module) total · 13B target path; DSpark draft module attached active · 1M tokens

PrecisionMXFP4 routed experts with FP8/BF16 components; attached DSpark draft weightsServing pathsTransformers, vLLM, SGLang

Best-fit scenarios

The superseding DeepSeek V4 Flash checkpoint for high-frequency coding, reasoning and agent traffic, with an attached DSpark speculative draft module.

Not recommended
  • Treating Flash as quality-equivalent to Pro on every task
  • Counting 13B active parameters as the checkpoint memory requirement
  • Launching a public API without measured parser, speculative-decoding and long-context behavior

High-frequency coding agents

Public evidence

The 0731 card supersedes the preview checkpoint and reports coding and agent evaluations with the attached DSpark draft module. The scores remain vendor evaluations, not ChinaAPI reproductions.

Capability source

Million-token analysis

Vendor-stated

The target model retains a one-million-token context and a 284B/13B-active MoE shape; the published 304B checkpoint metadata also counts the attached DSpark module.

Capability source

Throughput-sensitive private endpoint

ChinaAPI inference

The official vLLM and SGLang launch paths make the superseding checkpoint a deployment candidate where concurrency matters more than the V4 Pro quality ceiling.

Official launch / minimum

Evidence B

Official framework floor

The verified vLLM recipe serves the checkpoint on one four-GPU GB300 node at TP4. This is a framework-validated launch shape, not a throughput or concurrency result. Primary recipe

200-person team

Evidence B

Validated worker benchmark required

Start with one measured TP4 GB300 worker and replay the defined 200-seat context, concurrency and reasoning mix; add an independent worker when maintenance or failure recovery is required. Primary recipe

Commercial API

Evidence B

Official framework paths

vLLM publishes a verified four-GPU GB300 recipe and the vendor documents SGLang TP4. A commercial service still needs independent replicas, admission control, pinned encoding and speculative-decoding validation. Primary recipe

Commercial-use check

MIT

MaaS / hosted service: No model-specific MaaS restriction identified in the MIT license.

Attribution: Retain the copyright and permission notice.

Read the primary license text

Known limitations and open questions
  • No ChinaAPI hardware reproduction or traffic benchmark
  • The 304B checkpoint metadata includes the attached draft module and must not be compared directly with the target model's 284B architecture
  • The release uses a dedicated encoding implementation rather than a Jinja chat template
  • Reasoning effort, parser behavior and DSpark acceptance require workload-specific validation

Evidence boundaries

A launch shape is not a production SLA.

The minimum tier records the smallest official or inference-framework configuration we found. The 200-person and commercial API tiers still require measurements against real prompt length, output length, concurrency, latency and redundancy targets.

Read the complete index methodology

  1. AChinaAPI reproduced
  2. Binference-framework official validated recipe
  3. Cmodel-vendor documented configuration
  4. Dthird-party reproduction
  5. Ecapacity estimate only

Compare before deploying

Review every verified model or compare hosted access.

The directory keeps model selection separate from the evidence and capacity details on this page.