Open-weight deployment dossier · verified 2026-09-12

Step 3.7 Flash: self-hosting hardware, scenarios and commercial license

High-frequency multimodal agents, financial-document parsing, verified search loops and concurrent coding.

198B
total parameters
256K tokens
context window
Apache-2.0
published license

Released 2026-05-28 · general agent

Deployment evidence

StepFun · 198B total · approximately 11B active · 256K tokens

PrecisionBF16, FP8, NVFP4 and GGUF releasesServing pathsvLLM, SGLang, Transformers, llama.cpp

Best-fit scenarios

High-frequency multimodal agents, financial-document parsing, verified search loops and concurrent coding.

Not recommended
  • Using the 128GB local path as a production throughput claim
  • Assuming vendor benchmark throughput transfers to long-context prefill

Visual documents and UI-to-code

Vendor-stated

Official materials highlight charts, GUIs, wireframes and structured-code extraction.

Capability source

Search and tool orchestration

Public evidence

The vendor publishes ClawEval, Toolathlon and tool-use evaluations.

Capability source

Concurrent coding agents

Vendor-stated

The model is explicitly engineered for high-frequency production agent workloads.

Capability source

Official launch / minimum

Evidence C

Official local minimum

The official GGUF path specifies about 120GB minimum unified memory/VRAM and recommends 128GB. Primary recipe

200-person team

Evidence E

Production candidate

Start with one TP4 NVFP4 or TP8 FP8 worker and benchmark the actual image and context mix.

Commercial API

Evidence C

Official launch shapes

Official examples publish TP4 NVFP4 and TP8 FP8/BF16 serving shapes; replicas are still required for HA. Primary recipe

Commercial-use check

Apache-2.0

MaaS / hosted service: No model-specific MaaS restriction identified in Apache-2.0.

Attribution: Provide the license and notices, preserve attribution notices, and mark modified files.

Read the primary license text

Known limitations and open questions
  • Published 400 tok/s is a vendor benchmark, not a universal SLA
  • NVFP4 support depends on recent engine and GPU paths

Evidence boundaries

A launch shape is not a production SLA.

The minimum tier records the smallest official or inference-framework configuration we found. The 200-person and commercial API tiers still require measurements against real prompt length, output length, concurrency, latency and redundancy targets.

Read the complete index methodology

  1. AChinaAPI reproduced
  2. Binference-framework official validated recipe
  3. Cmodel-vendor documented configuration
  4. Dthird-party reproduction
  5. Ecapacity estimate only

Compare before deploying

Review every verified model or compare hosted access.

The directory keeps model selection separate from the evidence and capacity details on this page.