Released 2026-05-28 · general agent
Deployment evidence
StepFun · 198B total · approximately 11B active · 256K tokens
Best-fit scenarios
High-frequency multimodal agents, financial-document parsing, verified search loops and concurrent coding.
Not recommended
- Using the 128GB local path as a production throughput claim
- Assuming vendor benchmark throughput transfers to long-context prefill
Visual documents and UI-to-code
Vendor-statedOfficial materials highlight charts, GUIs, wireframes and structured-code extraction.
Capability sourceSearch and tool orchestration
Public evidenceThe vendor publishes ClawEval, Toolathlon and tool-use evaluations.
Capability sourceConcurrent coding agents
Vendor-statedThe model is explicitly engineered for high-frequency production agent workloads.
Capability sourceOfficial launch / minimum
Evidence COfficial local minimum
The official GGUF path specifies about 120GB minimum unified memory/VRAM and recommends 128GB. Primary recipe
200-person team
Evidence EProduction candidate
Start with one TP4 NVFP4 or TP8 FP8 worker and benchmark the actual image and context mix.
Commercial API
Evidence COfficial launch shapes
Official examples publish TP4 NVFP4 and TP8 FP8/BF16 serving shapes; replicas are still required for HA. Primary recipe
Commercial-use check
Apache-2.0
MaaS / hosted service: No model-specific MaaS restriction identified in Apache-2.0.
Attribution: Provide the license and notices, preserve attribution notices, and mark modified files.
Known limitations and open questions
- Published 400 tok/s is a vendor benchmark, not a universal SLA
- NVFP4 support depends on recent engine and GPU paths
