# Step 3.7 Flash: Self-Hosting Hardware, Scenarios and Commercial License

> High-frequency multimodal agents, financial-document parsing, verified search loops and concurrent coding.

- Verified: 2026-09-12
- Released: 2026-05-28
- Canonical: https://chinaapi.ai/open-model-deployment/step-3.7-flash/
- Back to directory: https://chinaapi.ai/open-model-deployment/
- Machine-readable dataset: https://chinaapi.ai/data/open-model-deployment.json

## Model facts

- Vendor: StepFun
- Parameters: 198B total; approximately 11B active
- Native precision: BF16, FP8, NVFP4 and GGUF releases
- Context: 256K tokens
- Category: General and agentic foundation models
- Modalities: text, image → text
- Serving paths: vLLM, SGLang, Transformers, llama.cpp
- Official weights: https://huggingface.co/stepfun-ai/Step-3.7-Flash?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- Model documentation: https://github.com/stepfun-ai/Step-3.7-Flash?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment

## Deployment evidence

- **Official launch / minimum — evidence C:** The official GGUF path specifies about 120GB minimum unified memory/VRAM and recommends 128GB. Source: https://github.com/stepfun-ai/Step-3.7-Flash?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **200-person team — evidence E:** Start with one TP4 NVFP4 or TP8 FP8 worker and benchmark the actual image and context mix.
- **Commercial API — evidence C:** Official examples publish TP4 NVFP4 and TP8 FP8/BF16 serving shapes; replicas are still required for HA. Source: https://github.com/stepfun-ai/Step-3.7-Flash?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment

## Scenario evidence

- **Visual documents and UI-to-code — Vendor-stated:** Official materials highlight charts, GUIs, wireframes and structured-code extraction. Source: https://github.com/stepfun-ai/Step-3.7-Flash?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Search and tool orchestration — Public evidence:** The vendor publishes ClawEval, Toolathlon and tool-use evaluations. Source: https://github.com/stepfun-ai/Step-3.7-Flash?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Concurrent coding agents — Vendor-stated:** The model is explicitly engineered for high-frequency production agent workloads. Source: https://github.com/stepfun-ai/Step-3.7-Flash?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment

## Commercial-use check

- License: Apache-2.0 — https://github.com/stepfun-ai/Step-3.7-Flash/blob/main/LICENSE?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- MaaS / hosted service: No model-specific MaaS restriction identified in Apache-2.0.
- Attribution: Provide the license and notices, preserve attribution notices, and mark modified files.

## Avoid or validate first

- Using the 128GB local path as a production throughput claim
- Assuming vendor benchmark throughput transfers to long-context prefill

## Known limitations and open questions

- Published 400 tok/s is a vendor benchmark, not a universal SLA
- NVFP4 support depends on recent engine and GPU paths

## Evidence boundary

Loading weights, completing a first forward pass and meeting a production latency/SLA are separate thresholds. The 200-person and commercial tiers require workload-specific measurement. License summaries are product research, not legal advice.

Generated by GENERATED BY scripts/gen_open_model_deployment.py.
