Released 2026-07-31 · general agent
Deployment evidence
Meituan · 69B total · approximately 3B active · 1M tokens
Best-fit scenarios
Sparse long-context coding, search and tool agents with an unusually small active path and an official single-H20 serving example.
Not recommended
- Assuming the 3B active path means 3B of weight memory
- Promising one-million-token concurrency from a single-request launch example
Long-context coding agents
Vendor-statedThe model card positions the sparse checkpoint for coding and agent tasks with up to one-million-token context.
Capability sourceSearch and tool use
Public evidenceMeituan publishes search, coding and tool-use evaluations; these are vendor results rather than ChinaAPI reproductions.
Capability sourceCost-sensitive private endpoint
ChinaAPI inferenceThe 69B/approximately-3B-active shape and single-node recipe make it a candidate for measured private serving where long context matters.
Official launch / minimum
Evidence COfficial single accelerator path
The vendor documents SGLang serving on one H20-141G accelerator. This is an official runnable path, not a claim that smaller memory shapes cannot work. Primary recipe
200-person team
Evidence EBenchmark required
Start with one measured H20-141G worker and replay normal and long-context traffic before selecting worker count.
Commercial API
Evidence CVendor sglang path
The official SGLang path can seed a service, but production requires replicas, rate limits and measured cache behavior. Primary recipe
Commercial-use check
MIT
MaaS / hosted service: No model-specific MaaS restriction identified in the MIT license.
Attribution: Retain the copyright and permission notice in copies or substantial portions of the weights.
Known limitations and open questions
- No ChinaAPI reproduction
- The single-accelerator example is not a throughput benchmark
- The vendor card does not publish a full production topology
