Released 2026-08-02 · general agent
Deployment evidence
Ant Group InclusionAI · 124B total · 5.1B active · 256K tokens
Best-fit scenarios
Hybrid-linear sparse model for coding, research and long-horizon agents with official four- and eight-GPU serving shapes.
Not recommended
- Sizing storage or GPU memory from 5.1B active parameters
- Generalizing the vendor low-latency configuration to all prompt lengths
Coding and general agents
Vendor-statedThe model card positions Ling-3.0-flash for coding, reasoning and general agent workloads. The reported quality numbers are vendor results.
Capability sourceLong-horizon research workflows
Vendor-statedThe 256K hybrid-linear design and documented HiCache serving path target long-context and long-horizon workloads.
Capability sourceLow-latency sparse endpoint
ChinaAPI inferenceThe 124B/5.1B-active shape is a candidate for measured high-frequency serving, but the active count is not a memory requirement.
Official launch / minimum
Evidence COfficial vendor floor
The vendor documents a low-latency path on four 141GB-class H20-3e or Blackwell GPUs; 80GB H100/H800 systems use TP8. These are launch shapes, not capacity guarantees. Primary recipe
200-person team
Evidence EBenchmark required
Benchmark one official-shape worker against the 200-seat context distribution, then add a second worker for availability if required.
Commercial API
Evidence COfficial framework paths
Official SGLang and vLLM commands are published; public service still needs redundant workers, cache monitoring and admission control. Primary recipe
Commercial-use check
MIT
MaaS / hosted service: No model-specific MaaS restriction identified in the MIT license.
Attribution: Retain the copyright and permission notice in copies or substantial portions of the weights.
Known limitations and open questions
- No ChinaAPI hardware or quality reproduction
- Published latency and quality results are vendor measurements
- Long-context cache efficiency depends on the actual traffic distribution
