Released 2026-08-10 · general agent
Deployment evidence
Ant Group InclusionAI · 7.9B total · 1.3B active · 256K tokens
Best-fit scenarios
Compact sparse model for local coding and agent assistants on Apple silicon, DGX Spark and conventional GPU servers.
Not recommended
- Reading the published 8K memory result as proof that 256K context fits the same device
- Treating desktop token rates as a multi-user API SLA
Apple-silicon local assistant
Vendor-statedThe official card includes MacBook and Mac mini demonstrations and reports M4 Pro measurements for local use.
Capability sourceDGX Spark coding agent
Vendor-statedThe vendor publishes DGX Spark FP8 measurements and local agent demonstrations; results remain vendor-measured.
Capability sourceLow-cost private endpoint
ChinaAPI inferenceThe 7.9B/1.3B-active shape and multiple precisions make it a practical first benchmark before larger Ling variants.
Official launch / minimum
Evidence COfficial local paths
The vendor validates M4 Pro and DGX Spark local paths at 8K, including roughly 8.34 GiB peak memory in one published setup. Its separate full-256K SGLang example uses one 141GB accelerator. Primary recipe
200-person team
Evidence EBenchmark required
Use a server-class measured worker for shared use; desktop results are useful for individual deployment, not 200-seat sizing.
Commercial API
Evidence COfficial framework paths
Official SGLang and vLLM serving paths exist, but replicas and workload-specific context limits remain necessary. Primary recipe
Commercial-use check
MIT
MaaS / hosted service: No model-specific MaaS restriction identified in the MIT license.
Attribution: Retain the copyright and permission notice in copies or substantial portions of the weights.
Known limitations and open questions
- Vendor speed and memory results were not reproduced by ChinaAPI
- The 8K local benchmark does not validate full 256K context
- INT4 and Apple-silicon paths may use different runtimes than the server examples
