Open-weight deployment dossier · verified 2026-09-12

Ling-3.0-tiny: self-hosting hardware, scenarios and commercial license

Compact sparse model for local coding and agent assistants on Apple silicon, DGX Spark and conventional GPU servers.

7.9B
total parameters
256K tokens
context window
MIT
published license

Released 2026-08-10 · general agent

Deployment evidence

Ant Group InclusionAI · 7.9B total · 1.3B active · 256K tokens

PrecisionBF16, FP8 and INT4Serving pathsSGLang, vLLM, Ollama

Best-fit scenarios

Compact sparse model for local coding and agent assistants on Apple silicon, DGX Spark and conventional GPU servers.

Not recommended
  • Reading the published 8K memory result as proof that 256K context fits the same device
  • Treating desktop token rates as a multi-user API SLA

Apple-silicon local assistant

Vendor-stated

The official card includes MacBook and Mac mini demonstrations and reports M4 Pro measurements for local use.

Capability source

DGX Spark coding agent

Vendor-stated

The vendor publishes DGX Spark FP8 measurements and local agent demonstrations; results remain vendor-measured.

Capability source

Low-cost private endpoint

ChinaAPI inference

The 7.9B/1.3B-active shape and multiple precisions make it a practical first benchmark before larger Ling variants.

Official launch / minimum

Evidence C

Official local paths

The vendor validates M4 Pro and DGX Spark local paths at 8K, including roughly 8.34 GiB peak memory in one published setup. Its separate full-256K SGLang example uses one 141GB accelerator. Primary recipe

200-person team

Evidence E

Benchmark required

Use a server-class measured worker for shared use; desktop results are useful for individual deployment, not 200-seat sizing.

Commercial API

Evidence C

Official framework paths

Official SGLang and vLLM serving paths exist, but replicas and workload-specific context limits remain necessary. Primary recipe

Commercial-use check

MIT

MaaS / hosted service: No model-specific MaaS restriction identified in the MIT license.

Attribution: Retain the copyright and permission notice in copies or substantial portions of the weights.

Read the primary license text

Known limitations and open questions
  • Vendor speed and memory results were not reproduced by ChinaAPI
  • The 8K local benchmark does not validate full 256K context
  • INT4 and Apple-silicon paths may use different runtimes than the server examples

Evidence boundaries

A launch shape is not a production SLA.

The minimum tier records the smallest official or inference-framework configuration we found. The 200-person and commercial API tiers still require measurements against real prompt length, output length, concurrency, latency and redundancy targets.

Read the complete index methodology

  1. AChinaAPI reproduced
  2. Binference-framework official validated recipe
  3. Cmodel-vendor documented configuration
  4. Dthird-party reproduction
  5. Ecapacity estimate only

Compare before deploying

Review every verified model or compare hosted access.

The directory keeps model selection separate from the evidence and capacity details on this page.