Open-weight deployment dossier · verified 2026-09-12

Ling-3.0-flash: self-hosting hardware, scenarios and commercial license

Hybrid-linear sparse model for coding, research and long-horizon agents with official four- and eight-GPU serving shapes.

124B
total parameters
256K tokens
context window
MIT
published license

Released 2026-08-02 · general agent

Deployment evidence

Ant Group InclusionAI · 124B total · 5.1B active · 256K tokens

PrecisionBF16 and FP8Serving pathsSGLang, vLLM

Best-fit scenarios

Hybrid-linear sparse model for coding, research and long-horizon agents with official four- and eight-GPU serving shapes.

Not recommended
  • Sizing storage or GPU memory from 5.1B active parameters
  • Generalizing the vendor low-latency configuration to all prompt lengths

Coding and general agents

Vendor-stated

The model card positions Ling-3.0-flash for coding, reasoning and general agent workloads. The reported quality numbers are vendor results.

Capability source

Long-horizon research workflows

Vendor-stated

The 256K hybrid-linear design and documented HiCache serving path target long-context and long-horizon workloads.

Capability source

Low-latency sparse endpoint

ChinaAPI inference

The 124B/5.1B-active shape is a candidate for measured high-frequency serving, but the active count is not a memory requirement.

Official launch / minimum

Evidence C

Official vendor floor

The vendor documents a low-latency path on four 141GB-class H20-3e or Blackwell GPUs; 80GB H100/H800 systems use TP8. These are launch shapes, not capacity guarantees. Primary recipe

200-person team

Evidence E

Benchmark required

Benchmark one official-shape worker against the 200-seat context distribution, then add a second worker for availability if required.

Commercial API

Evidence C

Official framework paths

Official SGLang and vLLM commands are published; public service still needs redundant workers, cache monitoring and admission control. Primary recipe

Commercial-use check

MIT

MaaS / hosted service: No model-specific MaaS restriction identified in the MIT license.

Attribution: Retain the copyright and permission notice in copies or substantial portions of the weights.

Read the primary license text

Known limitations and open questions
  • No ChinaAPI hardware or quality reproduction
  • Published latency and quality results are vendor measurements
  • Long-context cache efficiency depends on the actual traffic distribution

Evidence boundaries

A launch shape is not a production SLA.

The minimum tier records the smallest official or inference-framework configuration we found. The 200-person and commercial API tiers still require measurements against real prompt length, output length, concurrency, latency and redundancy targets.

Read the complete index methodology

  1. AChinaAPI reproduced
  2. Binference-framework official validated recipe
  3. Cmodel-vendor documented configuration
  4. Dthird-party reproduction
  5. Ecapacity estimate only

Compare before deploying

Review every verified model or compare hosted access.

The directory keeps model selection separate from the evidence and capacity details on this page.