# Ling-3.0-tiny: Self-Hosting Hardware, Scenarios and Commercial License

> Compact sparse model for local coding and agent assistants on Apple silicon, DGX Spark and conventional GPU servers.

- Verified: 2026-09-12
- Released: 2026-08-10
- Canonical: https://chinaapi.ai/open-model-deployment/ling-3.0-tiny/
- Back to directory: https://chinaapi.ai/open-model-deployment/
- Machine-readable dataset: https://chinaapi.ai/data/open-model-deployment.json

## Model facts

- Vendor: Ant Group InclusionAI
- Parameters: 7.9B total; 1.3B active
- Native precision: BF16, FP8 and INT4
- Context: 256K tokens
- Category: General and agentic foundation models
- Modalities: text → text
- Serving paths: SGLang, vLLM, Ollama
- Official weights: https://huggingface.co/inclusionAI/Ling-3.0-tiny?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- Model documentation: https://huggingface.co/inclusionAI/Ling-3.0-tiny?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment

## Deployment evidence

- **Official launch / minimum — evidence C:** The vendor validates M4 Pro and DGX Spark local paths at 8K, including roughly 8.34 GiB peak memory in one published setup. Its separate full-256K SGLang example uses one 141GB accelerator. Source: https://huggingface.co/inclusionAI/Ling-3.0-tiny?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **200-person team — evidence E:** Use a server-class measured worker for shared use; desktop results are useful for individual deployment, not 200-seat sizing.
- **Commercial API — evidence C:** Official SGLang and vLLM serving paths exist, but replicas and workload-specific context limits remain necessary. Source: https://huggingface.co/inclusionAI/Ling-3.0-tiny?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment

## Scenario evidence

- **Apple-silicon local assistant — Vendor-stated:** The official card includes MacBook and Mac mini demonstrations and reports M4 Pro measurements for local use. Source: https://huggingface.co/inclusionAI/Ling-3.0-tiny?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **DGX Spark coding agent — Vendor-stated:** The vendor publishes DGX Spark FP8 measurements and local agent demonstrations; results remain vendor-measured. Source: https://huggingface.co/inclusionAI/Ling-3.0-tiny?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Low-cost private endpoint — ChinaAPI inference:** The 7.9B/1.3B-active shape and multiple precisions make it a practical first benchmark before larger Ling variants.

## Commercial-use check

- License: MIT — https://huggingface.co/inclusionAI/Ling-3.0-tiny/blob/main/LICENSE?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- MaaS / hosted service: No model-specific MaaS restriction identified in the MIT license.
- Attribution: Retain the copyright and permission notice in copies or substantial portions of the weights.

## Avoid or validate first

- Reading the published 8K memory result as proof that 256K context fits the same device
- Treating desktop token rates as a multi-user API SLA

## Known limitations and open questions

- Vendor speed and memory results were not reproduced by ChinaAPI
- The 8K local benchmark does not validate full 256K context
- INT4 and Apple-silicon paths may use different runtimes than the server examples

## Evidence boundary

Loading weights, completing a first forward pass and meeting a production latency/SLA are separate thresholds. The 200-person and commercial tiers require workload-specific measurement. License summaries are product research, not legal advice.

Generated by GENERATED BY scripts/gen_open_model_deployment.py.
