# Spark X2.5: Self-Hosting Hardware, Scenarios and Commercial License

> Compact million-token multilingual model family for local agents, coding assistants and edge or desktop deployment.

- Verified: 2026-09-12
- Released: 2026-09-01
- Canonical: https://chinaapi.ai/open-model-deployment/spark-x2.5/
- Back to directory: https://chinaapi.ai/open-model-deployment/
- Machine-readable dataset: https://chinaapi.ai/data/open-model-deployment.json

## Model facts

- Vendor: XHToken
- Parameters: 4B primary; 1.7B official variant total; 4B dense primary active
- Native precision: BF16 with official FP8, INT8 and GGUF releases
- Context: 1M tokens
- Category: General and agentic foundation models
- Modalities: text → text
- Serving paths: vLLM, SGLang, llama.cpp, MLX, Ollama, LM Studio
- Official weights: Official 4B / primary: https://huggingface.co/XHToken/Spark-X2.5-4B?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment; Official 1.7B: https://huggingface.co/XHToken/Spark-X2.5-1.7B?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- Model documentation: https://github.com/XHToken/Spark-X2.5?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment

## Deployment evidence

- **Official launch / minimum — evidence C:** Official SGLang and local-runtime paths run on one device, but the vendor only says that sufficient memory is required and publishes no absolute VRAM floor for one-million-token context. Source: https://github.com/XHToken/Spark-X2.5?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **200-person team — evidence E:** Benchmark a 4B server worker and the actual language mix; use the 1.7B variant only after task-quality acceptance.
- **Commercial API — evidence C:** Official vLLM and SGLang paths exist; production still requires replicas, parser pinning and per-language evaluation. Source: https://github.com/XHToken/Spark-X2.5?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment

## Scenario evidence

- **Multilingual local agents — Vendor-stated:** The official release advertises support for more than 200 languages and a one-million-token context across the 4B and 1.7B variants. Source: https://github.com/XHToken/Spark-X2.5?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Coding and tool workflows — Vendor-stated:** The vendor positions the family for coding, function calling and agent use across server and local runtimes. Source: https://github.com/XHToken/Spark-X2.5?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Desktop and edge assistant — ChinaAPI inference:** The small dense variants and official GGUF, MLX, Ollama and LM Studio paths make local evaluation practical.

## Commercial-use check

- License: Apache-2.0 — https://github.com/XHToken/Spark-X2.5/blob/main/LICENSE?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- MaaS / hosted service: No model-specific MaaS restriction identified in Apache-2.0.
- Attribution: Provide the license and required notices, preserve attribution notices, and mark modified files; trademark rights are not granted.

## Avoid or validate first

- Assuming one-million-token context fits a small device because the weights do
- Treating more than 200 supported languages as equal quality in every language

## Known limitations and open questions

- No ChinaAPI reproduction or official absolute hardware minimum
- The one-million-token setting has workload-dependent memory and latency
- Multilingual breadth is a vendor claim and needs target-language validation

## Evidence boundary

Loading weights, completing a first forward pass and meeting a production latency/SLA are separate thresholds. The 200-person and commercial tiers require workload-specific measurement. License summaries are product research, not legal advice.

Generated by GENERATED BY scripts/gen_open_model_deployment.py.
