Open-weight deployment dossier · verified 2026-09-12

Spark X2.5: self-hosting hardware, scenarios and commercial license

Compact million-token multilingual model family for local agents, coding assistants and edge or desktop deployment.

4B primary; 1.7B official variant
total parameters
1M tokens
context window
Apache-2.0
published license

Released 2026-09-01 · general agent

Deployment evidence

XHToken · 4B primary; 1.7B official variant total · 4B dense primary active · 1M tokens

PrecisionBF16 with official FP8, INT8 and GGUF releasesServing pathsvLLM, SGLang, llama.cpp, MLX, Ollama, LM Studio

Best-fit scenarios

Compact million-token multilingual model family for local agents, coding assistants and edge or desktop deployment.

Not recommended
  • Assuming one-million-token context fits a small device because the weights do
  • Treating more than 200 supported languages as equal quality in every language

Multilingual local agents

Vendor-stated

The official release advertises support for more than 200 languages and a one-million-token context across the 4B and 1.7B variants.

Capability source

Coding and tool workflows

Vendor-stated

The vendor positions the family for coding, function calling and agent use across server and local runtimes.

Capability source

Desktop and edge assistant

ChinaAPI inference

The small dense variants and official GGUF, MLX, Ollama and LM Studio paths make local evaluation practical.

Official launch / minimum

Evidence C

Official single device path no vram floor

Official SGLang and local-runtime paths run on one device, but the vendor only says that sufficient memory is required and publishes no absolute VRAM floor for one-million-token context. Primary recipe

200-person team

Evidence E

Benchmark required

Benchmark a 4B server worker and the actual language mix; use the 1.7B variant only after task-quality acceptance.

Commercial API

Evidence C

Official framework paths

Official vLLM and SGLang paths exist; production still requires replicas, parser pinning and per-language evaluation. Primary recipe

Commercial-use check

Apache-2.0

MaaS / hosted service: No model-specific MaaS restriction identified in Apache-2.0.

Attribution: Provide the license and required notices, preserve attribution notices, and mark modified files; trademark rights are not granted.

Read the primary license text

Known limitations and open questions
  • No ChinaAPI reproduction or official absolute hardware minimum
  • The one-million-token setting has workload-dependent memory and latency
  • Multilingual breadth is a vendor claim and needs target-language validation

Evidence boundaries

A launch shape is not a production SLA.

The minimum tier records the smallest official or inference-framework configuration we found. The 200-person and commercial API tiers still require measurements against real prompt length, output length, concurrency, latency and redundancy targets.

Read the complete index methodology

  1. AChinaAPI reproduced
  2. Binference-framework official validated recipe
  3. Cmodel-vendor documented configuration
  4. Dthird-party reproduction
  5. Ecapacity estimate only

Compare before deploying

Review every verified model or compare hosted access.

The directory keeps model selection separate from the evidence and capacity details on this page.