Open-weight deployment dossier · verified 2026-09-12

LongCat-Flash-Lite-Sparse: self-hosting hardware, scenarios and commercial license

Sparse long-context coding, search and tool agents with an unusually small active path and an official single-H20 serving example.

69B
total parameters
1M tokens
context window
MIT
published license

Released 2026-07-31 · general agent

Deployment evidence

Meituan · 69B total · approximately 3B active · 1M tokens

PrecisionBF16 / F32 checkpoint componentsServing pathsTransformers, SGLang

Best-fit scenarios

Sparse long-context coding, search and tool agents with an unusually small active path and an official single-H20 serving example.

Not recommended
  • Assuming the 3B active path means 3B of weight memory
  • Promising one-million-token concurrency from a single-request launch example

Long-context coding agents

Vendor-stated

The model card positions the sparse checkpoint for coding and agent tasks with up to one-million-token context.

Capability source

Search and tool use

Public evidence

Meituan publishes search, coding and tool-use evaluations; these are vendor results rather than ChinaAPI reproductions.

Capability source

Cost-sensitive private endpoint

ChinaAPI inference

The 69B/approximately-3B-active shape and single-node recipe make it a candidate for measured private serving where long context matters.

Official launch / minimum

Evidence C

Official single accelerator path

The vendor documents SGLang serving on one H20-141G accelerator. This is an official runnable path, not a claim that smaller memory shapes cannot work. Primary recipe

200-person team

Evidence E

Benchmark required

Start with one measured H20-141G worker and replay normal and long-context traffic before selecting worker count.

Commercial API

Evidence C

Vendor sglang path

The official SGLang path can seed a service, but production requires replicas, rate limits and measured cache behavior. Primary recipe

Commercial-use check

MIT

MaaS / hosted service: No model-specific MaaS restriction identified in the MIT license.

Attribution: Retain the copyright and permission notice in copies or substantial portions of the weights.

Read the primary license text

Known limitations and open questions
  • No ChinaAPI reproduction
  • The single-accelerator example is not a throughput benchmark
  • The vendor card does not publish a full production topology

Evidence boundaries

A launch shape is not a production SLA.

The minimum tier records the smallest official or inference-framework configuration we found. The 200-person and commercial API tiers still require measurements against real prompt length, output length, concurrency, latency and redundancy targets.

Read the complete index methodology

  1. AChinaAPI reproduced
  2. Binference-framework official validated recipe
  3. Cmodel-vendor documented configuration
  4. Dthird-party reproduction
  5. Ecapacity estimate only

Compare before deploying

Review every verified model or compare hosted access.

The directory keeps model selection separate from the evidence and capacity details on this page.