# LongCat-Flash-Lite-Sparse: Self-Hosting Hardware, Scenarios and Commercial License

> Sparse long-context coding, search and tool agents with an unusually small active path and an official single-H20 serving example.

- Verified: 2026-09-12
- Released: 2026-07-31
- Canonical: https://chinaapi.ai/open-model-deployment/longcat-flash-lite-sparse/
- Back to directory: https://chinaapi.ai/open-model-deployment/
- Machine-readable dataset: https://chinaapi.ai/data/open-model-deployment.json

## Model facts

- Vendor: Meituan
- Parameters: 69B total; approximately 3B active
- Native precision: BF16 / F32 checkpoint components
- Context: 1M tokens
- Category: General and agentic foundation models
- Modalities: text → text
- Serving paths: Transformers, SGLang
- Official weights: https://huggingface.co/meituan-longcat/LongCat-Flash-Lite-Sparse?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- Model documentation: https://huggingface.co/meituan-longcat/LongCat-Flash-Lite-Sparse?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment

## Deployment evidence

- **Official launch / minimum — evidence C:** The vendor documents SGLang serving on one H20-141G accelerator. This is an official runnable path, not a claim that smaller memory shapes cannot work. Source: https://huggingface.co/meituan-longcat/LongCat-Flash-Lite-Sparse?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **200-person team — evidence E:** Start with one measured H20-141G worker and replay normal and long-context traffic before selecting worker count.
- **Commercial API — evidence C:** The official SGLang path can seed a service, but production requires replicas, rate limits and measured cache behavior. Source: https://huggingface.co/meituan-longcat/LongCat-Flash-Lite-Sparse?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment

## Scenario evidence

- **Long-context coding agents — Vendor-stated:** The model card positions the sparse checkpoint for coding and agent tasks with up to one-million-token context. Source: https://huggingface.co/meituan-longcat/LongCat-Flash-Lite-Sparse?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Search and tool use — Public evidence:** Meituan publishes search, coding and tool-use evaluations; these are vendor results rather than ChinaAPI reproductions. Source: https://huggingface.co/meituan-longcat/LongCat-Flash-Lite-Sparse?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Cost-sensitive private endpoint — ChinaAPI inference:** The 69B/approximately-3B-active shape and single-node recipe make it a candidate for measured private serving where long context matters.

## Commercial-use check

- License: MIT — https://huggingface.co/meituan-longcat/LongCat-Flash-Lite-Sparse/blob/main/LICENSE?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- MaaS / hosted service: No model-specific MaaS restriction identified in the MIT license.
- Attribution: Retain the copyright and permission notice in copies or substantial portions of the weights.

## Avoid or validate first

- Assuming the 3B active path means 3B of weight memory
- Promising one-million-token concurrency from a single-request launch example

## Known limitations and open questions

- No ChinaAPI reproduction
- The single-accelerator example is not a throughput benchmark
- The vendor card does not publish a full production topology

## Evidence boundary

Loading weights, completing a first forward pass and meeting a production latency/SLA are separate thresholds. The 200-person and commercial tiers require workload-specific measurement. License summaries are product research, not legal advice.

Generated by GENERATED BY scripts/gen_open_model_deployment.py.
