Open-weight deployment dossier · verified 2026-09-12

LongCat 2.0: self-hosting hardware, scenarios and commercial license

Long-horizon coding, search, repository edits and tool-driven agents on GPU or NPU clusters.

1.6T
total parameters
1M tokens
context window
MIT
published license

Released 2026-07-02 · general agent

Deployment evidence

Meituan · 1.6T total · approximately 48B active · 1M tokens

PrecisionVendor checkpoint; exact serving precision varies by recipeServing pathsSGLang, SGLang-FluentLLM

Best-fit scenarios

Long-horizon coding, search, repository edits and tool-driven agents on GPU or NPU clusters.

Not recommended
  • Workstation deployment
  • Quoting in-house benchmark results as ChinaAPI reproduction

Repository-scale coding

Vendor-stated

Official materials emphasize repository edits and integrations with coding-agent harnesses.

Capability source

Search and general agents

Public evidence

The vendor publishes coding, BrowseComp, RWSearch and agent evaluations.

Capability source

NPU-based sovereign deployment

Vendor-stated

An official SGLang-FluentLLM NPU serving path is linked.

Capability source

Official launch / minimum

Evidence C

Not published

GPU and NPU paths are documented, but the vendor does not publish an absolute minimum hardware shape.

200-person team

Evidence E

Cluster sizing required

A 1.6T model needs a measured cluster design; employee count alone is not a capacity input.

Commercial API

Evidence C

Official serving paths

Official GPU and NPU serving paths exist; topology, redundancy and throughput remain operator-specific. Primary recipe

Commercial-use check

MIT

MaaS / hosted service: No model-specific MaaS restriction identified in the MIT license.

Attribution: Retain the copyright and permission notice.

Read the primary license text

Known limitations and open questions
  • No official minimum GPU count
  • Most published evaluation values are vendor-measured

Evidence boundaries

A launch shape is not a production SLA.

The minimum tier records the smallest official or inference-framework configuration we found. The 200-person and commercial API tiers still require measurements against real prompt length, output length, concurrency, latency and redundancy targets.

Read the complete index methodology

  1. AChinaAPI reproduced
  2. Binference-framework official validated recipe
  3. Cmodel-vendor documented configuration
  4. Dthird-party reproduction
  5. Ecapacity estimate only

Compare before deploying

Review every verified model or compare hosted access.

The directory keeps model selection separate from the evidence and capacity details on this page.