# LLaDA2.2-flash: Self-Hosting Hardware, Scenarios and Commercial License

> Diffusion-language-model research checkpoint for agent reasoning, tool use and iterative editing, with an executable Transformers path but no vendor-confirmed production server yet.

- Verified: 2026-09-12
- Released: 2026-07-16
- Canonical: https://chinaapi.ai/open-model-deployment/llada2.2-flash/
- Back to directory: https://chinaapi.ai/open-model-deployment/
- Machine-readable dataset: https://chinaapi.ai/data/open-model-deployment.json

## Model facts

- Vendor: Ant Group InclusionAI
- Parameters: 100B non-embedding; 103B repository metadata total; sparse MoE; exact active count not published active
- Native precision: BF16
- Context: 128K tokens
- Category: General and agentic foundation models
- Modalities: text → text
- Serving paths: Transformers
- Official weights: https://huggingface.co/inclusionAI/LLaDA2.2-flash?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- Model documentation: https://huggingface.co/inclusionAI/LLaDA2.2-flash?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment

## Deployment evidence

- **Official launch / minimum — evidence C:** The official Transformers generation path is executable, but the roughly 206GB BF16 repository and runtime overhead require operator planning; the vendor publishes no minimum GPU topology. Source: https://huggingface.co/inclusionAI/LLaDA2.2-flash?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **200-person team — evidence E:** Do not size a shared service until a supported server path is validated against the actual diffusion steps, context and concurrency.
- **Commercial API — evidence C:** The vendor card marks SGLang deployment support as coming soon. Wait for an official serving path and then validate redundancy and admission control. Source: https://huggingface.co/inclusionAI/LLaDA2.2-flash?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment

## Scenario evidence

- **Diffusion agent research — Vendor-stated:** The vendor presents LLaDA2.2-flash as a diffusion-language-model checkpoint for reasoning, coding and agent tasks. Source: https://huggingface.co/inclusionAI/LLaDA2.2-flash?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Iterative editing and error correction — Vendor-stated:** Bidirectional denoising is positioned for revision and correction workflows; the claimed quality remains vendor-evaluated. Source: https://huggingface.co/inclusionAI/LLaDA2.2-flash?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Research-only local generation — ChinaAPI inference:** The official Transformers example is executable, but the lack of a vendor-confirmed production server keeps this below a normal commercial-serving recommendation.

## Commercial-use check

- License: Apache-2.0 — https://huggingface.co/inclusionAI/LLaDA2.2-flash/blob/main/LICENSE?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- MaaS / hosted service: No model-specific MaaS restriction identified in Apache-2.0.
- Attribution: Provide the license and required notices, preserve attribution notices, and mark modified files; trademark rights are not granted.

## Avoid or validate first

- Advertising SGLang as vendor-supported while the card says support is coming soon
- Treating the research generation loop as a production API topology

## Known limitations and open questions

- No ChinaAPI reproduction
- Vendor-confirmed SGLang deployment is still marked coming soon
- The active-parameter count and official minimum hardware topology are not published
- Diffusion decoding has different latency and batching behavior from autoregressive servers

## Evidence boundary

Loading weights, completing a first forward pass and meeting a production latency/SLA are separate thresholds. The 200-person and commercial tiers require workload-specific measurement. License summaries are product research, not legal advice.

Generated by GENERATED BY scripts/gen_open_model_deployment.py.
