Open-weight deployment dossier · verified 2026-09-12

LLaDA2.2-flash: self-hosting hardware, scenarios and commercial license

Diffusion-language-model research checkpoint for agent reasoning, tool use and iterative editing, with an executable Transformers path but no vendor-confirmed production server yet.

100B non-embedding; 103B repository metadata
total parameters
128K tokens
context window
Apache-2.0
published license

Released 2026-07-16 · general agent

Deployment evidence

Ant Group InclusionAI · 100B non-embedding; 103B repository metadata total · sparse MoE; exact active count not published active · 128K tokens

PrecisionBF16Serving pathsTransformers

Best-fit scenarios

Diffusion-language-model research checkpoint for agent reasoning, tool use and iterative editing, with an executable Transformers path but no vendor-confirmed production server yet.

Not recommended
  • Advertising SGLang as vendor-supported while the card says support is coming soon
  • Treating the research generation loop as a production API topology

Diffusion agent research

Vendor-stated

The vendor presents LLaDA2.2-flash as a diffusion-language-model checkpoint for reasoning, coding and agent tasks.

Capability source

Iterative editing and error correction

Vendor-stated

Bidirectional denoising is positioned for revision and correction workflows; the claimed quality remains vendor-evaluated.

Capability source

Research-only local generation

ChinaAPI inference

The official Transformers example is executable, but the lack of a vendor-confirmed production server keeps this below a normal commercial-serving recommendation.

Official launch / minimum

Evidence C

Transformers path no hardware floor

The official Transformers generation path is executable, but the roughly 206GB BF16 repository and runtime overhead require operator planning; the vendor publishes no minimum GPU topology. Primary recipe

200-person team

Evidence E

Production server not confirmed

Do not size a shared service until a supported server path is validated against the actual diffusion steps, context and concurrency.

Commercial API

Evidence C

Not ready for recommendation

The vendor card marks SGLang deployment support as coming soon. Wait for an official serving path and then validate redundancy and admission control. Primary recipe

Commercial-use check

Apache-2.0

MaaS / hosted service: No model-specific MaaS restriction identified in Apache-2.0.

Attribution: Provide the license and required notices, preserve attribution notices, and mark modified files; trademark rights are not granted.

Read the primary license text

Known limitations and open questions
  • No ChinaAPI reproduction
  • Vendor-confirmed SGLang deployment is still marked coming soon
  • The active-parameter count and official minimum hardware topology are not published
  • Diffusion decoding has different latency and batching behavior from autoregressive servers

Evidence boundaries

A launch shape is not a production SLA.

The minimum tier records the smallest official or inference-framework configuration we found. The 200-person and commercial API tiers still require measurements against real prompt length, output length, concurrency, latency and redundancy targets.

Read the complete index methodology

  1. AChinaAPI reproduced
  2. Binference-framework official validated recipe
  3. Cmodel-vendor documented configuration
  4. Dthird-party reproduction
  5. Ecapacity estimate only

Compare before deploying

Review every verified model or compare hosted access.

The directory keeps model selection separate from the evidence and capacity details on this page.