Open-weight deployment dossier · verified 2026-09-12

AREX-Turbo: self-hosting hardware, scenarios and commercial license

Compact deep-research model for long-horizon information seeking, evidence aggregation and multi-constraint verification.

4B dense
total parameters
256K tokens
context window
Apache-2.0
published license

Released 2026-07-23 · general agent

Deployment evidence

BAAI · 4B dense total · 4B dense active · 256K tokens

PrecisionBF16Serving pathsTransformers, vLLM, SGLang

Best-fit scenarios

Compact deep-research model for long-horizon information seeking, evidence aggregation and multi-constraint verification.

Not recommended
  • Treating long-context support as proof of source reliability
  • Deploying research answers without citation and retrieval validation

Deep research agents

Vendor-stated

BAAI positions AREX-Turbo for long-horizon information seeking, evidence aggregation and multi-step research.

Capability source

Multi-constraint verification

Public evidence

The official release publishes deep-research and verification evaluations; results are vendor measurements rather than ChinaAPI reproductions.

Capability source

Private research endpoint

ChinaAPI inference

The dense 4B shape and three official inference paths make it a practical private research-agent candidate.

Official launch / minimum

Evidence C

Official paths no hardware floor

Transformers, vLLM and SGLang launch paths are documented, but the vendor does not publish an absolute minimum GPU or memory configuration. Primary recipe

200-person team

Evidence E

Benchmark required

Start with one measured worker and size by concurrent research trajectories, tool calls and context accumulation rather than employee count.

Commercial API

Evidence C

Official framework paths

Official vLLM and SGLang servers are available; public use still needs replicas, retrieval observability and citation-quality controls. Primary recipe

Commercial-use check

Apache-2.0

MaaS / hosted service: No model-specific MaaS restriction identified in Apache-2.0.

Attribution: Provide the license and required notices, preserve attribution notices, and mark modified files; trademark rights are not granted.

Read the primary license text

Known limitations and open questions
  • No ChinaAPI reproduction or official hardware minimum
  • Vendor evaluations do not validate a specific retrieval stack
  • Long-horizon quality depends on tools, sources and orchestration outside the weights

Evidence boundaries

A launch shape is not a production SLA.

The minimum tier records the smallest official or inference-framework configuration we found. The 200-person and commercial API tiers still require measurements against real prompt length, output length, concurrency, latency and redundancy targets.

Read the complete index methodology

  1. AChinaAPI reproduced
  2. Binference-framework official validated recipe
  3. Cmodel-vendor documented configuration
  4. Dthird-party reproduction
  5. Ecapacity estimate only

Compare before deploying

Review every verified model or compare hosted access.

The directory keeps model selection separate from the evidence and capacity details on this page.