Open-weight deployment dossier · verified 2026-09-12

Unlimited-OCR: self-hosting hardware, scenarios and commercial license

Compact document-vision model for one-shot long-document OCR, multi-page PDF parsing and structured layout extraction.

3B
total parameters
32K tokens
context window
MIT
published license

Released 2026-06-22 · document vision

Deployment evidence

Baidu · 3B total · 3B dense active · 32K tokens

PrecisionBF16Serving pathsTransformers, vLLM, SGLang

Best-fit scenarios

Compact document-vision model for one-shot long-document OCR, multi-page PDF parsing and structured layout extraction.

Not recommended
  • Sending unbounded PDFs without page, pixel and token admission limits
  • Assuming OCR benchmark scores cover every language, handwriting or layout

Long-document OCR

Vendor-stated

Baidu positions the model for one-shot long-horizon OCR rather than page-fragment-only extraction.

Capability source

Multi-page PDF parsing

Vendor-stated

The official processor and examples cover document images and PDF-oriented workflows with a 32K generation configuration.

Capability source

Structured layout extraction API

ChinaAPI inference

The official vLLM and SGLang OpenAI-compatible paths make queued document extraction plausible after processor and post-processing validation.

Official launch / minimum

Evidence C

Official single gpu path no vram floor

The official Transformers example loads the model on one CUDA GPU, but the vendor does not publish an absolute minimum VRAM configuration. Primary recipe

200-person team

Evidence E

Document queue benchmark required

Size by pages, pixels, document length and peak jobs rather than seats; start with one measured queue worker.

Commercial API

Evidence C

Official framework paths

Official vLLM and SGLang serving paths exist; production needs replicated workers, document limits, processor pinning and post-processing observability. Primary recipe

Commercial-use check

MIT

MaaS / hosted service: No model-specific MaaS restriction identified in the MIT license.

Attribution: Retain the copyright and permission notice in copies or substantial portions of the weights.

Read the primary license text

Known limitations and open questions
  • No ChinaAPI hardware or OCR-quality reproduction
  • PDFs still require deterministic image conversion and document post-processing
  • The 32K setting is taken from the official inference configuration, not a full-context concurrency result

Evidence boundaries

A launch shape is not a production SLA.

The minimum tier records the smallest official or inference-framework configuration we found. The 200-person and commercial API tiers still require measurements against real prompt length, output length, concurrency, latency and redundancy targets.

Read the complete index methodology

  1. AChinaAPI reproduced
  2. Binference-framework official validated recipe
  3. Cmodel-vendor documented configuration
  4. Dthird-party reproduction
  5. Ecapacity estimate only

Compare before deploying

Review every verified model or compare hosted access.

The directory keeps model selection separate from the evidence and capacity details on this page.