Open-weight deployment dossier · verified 2026-09-12

NAVA: self-hosting hardware, scenarios and commercial license

Synchronized audio-video generation with multi-speaker timbre control, camera prompting and image continuation.

6.3B backbone
total parameters
Not applicable
context window
Apache-2.0
published license

Released 2026-05-28 · audio video generation

Deployment evidence

Baidu ERNIE Team · 6.3B backbone total · 6.3B active · Not applicable

PrecisionBF16Serving pathsPyTorch, Ulysses sequence parallel

Best-fit scenarios

Synchronized audio-video generation with multi-speaker timbre control, camera prompting and image continuation.

Not recommended
  • Unconsented face or voice cloning
  • Low-latency interactive video generation

Synchronized short-form audio-video

Vendor-stated

NAVA jointly generates video, scene audio and speech rather than aligning separate outputs after generation.

Capability source

Multi-speaker and reference-timbre scenes

Vendor-stated

The official checkpoint supports up to two reference voices bound to speech spans.

Capability source

720p creative generation

Public evidence

The vendor reports VerseBench synchronization and quality results plus an 8-GPU fast path.

Capability source

Official launch / minimum

Evidence C

Single gpu supported no vram

The model card supports single-GPU inference but does not publish exact minimum VRAM. Primary recipe

200-person team

Evidence E

Media queue required

Size a queued render farm by jobs per hour, resolution and duration; office-seat assumptions do not apply.

Commercial API

Evidence C

Official 8gpu reference

The official Ulysses SP8 path reports roughly one minute for a 720p synchronized clip; HA needs additional workers. Primary recipe

Commercial-use check

Apache-2.0

MaaS / hosted service: The model card states Apache-2.0, but bundled LTX audio-VAE artifacts carry an additional community license that must be reviewed for the shipped stack.

Attribution: Preserve Apache notices and the notices/licenses for bundled upstream components.

Prohibited uses: The model card prohibits depicting real persons without consent, including face or voice likeness reproduction.

Read the primary license text

Known limitations and open questions
  • Default clips are about 6–10 seconds
  • The full dependency stack includes component-specific notices beyond the headline Apache license

Evidence boundaries

A launch shape is not a production SLA.

The minimum tier records the smallest official or inference-framework configuration we found. The 200-person and commercial API tiers still require measurements against real prompt length, output length, concurrency, latency and redundancy targets.

Read the complete index methodology

  1. AChinaAPI reproduced
  2. Binference-framework official validated recipe
  3. Cmodel-vendor documented configuration
  4. Dthird-party reproduction
  5. Ecapacity estimate only

Compare before deploying

Review every verified model or compare hosted access.

The directory keeps model selection separate from the evidence and capacity details on this page.