# NAVA: Self-Hosting Hardware, Scenarios and Commercial License

> Synchronized audio-video generation with multi-speaker timbre control, camera prompting and image continuation.

- Verified: 2026-09-12
- Released: 2026-05-28
- Canonical: https://chinaapi.ai/open-model-deployment/nava/
- Back to directory: https://chinaapi.ai/open-model-deployment/
- Machine-readable dataset: https://chinaapi.ai/data/open-model-deployment.json

## Model facts

- Vendor: Baidu ERNIE Team
- Parameters: 6.3B backbone total; 6.3B active
- Native precision: BF16
- Context: Not applicable
- Category: Joint audio-video generation
- Modalities: text, image, reference voice → video, stereo audio, speech
- Serving paths: PyTorch, Ulysses sequence parallel
- Official weights: https://huggingface.co/baidu/NAVA?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- Model documentation: https://github.com/ernie-research/NAVA?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment

## Deployment evidence

- **Official launch / minimum — evidence C:** The model card supports single-GPU inference but does not publish exact minimum VRAM. Source: https://huggingface.co/baidu/NAVA?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **200-person team — evidence E:** Size a queued render farm by jobs per hour, resolution and duration; office-seat assumptions do not apply.
- **Commercial API — evidence C:** The official Ulysses SP8 path reports roughly one minute for a 720p synchronized clip; HA needs additional workers. Source: https://huggingface.co/baidu/NAVA?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment

## Scenario evidence

- **Synchronized short-form audio-video — Vendor-stated:** NAVA jointly generates video, scene audio and speech rather than aligning separate outputs after generation. Source: https://huggingface.co/baidu/NAVA?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **Multi-speaker and reference-timbre scenes — Vendor-stated:** The official checkpoint supports up to two reference voices bound to speech spans. Source: https://huggingface.co/baidu/NAVA?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- **720p creative generation — Public evidence:** The vendor reports VerseBench synchronization and quality results plus an 8-GPU fast path. Source: https://huggingface.co/baidu/NAVA?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment

## Commercial-use check

- License: Apache-2.0 — https://huggingface.co/baidu/NAVA/blob/main/LICENSE?utm_source=chinaapi&utm_medium=research&utm_campaign=open-model-deployment
- MaaS / hosted service: The model card states Apache-2.0, but bundled LTX audio-VAE artifacts carry an additional community license that must be reviewed for the shipped stack.
- Attribution: Preserve Apache notices and the notices/licenses for bundled upstream components.
- Prohibited uses: The model card prohibits depicting real persons without consent, including face or voice likeness reproduction.

## Avoid or validate first

- Unconsented face or voice cloning
- Low-latency interactive video generation

## Known limitations and open questions

- Default clips are about 6–10 seconds
- The full dependency stack includes component-specific notices beyond the headline Apache license

## Evidence boundary

Loading weights, completing a first forward pass and meeting a production latency/SLA are separate thresholds. The 200-person and commercial tiers require workload-specific measurement. License summaries are product research, not legal advice.

Generated by GENERATED BY scripts/gen_open_model_deployment.py.
