Released 2026-05-28 · audio video generation
Deployment evidence
Baidu ERNIE Team · 6.3B backbone total · 6.3B active · Not applicable
Best-fit scenarios
Synchronized audio-video generation with multi-speaker timbre control, camera prompting and image continuation.
Not recommended
- Unconsented face or voice cloning
- Low-latency interactive video generation
Synchronized short-form audio-video
Vendor-statedNAVA jointly generates video, scene audio and speech rather than aligning separate outputs after generation.
Capability sourceMulti-speaker and reference-timbre scenes
Vendor-statedThe official checkpoint supports up to two reference voices bound to speech spans.
Capability source720p creative generation
Public evidenceThe vendor reports VerseBench synchronization and quality results plus an 8-GPU fast path.
Capability sourceOfficial launch / minimum
Evidence CSingle gpu supported no vram
The model card supports single-GPU inference but does not publish exact minimum VRAM. Primary recipe
200-person team
Evidence EMedia queue required
Size a queued render farm by jobs per hour, resolution and duration; office-seat assumptions do not apply.
Commercial API
Evidence COfficial 8gpu reference
The official Ulysses SP8 path reports roughly one minute for a 720p synchronized clip; HA needs additional workers. Primary recipe
Commercial-use check
Apache-2.0
MaaS / hosted service: The model card states Apache-2.0, but bundled LTX audio-VAE artifacts carry an additional community license that must be reviewed for the shipped stack.
Attribution: Preserve Apache notices and the notices/licenses for bundled upstream components.
Prohibited uses: The model card prohibits depicting real persons without consent, including face or voice likeness reproduction.
Known limitations and open questions
- Default clips are about 6–10 seconds
- The full dependency stack includes component-specific notices beyond the headline Apache license
