Inputs
Text, plus a built-in voice or explicitly supported voice design/clone controls.
audio-speech capability center
Compare TTS models by endpoint, voice controls, billing unit, output workflow, and an executable speech request.
Contract
Text, plus a built-in voice or explicitly supported voice design/clone controls.
Speech audio in a model-supported format, or an audio field in the chat response.
POST /v1/audio/speech or POST /v1/chat/completions, according to api_mode
Synchronous. Save the binary audio response or decode the audio field returned by chat mode.
Supported tasks
text-to-speech, voice-cloning, voice-design
Limit: Voice IDs, cloning consent, language coverage, streaming, character limits, and output formats are model-specific.
Matching catalog
Unknown capabilities are omitted, never guessed from a model name. Pricing examples use the synchronized public data; live Dash Pricing remains authoritative.
Xiaomi MiMo-V2.5-TTS — expressive multilingual speech synthesis with built-in voices, natural-language style instructions, emotion, pace, tone, dialect, and character control.
Tasks: video-with-audio, text-to-speech
Endpoint: POST /v1/chat/completions
Cost example: 1 10k characters × $0 = $0.
StepAudio 2.5 TTS — contextual speech synthesis with inline natural-language direction, expressive delivery, built-in voices, and OpenAI-compatible audio output.
Tasks: video-with-audio, text-to-speech
Endpoint: POST /v1/audio/speech
Cost example: 1 10k characters × $0.85 = $0.85.
GLM-TTS — context-aware speech synthesis with expressive delivery, streaming output, fast first audio, and controls for voice, speed, and volume.
Tasks: video-with-audio, text-to-speech
Endpoint: POST /v1/audio/speech
Cost example: 1 10k characters × $0.307692 = $0.307692.
Qwen3-TTS-Flash — low-latency multilingual speech synthesis with mixed-language input, built-in voices, automatic language matching, and character-based billing.
Tasks: video-with-audio, text-to-speech
Endpoint: POST /v1/audio/speech
Cost example: 1 10k characters × $0.123077 = $0.123077.
MiniMax Speech 2.8 HD — high-fidelity speech synthesis with natural emotional delivery, language enhancement, voice controls, and production audio formats.
Tasks: video-with-audio, text-to-speech
Endpoint: POST /v1/audio/speech
Cost example: 1 10k characters × $0.538462 = $0.538462.
MiniMax Speech 2.8 Turbo — speed-focused speech synthesis with natural output, emotion controls, language enhancement, and common production audio formats.
Tasks: video-with-audio, text-to-speech
Endpoint: POST /v1/audio/speech
Cost example: 1 10k characters × $0.307692 = $0.307692.
Step TTS 2 — expressive speech synthesis with official and cloned voices, speed and volume control, pronunciation mapping, and multiple output formats.
Tasks: video-with-audio, text-to-speech
Endpoint: POST /v1/audio/speech
Cost example: 1 10k characters × $0.4 = $0.4.
Step TTS Mini — cost-efficient expressive speech synthesis for Chinese, English, Japanese, Cantonese, and Sichuanese, with official and cloned voices.
Tasks: video-with-audio, text-to-speech
Endpoint: POST /v1/audio/speech
Cost example: 1 10k characters × $0.15 = $0.15.
Calculate API cost · Generate curl, Python, or JavaScript · Install the Seedance 2 Agent Recipe
Source and method: public model pricing dataset, updated 2026-08-08. Limitations are stated above; third-party claims remain next to the model pages that cite them.