Chinese speech API directory

Chinese speech APIs for ASR and text to speech.

Compare speech recognition and speech synthesis models available through one ChinaAPI key. Each model keeps its real request path. stepaudio-2.5-tts, glm-tts, qwen3-tts-flash, speech-2.8-hd, speech-2.8-turbo, step-tts-2, step-tts-mini uses /v1/audio/speech. stepaudio-2-asr-pro, step-asr, step-asr-1.1, qwen3-asr-flash, stepaudio-2.5-asr use /v1/audio/transcriptions. mimo-v2.5-asr, mimo-v2.5-tts use /v1/chat/completions with model-specific audio payloads.

Live catalog

14 speech models, generated from one data source.

Model IDs, endpoint modes, billing units, capabilities, and displayed prices come from scripts/models-data.json, whose audio prices are reconciled with the ChinaAPI gateway. The displayed media rate may include a service margin covering provider input/output billing, payment processing, chargeback exposure, and operations. Any margin is included in the displayed rate and is not added separately. Check live pricing before production rollout.

stepaudio-2-asr-pro

Speech recognition (ASR) · StepFun

Endpoint: POST /v1/audio/transcriptions

Request mode: OpenAI-compatible transcription endpoint

Billing: $0.35 per audio hour

AudioSpeech-to-Text

step-asr

Speech recognition (ASR) · StepFun

Endpoint: POST /v1/audio/transcriptions

Request mode: OpenAI-compatible transcription endpoint

Billing: $0.15 per audio hour

AudioSpeech-to-TextChinese Dialects

step-asr-1.1

Speech recognition (ASR) · StepFun

Endpoint: POST /v1/audio/transcriptions

Request mode: OpenAI-compatible transcription endpoint

Billing: $0.35 per audio hour

AudioSpeech-to-TextChinese Dialects

mimo-v2.5-asr

Speech recognition (ASR) · Xiaomi

Endpoint: POST /v1/chat/completions

Request mode: Chat Completions with an input_audio payload

Billing: $0.074 per audio hour

AudioSpeech-to-TextMultilingualChinese DialectsNoise Robustness

mimo-v2.5-tts

Text to speech (TTS) · Xiaomi

Endpoint: POST /v1/chat/completions

Request mode: Chat Completions with messages plus an audio output configuration

Billing: $0 per 10k characters

Limited-time free upstream pricing; check live pricing before production rollout.

AudioText-to-SpeechMultilingualStyle ControlBuilt-in Voices

stepaudio-2.5-tts

Text to speech (TTS) · StepFun

Endpoint: POST /v1/audio/speech

Request mode: OpenAI-compatible speech endpoint

Billing: $0.85 per 10k characters

AudioText-to-SpeechContextual TTSStyle ControlBuilt-in Voices

glm-tts

Text to speech (TTS) · Zhipu AI

Endpoint: POST /v1/audio/speech

Request mode: OpenAI-compatible speech endpoint

Billing: $0.3077 per 10k characters

AudioText-to-SpeechStreamingEmotionVoice Control

qwen3-asr-flash

Speech recognition (ASR) · Alibaba

Endpoint: POST /v1/audio/transcriptions

Request mode: OpenAI-compatible transcription endpoint

Billing: $0.1235 per audio hour

AudioSpeech-to-TextMultilingualEmotion RecognitionITN

qwen3-tts-flash

Text to speech (TTS) · Alibaba

Endpoint: POST /v1/audio/speech

Request mode: OpenAI-compatible speech endpoint

Billing: $0.1231 per 10k characters

AudioText-to-SpeechMultilingualBuilt-in VoicesStreaming

speech-2.8-hd

Text to speech (TTS) · MiniMax

Endpoint: POST /v1/audio/speech

Request mode: OpenAI-compatible speech endpoint

Billing: $0.5385 per 10k characters

AudioText-to-SpeechHigh FidelityEmotionMultilingual

speech-2.8-turbo

Text to speech (TTS) · MiniMax

Endpoint: POST /v1/audio/speech

Request mode: OpenAI-compatible speech endpoint

Billing: $0.3077 per 10k characters

AudioText-to-SpeechLow LatencyEmotionMultilingual

step-tts-2

Text to speech (TTS) · StepFun

Endpoint: POST /v1/audio/speech

Request mode: OpenAI-compatible speech endpoint

Billing: $0.4 per 10k characters

AudioText-to-SpeechVoice CloningEmotionPronunciation Control

step-tts-mini

Text to speech (TTS) · StepFun

Endpoint: POST /v1/audio/speech

Request mode: OpenAI-compatible speech endpoint

Billing: $0.15 per 10k characters

AudioText-to-SpeechMultilingualVoice CloningStyle Control

stepaudio-2.5-asr

Speech recognition (ASR) · StepFun

Endpoint: POST /v1/audio/transcriptions

Request mode: OpenAI-compatible transcription endpoint

Billing: $0.022 per audio hour

AudioSpeech-to-TextStreamingChinese and EnglishITN

Choose by workload

ASR for audio understanding; TTS for voice output.

Transcription and audio understanding

stepaudio-2-asr-pro: StepFun StepAudio 2 ASR Pro — 32B-parameter speech recognition, the accuracy-first option in the StepFun ASR line where stepaudio-2.5-asr is the speed-and-cost option. Accepts ogg, mp3, and wav uploads; transcript output is not billed.

Transcription and audio understanding

step-asr: StepFun step-asr — general-purpose speech recognition covering Chinese, English, and Chinese dialects, for transcription, meeting minutes, call quality review, and voice search. Accepts ogg, mp3, and wav uploads; transcript output is not billed.

Transcription and audio understanding

step-asr-1.1: StepFun step-asr-1.1 — the updated general-purpose recognizer in the step-asr line, covering Chinese, English, and Chinese dialects. Accepts ogg, mp3, and wav uploads; transcript output is not billed.

Transcription and audio understanding

mimo-v2.5-asr: Xiaomi MiMo-V2.5-ASR — speech recognition for Chinese, English, code-switching, regional dialects, noisy recordings, far-field audio, multiple speakers, and lyrics.

Controllable voice output

mimo-v2.5-tts: Xiaomi MiMo-V2.5-TTS — expressive multilingual speech synthesis with built-in voices, natural-language style instructions, emotion, pace, tone, dialect, and character control.

Speech output through a dedicated route

stepaudio-2.5-tts: StepAudio 2.5 TTS — contextual speech synthesis with inline natural-language direction, expressive delivery, built-in voices, and OpenAI-compatible audio output.

Speech output through a dedicated route

glm-tts: GLM-TTS — context-aware speech synthesis with expressive delivery, streaming output, fast first audio, and controls for voice, speed, and volume.

Transcription and audio understanding

qwen3-asr-flash: Qwen3-ASR-Flash — multilingual file transcription with language hints, inverse text normalization, emotion recognition, and OpenAI-compatible audio input.

Speech output through a dedicated route

qwen3-tts-flash: Qwen3-TTS-Flash — low-latency multilingual speech synthesis with mixed-language input, built-in voices, automatic language matching, and character-based billing.

Speech output through a dedicated route

speech-2.8-hd: MiniMax Speech 2.8 HD — high-fidelity speech synthesis with natural emotional delivery, language enhancement, voice controls, and production audio formats.

Speech output through a dedicated route

speech-2.8-turbo: MiniMax Speech 2.8 Turbo — speed-focused speech synthesis with natural output, emotion controls, language enhancement, and common production audio formats.

Speech output through a dedicated route

step-tts-2: Step TTS 2 — expressive speech synthesis with official and cloned voices, speed and volume control, pronunciation mapping, and multiple output formats.

Speech output through a dedicated route

step-tts-mini: Step TTS Mini — cost-efficient expressive speech synthesis for Chinese, English, Japanese, Cantonese, and Sichuanese, with official and cloned voices.

Transcription and audio understanding

stepaudio-2.5-asr: StepAudio 2.5 ASR — a 4B MTP streaming transcription model for fast Chinese-English recognition, ITN normalization, subtitles, meetings, and voice agents.

Compatibility boundary

OpenAI-compatible does not mean one universal audio route.

The route names and request envelopes follow OpenAI-compatible patterns, but clients must send the audio fields required by each model. stepaudio-2.5-tts, glm-tts, qwen3-tts-flash, speech-2.8-hd, speech-2.8-turbo, step-tts-2, step-tts-mini uses /v1/audio/speech. stepaudio-2-asr-pro, step-asr, step-asr-1.1, qwen3-asr-flash, stepaudio-2.5-asr use /v1/audio/transcriptions. mimo-v2.5-asr, mimo-v2.5-tts use /v1/chat/completions with model-specific audio payloads.

Quickstarts

Copy the request for the model and transport you need.

stepaudio-2-asr-pro

Speech recognition (ASR) via /v1/audio/transcriptions

curl -X POST https://api.chinaapi.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $CHINAAPI_API_KEY" \
  -F "model=stepaudio-2-asr-pro" \
  -F "[email protected]" \
  -F "response_format=json"

step-asr

Speech recognition (ASR) via /v1/audio/transcriptions

curl -X POST https://api.chinaapi.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $CHINAAPI_API_KEY" \
  -F "model=step-asr" \
  -F "[email protected]" \
  -F "response_format=json"

step-asr-1.1

Speech recognition (ASR) via /v1/audio/transcriptions

curl -X POST https://api.chinaapi.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $CHINAAPI_API_KEY" \
  -F "model=step-asr-1.1" \
  -F "[email protected]" \
  -F "response_format=json"

mimo-v2.5-asr

Speech recognition (ASR) via /v1/chat/completions

AUDIO_BASE64=$(base64 < meeting.wav | tr -d '\n')

curl -X POST https://api.chinaapi.ai/v1/chat/completions \
  -H "Authorization: Bearer $CHINAAPI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "mimo-v2.5-asr",
    "messages": [{
      "role": "user",
      "content": [{
        "type": "input_audio",
        "input_audio": {"data": "data:audio/wav;base64,'"$AUDIO_BASE64"'"}
      }]
    }]
  }'

mimo-v2.5-tts

Text to speech (TTS) via /v1/chat/completions

curl -X POST https://api.chinaapi.ai/v1/chat/completions \
  -H "Authorization: Bearer $CHINAAPI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "mimo-v2.5-tts",
    "messages": [
      {"role": "user", "content": "Natural, clear, and friendly delivery"},
      {"role": "assistant", "content": "Welcome to ChinaAPI. 欢迎使用语音模型。"}
    ],
    "audio": {"format": "wav", "voice": "冰糖"}
  }' > response.json

stepaudio-2.5-tts

Text to speech (TTS) via /v1/audio/speech

curl -X POST https://api.chinaapi.ai/v1/audio/speech \
  -H "Authorization: Bearer $CHINAAPI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "stepaudio-2.5-tts",
    "voice": "cixingnansheng",
    "input": "(轻声)Welcome to ChinaAPI. 今天我们来测试自然语音。",
    "response_format": "mp3"
  }' \
  --output speech.mp3

glm-tts

Text to speech (TTS) via /v1/audio/speech

curl -X POST https://api.chinaapi.ai/v1/audio/speech \
  -H "Authorization: Bearer $CHINAAPI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "glm-tts",
    "voice": "tongtong",
    "input": "(轻声)Welcome to ChinaAPI. 今天我们来测试自然语音。",
    "response_format": "mp3"
  }' \
  --output speech.mp3

qwen3-asr-flash

Speech recognition (ASR) via /v1/audio/transcriptions

curl -X POST https://api.chinaapi.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $CHINAAPI_API_KEY" \
  -F "model=qwen3-asr-flash" \
  -F "[email protected]" \
  -F "response_format=json"

qwen3-tts-flash

Text to speech (TTS) via /v1/audio/speech

curl -X POST https://api.chinaapi.ai/v1/audio/speech \
  -H "Authorization: Bearer $CHINAAPI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3-tts-flash",
    "voice": "Cherry",
    "input": "(轻声)Welcome to ChinaAPI. 今天我们来测试自然语音。",
    "response_format": "mp3"
  }' \
  --output speech.mp3

speech-2.8-hd

Text to speech (TTS) via /v1/audio/speech

curl -X POST https://api.chinaapi.ai/v1/audio/speech \
  -H "Authorization: Bearer $CHINAAPI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "speech-2.8-hd",
    "voice": "Chinese (Mandarin)_Warm_Bestie",
    "input": "(轻声)Welcome to ChinaAPI. 今天我们来测试自然语音。",
    "response_format": "mp3"
  }' \
  --output speech.mp3

speech-2.8-turbo

Text to speech (TTS) via /v1/audio/speech

curl -X POST https://api.chinaapi.ai/v1/audio/speech \
  -H "Authorization: Bearer $CHINAAPI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "speech-2.8-turbo",
    "voice": "Chinese (Mandarin)_Warm_Bestie",
    "input": "(轻声)Welcome to ChinaAPI. 今天我们来测试自然语音。",
    "response_format": "mp3"
  }' \
  --output speech.mp3

step-tts-2

Text to speech (TTS) via /v1/audio/speech

curl -X POST https://api.chinaapi.ai/v1/audio/speech \
  -H "Authorization: Bearer $CHINAAPI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "step-tts-2",
    "voice": "cixingnansheng",
    "input": "(轻声)Welcome to ChinaAPI. 今天我们来测试自然语音。",
    "response_format": "mp3"
  }' \
  --output speech.mp3

step-tts-mini

Text to speech (TTS) via /v1/audio/speech

curl -X POST https://api.chinaapi.ai/v1/audio/speech \
  -H "Authorization: Bearer $CHINAAPI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "step-tts-mini",
    "voice": "cixingnansheng",
    "input": "(轻声)Welcome to ChinaAPI. 今天我们来测试自然语音。",
    "response_format": "mp3"
  }' \
  --output speech.mp3

stepaudio-2.5-asr

Speech recognition (ASR) via /v1/audio/transcriptions

curl -X POST https://api.chinaapi.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $CHINAAPI_API_KEY" \
  -F "model=stepaudio-2.5-asr" \
  -F "[email protected]" \
  -F "response_format=json"

Do all ChinaAPI speech models use the same audio endpoint?

No. stepaudio-2.5-tts, glm-tts, qwen3-tts-flash, speech-2.8-hd, speech-2.8-turbo, step-tts-2, step-tts-mini uses /v1/audio/speech. stepaudio-2-asr-pro, step-asr, step-asr-1.1, qwen3-asr-flash, stepaudio-2.5-asr use /v1/audio/transcriptions. mimo-v2.5-asr, mimo-v2.5-tts use /v1/chat/completions with model-specific audio payloads.

How are the speech models billed?

stepaudio-2-asr-pro: $0.35 per audio hour; step-asr: $0.15 per audio hour; step-asr-1.1: $0.35 per audio hour; mimo-v2.5-asr: $0.074 per audio hour; mimo-v2.5-tts: $0 per 10k characters; stepaudio-2.5-tts: $0.85 per 10k characters; glm-tts: $0.3077 per 10k characters; qwen3-asr-flash: $0.1235 per audio hour; qwen3-tts-flash: $0.1231 per 10k characters; speech-2.8-hd: $0.5385 per 10k characters; speech-2.8-turbo: $0.3077 per 10k characters; step-tts-2: $0.4 per 10k characters; step-tts-mini: $0.15 per 10k characters; stepaudio-2.5-asr: $0.022 per audio hour. The displayed media rate may include a service margin covering provider input/output billing, payment processing, chargeback exposure, and operations. Any margin is included in the displayed rate and is not added separately. These values are rendered from the model catalog rather than copied into this page.

Can I use these models without a mainland-China account?

Yes. ChinaAPI provides access without a mainland-China account or phone number. Register for a key, then use the exact route and payload shown for your chosen model.