Type
Speech recognition (ASR)
StepFun API model
StepFun step-asr — general-purpose speech recognition covering Chinese, English, and Chinese dialects, for transcription, meeting minutes, call quality review, and voice search. Accepts ogg, mp3, and wav uploads; transcript output is not billed.
Specs
Transparent USD pricing. No mainland-China account or phone required. The displayed media rate may include a service margin covering provider input/output billing, payment processing, chargeback exposure, and operations. Any margin is included in the displayed rate and is not added separately.
Speech recognition (ASR)
Audio input, text output
Audio, Speech-to-Text, Chinese Dialects
$0.15 per audio hour
/v1/audio/transcriptions
StepFun step-asr — general-purpose speech recognition covering Chinese, English, and Chinese dialects, for transcription, meeting minutes, call quality review, and voice search. Accepts ogg, mp3, and wav uploads; transcript output is not billed.
Pricing
$0.15 per audio hour. Paid in USD. The displayed media rate may include a service margin covering provider input/output billing, payment processing, chargeback exposure, and operations. Any margin is included in the displayed rate and is not added separately. View live pricing
Quickstart
curl -X POST https://api.chinaapi.ai/v1/audio/transcriptions \
-H "Authorization: Bearer $CHINAAPI_API_KEY" \
-F "model=step-asr" \
-F "[email protected]" \
-F "response_format=json"
Internal links
Use step-asr through ChinaAPI with an API key and the OpenAI-compatible endpoint shown above.
step-asr is $0.15 per audio hour, paid in USD. The live pricing page is authoritative. The displayed media rate may include a service margin covering provider input/output billing, payment processing, chargeback exposure, and operations. Any margin is included in the displayed rate and is not added separately.
No. ChinaAPI provides access without a mainland-China account or phone number and includes a $2 free trial. Check live pricing for the displayed USD rate.