Alibaba Cloud released qwen3.8-max on 2026-08-03 as Qwen's most capable flagship model to date. The native vision-language model uses a 2.4T-parameter Mixture-of-Experts architecture and supports reasoning, text generation, and visual understanding, with significant improvements over the Qwen3.7 series. Official Model Studio documentation lists OpenAI-compatible, Anthropic-compatible, and DashScope access across Beijing and multiple global regions. The route is now available through ChinaAPI.
Model update wall
China AI model release timeline.
A data-driven wall of Chinese AI model releases. Dates come from the linked vendor sources; ChinaAPI availability is shown on each entry as a secondary status. 3 estimated dates and 5 unknown dates are flagged in the source data.
Newest first
Release wall.
MiniMax released H3 (Hailuo 3.0). The official release notes describe it as a new-generation open general-purpose multimodal video model that understands creative intent across text, image, video, and audio context and delivers more natural, coherent generation. ChinaAPI serves the route for 4-15 second clips at 768P and 2K with native stereo audio in a single pass, billed per second of output and callable through the OpenAI-compatible video endpoint.
Volcano Engine's Ark model catalog lists Doubao-Seedance-2.5 under model ID doubao-seedance-2-5-260628. Released on 2026-07-31, the video model can generate up to 30 seconds, accept up to 50 reference assets (30 images, 10 videos, and 10 audio clips), and support multimodal reference, video editing, extension, and generation in more than 10 languages. The 260628 suffix remains the model version tag. ChinaAPI has not yet verified this route for customer calls, so no model page or pricing is published yet.
DeepSeek upgraded its existing deepseek-v4-flash API route with a re-post-trained checkpoint that keeps the preview architecture and model size while strengthening coding and agent capabilities. DeepSeek reports 82.7 on Terminal Bench 2.1, 54.4 on DeepSWE, and 70.3 on Toolathlon Verified using max reasoning effort and an unreleased minimal harness; these are vendor results, not ChinaAPI tests. The official route now natively supports the Responses API and Codex integration, while DeepSeek's web app, mobile app, and V4-Pro API were not part of this update.
Alibaba Cloud released its native vision-language Flash model with 1M context and up to 131K output, upgrading multimodal understanding, real-world and spatial perception, Search Agent and CI Agent execution, and multimodal coding over Qwen3.6-Flash. The official model accepts text, image, and video and exposes built-in tools through the Responses API. ChinaAPI's gateway does not currently list this route.
Moonshot released its 2.8T-parameter flagship with 1M context, always-on reasoning, and native image and video understanding. The announced open-source weights are due by July 27, 2026.
Tencent officially released Hy3, a 256K-context hybrid fast-and-slow-thinking model with improved stability and cost efficiency.
Meituan released its open MoE reasoning model with 1M context for coding agents and long-horizon tool use.
Lower-latency Doubao Seed 2.1 model for multimodal reasoning, coding, and production agent workflows.
Flagship Doubao Seed 2.1 model for multimodal reasoning, coding, tool use, and long-horizon agent tasks.
HappyHorse 1.1 text-to-video gateway model with improved semantic understanding and camera control.
HappyHorse 1.1 reference-to-video gateway model for multi-reference subject, scene, and style consistency.
HappyHorse 1.1 image-to-video gateway model with stronger texture and identity consistency.
Open flagship GLM with 1M context and stronger long-horizon coding and agent workflows.
Fast Seedance 2.0 mini video model for lower-cost short-form generation.
High-speed Kimi K2.7 gateway variant for lower-latency coding and agent tasks.
Coding-focused Kimi gateway model for long-horizon repository work with vision support.
MiMo speech recognition for Chinese-English code-switching, Cantonese, Wu, Minnan and Sichuan dialects, noisy and far-field recordings, overlapping speakers, and lyrics.
Multimodal Qwen gateway model for visual analysis, planning, files, and tool use.
Native multimodal M-series model with 1M context, MSA attention, coding, agent, image, and video input support.
Fast StepFun gateway model with 256K context, vision input, reasoning, and tool use.
Audio-driven avatar video technical release focused on lip sync, identity consistency, and long-video generation.
Flagship Qwen gateway model for demanding analysis, coding, and production agent workflows.
Qwen 3.6 fast vision-language model for visual reasoning, files, and agent workflows.
1.02T-parameter MiMo agent model with 1M context for demanding coding and long-horizon tasks.
Native omnimodal MiMo model with 1M context and text, image, video, and audio understanding.
Wan 2.7 image-to-video with first/last-frame control, continuation, driving audio, and 720P/1080P output.
Flagship DeepSeek V4 variant with 1.6T total parameters and 1M-token context.
Fast DeepSeek V4 variant for efficient chat, coding assistance, and agent loops.
MiMo expressive text-to-speech with premium built-in voices and natural-language control over pace, emotion, tone, dialect, and character performance.
Open native multimodal Kimi model for long-horizon coding, visual design, and agent-swarm workflows.
A 200K-context multimodal coding foundation model for image, video, file, GUI-agent, and tool-driven workflows.
Wan 2.7 natural-language video editing model for local/global edits and element replacement.
Wan 2.7 text-to-video model for prompt-driven video generation.
Wan 2.7 reference-to-video model for controlled generation from visual references.
GLM update for longer-horizon tasks, coding agents, and hybrid reasoning.
Wan 2.7 Pro image generation and editing with brand-color control, multi-reference consistency, and up to 4096×4096 output.
Wan 2.7 image model for text-to-image, editing, multi-reference generation, and stronger text rendering.
High-speed MiniMax variant for low-latency coding and agent workflows.
Open-weight MiniMax model for chat, coding, office work, and agentic tasks.
Native multimodal-agent Qwen release that preceded the served Qwen 3.6 and 3.7 gateway models.
Efficient GLM-5 turbo variant for faster reasoning, coding, and agent work.
GLM-5 release for coding, reasoning, and agent workflows.
Kling 3.0 Omni variant for reference-to-video and editing workflows.
Kling 3.0 generation with text-to-video, image-to-video, audio, and 720p to 4K output tiers.
Fast Kling 3.0 Turbo path for lower-latency video generation with audio.
Open StepFun flash model with 256K context and tool-use support.
Seedream 5.0-lite image model with controllable creation, retrieval, and stronger subject consistency.
Fast Seedance 2.0 video generation option for cost-efficient clips.
Seedance 2.0 flagship video generation model billed by output resolution.
MiniMax released Speech 2.8 with native sound tags, improved voice cloning, studio-grade audio, and cross-lingual speech; Turbo is its speed-focused option.
MiniMax released Speech 2.8 with native sound tags, improved voice cloning, studio-grade audio, and cross-lingual speech; HD is its high-fidelity option.
Alibaba added OpenAI-compatible mode to Qwen3-ASR-Flash, retaining multilingual transcription, ITN, and streaming support.
Open native multimodal Kimi agent model with 256K context and vision-language training.
Z.AI released its multilingual ASR model with improved custom-dictionary support and specialised-terminology recognition.
Small image-generation model focused on data quality over model size.
Fast Hailuo 2.3 path for lower-cost image-to-video generation.
Hailuo 2.3 improves complex motion, physical actions, stylization, and micro-expression performance.
Alibaba released Qwen3-TTS-Flash for offline speech synthesis with expressive voices, multilingual and dialect output, and low-latency API access.
Budget Hailuo 02 video model for text-to-video and image-to-video generation from 512p.
Context-aware text-to-speech with global and inline natural-language direction, expressive delivery, and direct OpenAI-compatible /v1/audio/speech output.
Fast 4B MTP streaming transcription for Chinese-English audio with ITN normalization and an OpenAI-compatible transcription route.
Cost-efficient multilingual speech synthesis with official and cloned voices, pronunciation mapping, and style controls.
Expressive speech synthesis with official and cloned voices, pronunciation mapping, style controls, and OpenAI-compatible output.
Context-aware expressive speech synthesis with streaming output and controls for voice, speed, and volume.
