Model update wall

China AI model release timeline.

A data-driven wall of Chinese AI model releases. Dates come from the linked vendor sources; ChinaAPI availability is shown on each entry as a secondary status. 3 estimated dates and 5 unknown dates are flagged in the source data.

Newest first

Release wall.

2026-08-03

Alibaba Cloud released qwen3.8-max on 2026-08-03 as Qwen's most capable flagship model to date. The native vision-language model uses a 2.4T-parameter Mixture-of-Experts architecture and supports reasoning, text generation, and visual understanding, with significant improvements over the Qwen3.7 series. Official Model Studio documentation lists OpenAI-compatible, Anthropic-compatible, and DashScope access across Beijing and multiple global regions. The route is now available through ChinaAPI.

LLM ✓ On ChinaAPI confirmed Source · help.aliyun.com
2026-07-31

MiniMax released H3 (Hailuo 3.0). The official release notes describe it as a new-generation open general-purpose multimodal video model that understands creative intent across text, image, video, and audio context and delivers more natural, coherent generation. ChinaAPI serves the route for 4-15 second clips at 768P and 2K with native stereo audio in a single pass, billed per second of output and callable through the OpenAI-compatible video endpoint.

Video ✓ On ChinaAPI confirmed Source · platform.minimax.io
2026-07-31
ByteDance Doubao-Seedance-2.5 Seedance 2.5

Volcano Engine's Ark model catalog lists Doubao-Seedance-2.5 under model ID doubao-seedance-2-5-260628. Released on 2026-07-31, the video model can generate up to 30 seconds, accept up to 50 reference assets (30 images, 10 videos, and 10 audio clips), and support multimodal reference, video editing, extension, and generation in more than 10 languages. The 260628 suffix remains the model version tag. ChinaAPI has not yet verified this route for customer calls, so no model page or pricing is published yet.

Video Not yet confirmed Source · console.volcengine.com
2026-07-31

DeepSeek upgraded its existing deepseek-v4-flash API route with a re-post-trained checkpoint that keeps the preview architecture and model size while strengthening coding and agent capabilities. DeepSeek reports 82.7 on Terminal Bench 2.1, 54.4 on DeepSWE, and 70.3 on Toolathlon Verified using max reasoning effort and an unreleased minimal harness; these are vendor results, not ChinaAPI tests. The official route now natively supports the Responses API and Codex integration, while DeepSeek's web app, mobile app, and V4-Pro API were not part of this update.

LLM ✓ On ChinaAPI confirmed Source · api-docs.deepseek.com
2026-07-25
Alibaba Cloud Qwen3.7-Flash Qwen3.7

Alibaba Cloud released its native vision-language Flash model with 1M context and up to 131K output, upgrading multimodal understanding, real-world and spatial perception, Search Agent and CI Agent execution, and multimodal coding over Qwen3.6-Flash. The official model accepts text, image, and video and exposes built-in tools through the Responses API. ChinaAPI's gateway does not currently list this route.

LLM Not yet confirmed Source · cn.aliyun.com
2026-07-16

Moonshot released its 2.8T-parameter flagship with 1M context, always-on reasoning, and native image and video understanding. The announced open-source weights are due by July 27, 2026.

LLM ✓ On ChinaAPI confirmed Source · kimi.com
2026-07-06
Tencent Hy3 Hunyuan

Tencent officially released Hy3, a 256K-context hybrid fast-and-slow-thinking model with improved stability and cost efficiency.

LLM ✓ On ChinaAPI confirmed Source · tencent.com
2026-06-22

HappyHorse 1.1 text-to-video gateway model with improved semantic understanding and camera control.

Video ✓ On ChinaAPI confirmed
2026-06-22

HappyHorse 1.1 reference-to-video gateway model for multi-reference subject, scene, and style consistency.

Video ✓ On ChinaAPI confirmed
2026-06-22

HappyHorse 1.1 image-to-video gateway model with stronger texture and identity consistency.

Video ✓ On ChinaAPI confirmed
2026-06-12

Coding-focused Kimi gateway model for long-horizon repository work with vision support.

LLM ✓ On ChinaAPI confirmed
2026-06-02

MiMo speech recognition for Chinese-English code-switching, Cantonese, Wu, Minnan and Sichuan dialects, noisy and far-field recordings, overlapping speakers, and lyrics.

Audio ✓ On ChinaAPI confirmed Source · mimo.mi.com
2026-06-01

Multimodal Qwen gateway model for visual analysis, planning, files, and tool use.

LLM ✓ On ChinaAPI confirmed
2026-06-01

Native multimodal M-series model with 1M context, MSA attention, coding, agent, image, and video input support.

LLM ✓ On ChinaAPI confirmed Source · minimax.io
2026-05-28

Fast StepFun gateway model with 256K context, vision input, reasoning, and tool use.

LLM ✓ On ChinaAPI confirmed
2026-05-26
Meituan LongCat-Video-Avatar 1.5 LongCat Video

Audio-driven avatar video technical release focused on lip sync, identity consistency, and long-video generation.

Video Not yet confirmed Source · arxiv.org
2026-05-21

Flagship Qwen gateway model for demanding analysis, coding, and production agent workflows.

LLM ✓ On ChinaAPI confirmed
2026-04-23

MiMo expressive text-to-speech with premium built-in voices and natural-language control over pace, emotion, tone, dialect, and character performance.

Audio ✓ On ChinaAPI confirmed Source · mimo.mi.com
2026-04-07
Zhipu AI GLM-5V-Turbo GLM

A 200K-context multimodal coding foundation model for image, video, file, GUI-agent, and tool-driven workflows.

LLM Not yet confirmed Source · docs.bigmodel.cn
2026-04-03

Wan 2.7 natural-language video editing model for local/global edits and element replacement.

Video ✓ On ChinaAPI confirmed
2026-04-03

Wan 2.7 text-to-video model for prompt-driven video generation.

Video ✓ On ChinaAPI confirmed
2026-04-03

Wan 2.7 reference-to-video model for controlled generation from visual references.

Video ✓ On ChinaAPI confirmed
2026-04-01

Wan 2.7 image model for text-to-image, editing, multi-reference generation, and stronger text rendering.

Image ✓ On ChinaAPI confirmed
2026-02-16
Alibaba Qwen3.5 Qwen

Native multimodal-agent Qwen release that preceded the served Qwen 3.6 and 3.7 gateway models.

LLM Not yet confirmed Source · en.wikipedia.org
2026-02-01
StepFun Step-3.5-Flash Step

Open StepFun flash model with 256K context and tool-use support.

LLM Not yet confirmed Source · huggingface.co
2026-01-28

Seedream 5.0-lite image model with controllable creation, retrieval, and stronger subject consistency.

Image ✓ On ChinaAPI confirmed
2026-01-23

MiniMax released Speech 2.8 with native sound tags, improved voice cloning, studio-grade audio, and cross-lingual speech; Turbo is its speed-focused option.

Audio ✓ On ChinaAPI confirmed Source · minimax.io
2026-01-23

MiniMax released Speech 2.8 with native sound tags, improved voice cloning, studio-grade audio, and cross-lingual speech; HD is its high-fidelity option.

Audio ✓ On ChinaAPI confirmed Source · minimax.io
2026-01-01
Moonshot kimi-k2.5 Kimi

Open native multimodal Kimi agent model with 256K context and vision-language training.

LLM Not yet confirmed Source · huggingface.co
2025-12-10
Zhipu AI GLM-ASR-2512 GLM Audio

Z.AI released its multilingual ASR model with improved custom-dictionary support and specialised-terminology recognition.

Audio Not yet confirmed Source · docs.z.ai
2025-12
Meituan LongCat-Image LongCat

Small image-generation model focused on data quality over model size.

Image Not yet confirmed Source · en.wikipedia.org
2025-09-22

Alibaba released Qwen3-TTS-Flash for offline speech synthesis with expressive voices, multilingual and dialect output, and low-latency API access.

Audio ✓ On ChinaAPI confirmed Source · help.aliyun.com
Release date not publicly confirmed

Context-aware text-to-speech with global and inline natural-language direction, expressive delivery, and direct OpenAI-compatible /v1/audio/speech output.

Audio ✓ On ChinaAPI unknown Source · platform.stepfun.com
Release date not publicly confirmed

Expressive speech synthesis with official and cloned voices, pronunciation mapping, style controls, and OpenAI-compatible output.

Audio ✓ On ChinaAPI unknown Source · platform.stepfun.com
Release date not publicly confirmed

Context-aware expressive speech synthesis with streaming output and controls for voice, speed, and volume.

Audio ✓ On ChinaAPI unknown Source · docs.bigmodel.cn