Model update wall

What shipped across Chinese AI.

A concise, source-linked feed of Chinese AI model releases, with ChinaAPI availability at a glance. 3 estimated dates and 5 unknown dates are flagged in the source data.

Newest first

Release wall.

Xiaomi MiMo-V2.6-Pro-UltraSpeed
LLM MiMo V2.6 Source · mimo.mi.com
Xiaomi released a low-latency serving tier that retains MiMo-V2.6-Pro's omni-modal input, flagship reasoning capability, 1M-token context window and 128K output limit. Xiaomi advertises output speeds up to 20 times faster than the standard Pro service; RPM and TPM are supplied as a custom service. The official API model ID is mimo-v2.6-pro-ultraspeed. ChinaAPI lists this route for customer calls; check the live catalog for current pricing.
2026-09-22
✓ On ChinaAPI confirmed
Xiaomi MiMo-V2.6-Pro
LLM MiMo V2.6 Source · mimo.mi.com
Xiaomi released its native omni-modal flagship reasoning model for complex projects, long-horizon work, cybersecurity and research. MiMo-V2.6-Pro accepts text, image, video and audio, returns text, provides a 1M-token context window and up to 128K output, and supports deep thinking, tool calling, streaming, web search, structured output and context caching. The official API model ID is mimo-v2.6-pro. ChinaAPI now offers this route for customer calls; check the live catalog for current pricing.
2026-09-22
✓ On ChinaAPI confirmed
Xiaomi MiMo-V2.6-Flash
LLM MiMo V2.6 Source · mimo.mi.com
Xiaomi released its native omni-modal reasoning model tuned for high-frequency professional work and scaled tasks at lower cost. MiMo-V2.6-Flash accepts text, image, video and audio, returns text, provides a 1M-token context window and up to 128K output, and supports deep thinking, tool calling, streaming, web search, structured output and context caching. The official API model ID is mimo-v2.6-flash. ChinaAPI now offers this route for customer calls; check the live catalog for current pricing.
2026-09-22
✓ On ChinaAPI confirmed
Xiaomi Batch API for MiMo-V2.6
LLM MiMo Batch API Source · mimo.mi.com
Xiaomi launched asynchronous Batch API inference for mimo-v2.6-pro and mimo-v2.6-flash at 50% of its real-time API prices. The service is designed for non-real-time jobs such as evaluation, labeling and bulk generation, accepts JSONL requests, and supports OpenAI Chat Completions, OpenAI Responses and Anthropic Messages request formats without streaming. ChinaAPI offers the two real-time model routes; that does not imply support for Xiaomi's separate Batch API.
2026-09-22
Not on ChinaAPI confirmed
StepFun released Step 5 Preview on 2026-09-20 as its flagship model for software engineering, long-horizon agent tasks, professional knowledge work and finance. The sparse Mixture-of-Experts model has 600B total parameters with 27B activated per token and a 1M-token context window. The official API model ID is step-5-preview; it accepts text, image and video input, returns text, and supports up to 64K output tokens. The API also supports streaming, tool calling, JSON Mode and JSON Schema, low/medium/high reasoning effort, and prompt caching. StepFun reports an Artificial Analysis Intelligence Index score of 44 and a range of coding, agent, finance and multimodal benchmark results; these are vendor-reported results, not ChinaAPI tests. StepFun says open weights will follow on 2026-10-15. ChinaAPI's live catalog now exposes step-5-preview as a customer-callable route at the vendor list price of $1.00/M input, $2.70/M output and $0.05/M prompt-cache reads; the live gateway remains authoritative.
2026-09-20
✓ On ChinaAPI confirmed
Alibaba Qwen Qwen-Image-2.1 open weights
Image Qwen-Image Source · github.com
Qwen released Qwen-Image-2.1 weights on 2026-09-20. Its 7B-parameter visual generation component unifies text-to-image generation and image editing, with native RGBA transparency, up to 10 reference images, local edits and 2K output. These are Qwen-reported capabilities, not ChinaAPI tests. The weights are under the Qwen Research License, which permits only non-commercial research and evaluation without a separate commercial license. Alibaba Model Studio has not published a Qwen-Image-2.1 API model ID or price, and ChinaAPI does not offer this route; no ChinaAPI model page or price is published.
2026-09-20
Not on ChinaAPI confirmed
DeepSeek released and open-sourced DeepSeek-V4.1-Flash on 2026-09-10. The native image-and-text model uses a 552B-parameter Mixture-of-Experts Causal Encoder-Decoder architecture, activating 8B parameters during input prefill and 16B during output decoding, with a 1M context window and up to 384K output. DeepSeek says its vendor benchmark results exceed V4 Pro and that its new KV Cache design uses one quarter of the HBM and one eighth of the SSD storage required by the previous V4 Flash; these are vendor claims, not ChinaAPI tests. The direct DeepSeek API model ID is deepseek-flash. DeepSeek retired the standalone V4 Flash and Vision-Exp checkpoints but temporarily maps their legacy model IDs to V4.1 Flash. ChinaAPI's live catalog now exposes the canonical deepseek-flash ID in place of both legacy Flash IDs; the two retired Flash marketing URLs redirect to that canonical model page. Off-peak pricing is $0.15/M input, $0.60/M output and $0.003/M cache reads; Beijing-time weekday peak pricing is exactly 2x at $0.30/M input, $1.20/M output and $0.006/M cache reads. The live gateway remains authoritative. DeepSeek had also announced that direct deepseek-v4-pro calls would route to V4.1 Flash after 2026-09-14, but later withdrew that plan: V4 Pro remains available after that date with unchanged billing, and deepseek-v4-pro stays a separate ChinaAPI route.
2026-09-10
✓ On ChinaAPI confirmed
Tencent Hunyuan AuK and AuK-Flash open source
Tencent open-sourced AuK and AuK-Flash on 2026-09-09. AuK is a 1.5B speech generation and editing foundation model trained with 1.95 million hours of effective supervision across speech generation, content editing, enhancement and separation, paralinguistic editing, and acoustic editing. A shared natural-language instruction interface supports zero-shot and instruction-based TTS, speech and lyric editing, pitch, speed and volume changes, emotion and timbre editing, de-accenting, nonverbal edits, enhancement, and source separation. AuK-Flash is a distilled four-step variant; the paper reports a 4.5x wall-clock speedup over AuK under matched conditions, which is a vendor result rather than a ChinaAPI test. Both official repositories provide downloadable MIT-licensed weights and use Qwen2.5-Omni-3B as an external encoder. The official code and SGLang-Omni document self-hosted inference, but ChinaAPI does not currently expose AuK as a customer-callable route, so no API price or model page is published.
2026-09-09
Not on ChinaAPI confirmed

Tencent open-sourced Hy4 preview on 2026-08-28 as a 770B-parameter Mixture-of-Experts model with 49B backbone parameters activated per token, plus a native 10B MTP layer for speculative decoding. The text model supports a 1M context window and is published in two official weight repositories: the original checkpoint and an FP8 precision variant. Tencent documents OpenAI-compatible serving through vLLM and SGLang; the official SGLang matrix includes FP8 on 8x B200 or 4x B300/GB300 and BF16 on 16x H200. The release uses Apache-2.0 and is positioned for long-horizon software engineering, office analysis, game development and scientific research; the published evaluations are vendor results, not ChinaAPI tests. ChinaAPI lists hy4-preview as a customer-callable route, with a dedicated model page and deployment-cost analysis; the live catalog remains authoritative for current pricing.
2026-08-28
✓ On ChinaAPI confirmed
Z.ai released and open-sourced GLM-5.3-Flash (320B-A18B) on 2026-08-27 as the GLM-5 family's first native multimodal model. It accepts text, image, video, and file input, with a 1M context window and up to 128K output. The official Chat Completion API keeps thinking enabled and supports function calling, context caching, and structured output. Its low-cost architecture combines sparse and linear attention with Manifold-Constrained Hyper-Connections; Z.ai says the model was pretrained on 30T multimodal tokens and, compared with GLM-5.3, reduces attention computation by 3.01x and average per-layer KV cache by 4.44x. Z.ai also reports 57 on the Artificial Analysis Intelligence Index and coding performance comparable with Claude Opus 4.8; these are vendor claims, not ChinaAPI tests. The glm-5.3-flash route is now customer-callable on ChinaAPI through OpenAI-compatible and Anthropic-compatible endpoints, with text, image, and video input.
2026-08-27
✓ On ChinaAPI confirmed
Qwen released Qwen3.8-Flash-Next on 2026-08-26 as a multimodal Mixture-of-Experts preview of the Qwen4 architecture. The language model has a 125B-parameter main model, 51B additional N-gram embeddings, a 4B MTP module and 6B parameters activated per token, with native 262K context extendable to 1M through YaRN. Qwen publishes official BF16 and FP8 checkpoints; Hugging Face lists them at about 360GB and 186GB, while the vLLM recipe measures 335.28GiB and 172.78GiB. vLLM validates 2x GB300 as the minimum official FP8 topology and reports about 1,430 output tokens/s at concurrency 64 on 4x H100 with the N-gram table offloaded to host memory, using a random 1,024-input/256-output workload. A third-party NVFP4 recipe also reports a single 128GB DGX Spark run, but it uses a non-Qwen-official four-bit checkpoint and custom serving work, so it is not the official minimum. Qwen Community License 1.0 requires a separate license before commercial use for MaaS or a defined AI Work Assistant business. ChinaAPI lists qwen3.8-flash, the managed production model based on Flash-Next, as a customer-callable route; it is not represented as byte-identical to the downloadable preview weights.
2026-08-26
✓ On ChinaAPI confirmed
DeepSeek DeepSeek-V4-Flash-Vision-Exp
LLM DeepSeek V4 Flash Vision Source · api-docs.deepseek.com
DeepSeek launched DeepSeek-V4-Flash-Vision-Exp on its API platform on 2026-08-21 as a standalone experimental visual-understanding model, called with model='deepseek-v4-flash-vision-exp'. It accepts mixed text and image input through OpenAI-compatible Chat Completions and Responses APIs plus the Anthropic-compatible Messages API. Images can be supplied as base64 data, public URLs, or Files API references in JPEG, PNG, GIF, or WebP format. The model has a 1M context window and up to 384K output, supports thinking and non-thinking modes, JSON Output, and Tool Calls, but does not support FIM. DeepSeek says its text-only Agent, reasoning, and world-knowledge ability matches the production DeepSeek-V4-Flash model, while its visual Agent results improve substantially; these benchmark comparisons are vendor claims, not ChinaAPI tests. Each image is converted to at most 384 tokens and uses the official V4-Flash token rates. The deepseek-v4-flash-vision-exp route was customer-callable on ChinaAPI at release; it has since been retired from the live catalog in favor of deepseek-flash, and its old marketing URL redirects to the canonical page.
2026-08-21
Not on ChinaAPI confirmed
Alibaba Qwen Qwen3.8 open weights: 27B and 2.4T-A95B
Qwen published two deployable Qwen3.8 model specifications and four official weight repositories: Qwen3.8-27B and Qwen3.8-27B-FP8 are precision variants of one dense native vision-language model, while Qwen3.8-2.4T-A95B and its FP8 repository are variants of one text-only Mixture-of-Experts model with 95B active parameters. Both specifications support 262K native context and documented self-hosting paths. The 27B release uses Apache-2.0; the 2.4T release uses the custom Qwen3.8-Max License with attribution and high-revenue MaaS or AI Work Assistant conditions. ChinaAPI's live catalog lists customer-callable qwen3.8-27b and qwen3.8-2.4t-a95b routes; this does not claim they are byte-identical to the downloadable weights. Check live pricing for current rates.
2026-08-14
✓ On ChinaAPI confirmed
Z.ai released GLM-5.3 on 2026-08-14 and rolled it out to all GLM Coding Plan users. It uses the same base model as GLM-5.2, with all gains coming from post-training. The text-only model supports a 1M context window and up to 128K output; thinking is always enabled, with low, high, and max reasoning-effort levels. Z.ai reports a 50% improvement over GLM-5.2 on its internal Z.ai Code Bench plus stronger public coding and cybersecurity benchmark results; these are vendor claims, not ChinaAPI tests. The official Model API is now live, and the glm-5.3 route is available on ChinaAPI with OpenAI-compatible access. When migrating from a configuration that disabled thinking, enable thinking and set reasoning_effort to low before switching the model ID, as the official guide warns that otherwise the request will fail.
2026-08-14
✓ On ChinaAPI confirmed
DeepSeek updated the existing deepseek-v4-pro API alias to serve DeepSeek-V4-Pro-0813. The calling method and model ID are unchanged: requests using deepseek-v4-pro receive the latest version. The official quick-start page also confirms that deepseek-v4-flash remains mapped to DeepSeek-V4-Flash-0731. This was a serving-version rollover rather than a new API route. DeepSeek later withdrew a planned switch of deepseek-v4-pro to V4.1 Flash; V4 Pro remains available after 2026-09-14 with unchanged billing, so existing ChinaAPI integrations can keep the same model ID.
2026-08-13
✓ On ChinaAPI confirmed
Alibaba Cloud released qwen3.8-max on 2026-08-03 as Qwen's most capable flagship model to date. The native vision-language model uses a 2.4T-parameter Mixture-of-Experts architecture and supports reasoning, text generation, and visual understanding, with significant improvements over the Qwen3.7 series. Official Model Studio documentation lists OpenAI-compatible, Anthropic-compatible, and DashScope access across Beijing and multiple global regions. The route is now available through ChinaAPI.
2026-08-03
✓ On ChinaAPI confirmed

MiniMax released H3 (Hailuo 3.0). The official release notes describe it as a new-generation open general-purpose multimodal video model that understands creative intent across text, image, video, and audio context and delivers more natural, coherent generation. ChinaAPI serves the route for 4-15 second clips at 768P and 2K with native stereo audio in a single pass, billed per second of output and callable through the OpenAI-compatible video endpoint.
2026-07-31
✓ On ChinaAPI confirmed
ByteDance Doubao-Seedance-2.5
Volcano Engine's Ark model catalog lists Doubao-Seedance-2.5 under model ID doubao-seedance-2-5-260628. Released on 2026-07-31, the video model can generate up to 30 seconds, accept up to 50 reference assets (30 images, 10 videos, and 10 audio clips), and support multimodal reference, video editing, extension, and generation in more than 10 languages. The 260628 suffix remains the model version tag. ChinaAPI now lists this route for customer calls; check the live catalog for current pricing.
2026-07-31
✓ On ChinaAPI confirmed
DeepSeek DeepSeek-V4-Flash-0731
LLM DeepSeek V4 Flash Source · api-docs.deepseek.com
DeepSeek upgraded its existing deepseek-v4-flash API route with a re-post-trained checkpoint that keeps the preview architecture and model size while strengthening coding and agent capabilities. DeepSeek reports 82.7 on Terminal Bench 2.1, 54.4 on DeepSWE, and 70.3 on Toolathlon Verified using max reasoning effort and an unreleased minimal harness; these are vendor results, not ChinaAPI tests. The official route now natively supports the Responses API and Codex integration, while DeepSeek's web app, mobile app, and V4-Pro API were not part of this update.
2026-07-31
Not on ChinaAPI confirmed
Alibaba Cloud Qwen3.7-Flash
Alibaba Cloud released its native vision-language Flash model with 1M context and up to 131K output, upgrading multimodal understanding, real-world and spatial perception, Search Agent and CI Agent execution, and multimodal coding over Qwen3.6-Flash. The official model accepts text, image, and video and exposes built-in tools through the Responses API. ChinaAPI now lists qwen3.7-flash as a customer-callable route; check the live catalog for current pricing.
2026-07-25
✓ On ChinaAPI confirmed
Moonshot released its 2.8T-parameter flagship with 1M context, always-on reasoning, and native image and video understanding. The announced open-source weights are due by July 27, 2026.
2026-07-16
✓ On ChinaAPI confirmed
Tencent officially released Hy3, a 256K-context hybrid fast-and-slow-thinking model with improved stability and cost efficiency.
2026-07-06
✓ On ChinaAPI confirmed

Video HappyHorse
HappyHorse 1.1 text-to-video gateway model with improved semantic understanding and camera control.
2026-06-22
✓ On ChinaAPI confirmed
Video HappyHorse
HappyHorse 1.1 reference-to-video gateway model for multi-reference subject, scene, and style consistency.
2026-06-22
✓ On ChinaAPI confirmed
Video HappyHorse
HappyHorse 1.1 image-to-video gateway model with stronger texture and identity consistency.
2026-06-22
✓ On ChinaAPI confirmed
Coding-focused Kimi gateway model for long-horizon repository work with vision support.
2026-06-12
✓ On ChinaAPI confirmed
MiMo speech recognition for Chinese-English code-switching, Cantonese, Wu, Minnan and Sichuan dialects, noisy and far-field recordings, overlapping speakers, and lyrics.
2026-06-02
✓ On ChinaAPI confirmed
Multimodal Qwen gateway model for visual analysis, planning, files, and tool use.
2026-06-01
✓ On ChinaAPI confirmed
Native multimodal M-series model with 1M context, MSA attention, coding, agent, image, and video input support.
2026-06-01
✓ On ChinaAPI confirmed

Fast StepFun gateway model with 256K context, vision input, reasoning, and tool use.
2026-05-28
✓ On ChinaAPI confirmed
Meituan LongCat-Video-Avatar 1.5
Video LongCat Video Source · arxiv.org
Audio-driven avatar video technical release focused on lip sync, identity consistency, and long-video generation.
2026-05-26
Not on ChinaAPI confirmed
Flagship Qwen gateway model for demanding analysis, coding, and production agent workflows.
2026-05-21
✓ On ChinaAPI confirmed

DeepSeek deepseek-v4-flash
Fast DeepSeek V4 variant for efficient chat, coding assistance, and agent loops.
2026-04-24
Not on ChinaAPI confirmed
MiMo expressive text-to-speech with premium built-in voices and natural-language control over pace, emotion, tone, dialect, and character performance.
2026-04-23
✓ On ChinaAPI confirmed
Zhipu AI GLM-5V-Turbo
A 200K-context multimodal coding foundation model for image, video, file, GUI-agent, and tool-driven workflows.
2026-04-07
Not on ChinaAPI confirmed
Wan 2.7 natural-language video editing model for local/global edits and element replacement.
2026-04-03
✓ On ChinaAPI confirmed
Video Wan
Wan 2.7 text-to-video model for prompt-driven video generation.
2026-04-03
✓ On ChinaAPI confirmed
Video Wan
Wan 2.7 reference-to-video model for controlled generation from visual references.
2026-04-03
✓ On ChinaAPI confirmed
Wan 2.7 image model for text-to-image, editing, multi-reference generation, and stronger text rendering.
2026-04-01
✓ On ChinaAPI confirmed

Alibaba Qwen3.5
Native multimodal-agent Qwen release that preceded the served Qwen 3.6 and 3.7 gateway models.
2026-02-16
Not on ChinaAPI confirmed
StepFun Step-3.5-Flash
Open StepFun flash model with 256K context and tool-use support. ChinaAPI lists step-3.5-flash for customer calls; check the live catalog for current pricing.
2026-02-01
✓ On ChinaAPI confirmed

Seedream 5.0-lite image model with controllable creation, retrieval, and stronger subject consistency.
2026-01-28
✓ On ChinaAPI confirmed
MiniMax released Speech 2.8 with native sound tags, improved voice cloning, studio-grade audio, and cross-lingual speech; Turbo is its speed-focused option.
2026-01-23
✓ On ChinaAPI confirmed
MiniMax released Speech 2.8 with native sound tags, improved voice cloning, studio-grade audio, and cross-lingual speech; HD is its high-fidelity option.
2026-01-23
✓ On ChinaAPI confirmed
Moonshot kimi-k2.5
Open native multimodal Kimi agent model with 256K context and vision-language training.
2026-01-01
Not on ChinaAPI confirmed

Zhipu AI GLM-ASR-2512
Audio GLM Audio Source · docs.z.ai
Z.AI released its multilingual ASR model with improved custom-dictionary support and specialised-terminology recognition.
2025-12-10
Not on ChinaAPI confirmed
Meituan LongCat-Image
Small image-generation model focused on data quality over model size.
2025-12
Not on ChinaAPI confirmed

Alibaba released Qwen3-TTS-Flash for offline speech synthesis with expressive voices, multilingual and dialect output, and low-latency API access.
2025-09-22
✓ On ChinaAPI confirmed

Context-aware text-to-speech with global and inline natural-language direction, expressive delivery, and direct OpenAI-compatible /v1/audio/speech output.
Release date not publicly confirmed
✓ On ChinaAPI unknown
StepFun Step TTS Mini
Cost-efficient multilingual speech synthesis with official and cloned voices, pronunciation mapping, and style controls.
Release date not publicly confirmed
Not on ChinaAPI unknown
Expressive speech synthesis with official and cloned voices, pronunciation mapping, style controls, and OpenAI-compatible output.
Release date not publicly confirmed
✓ On ChinaAPI unknown
Context-aware expressive speech synthesis with streaming output and controls for voice, speed, and volume.
Release date not publicly confirmed
✓ On ChinaAPI unknown