Task-led research & comparison

Choose a Chinese AI model by the constraint that matters.

Start with coding, context, latency measurement, video, or value. Each route separates current model facts from the decision you still need to test. Standard token-model rates follow providers' international list prices at 100%; some models have a separate Paid rate after the first top-up. Media rows show transparent USD service rates; any service margin is included in the displayed rate.

This is a decision hub, not a popularity or quality leaderboard. We publish only current catalog facts: context window, listed USD price, modality, and declared capabilities. For independent quality benchmarks, see Artificial Analysis.

33Token-billed LLMs with current context, price, and capability facts.
18LLMs with 1M context windows in the current model data.
21Image and video generation models listed separately with media billing.

Start with the decision

Five routes, one evidence standard.

The route cards state the evidence we have, when it is useful, and where it stops. They lead to the relevant model facts and a single UTM-tagged registration path.

Coding & agents

Start with models whose current descriptions or capabilities explicitly mention coding, tools, agents, engineering, planning, or structured work.

  • Use it when: Use this route when tool use, code generation, or an agent loop is the bottleneck.
  • Limit: Capability tags describe declared fit; they are not a measured coding-quality score.
Open this decision route

Long context

Use the current context-window field to narrow candidates before evaluating retrieval quality, prompt structure, and total token cost.

  • Use it when: Use this route for large repositories, long documents, and multi-step knowledge work.
  • Limit: A larger context window does not prove better recall, reasoning, or latency.
Open this decision route

Low latency

ChinaAPI does not publish a cross-region latency ranking yet. This route gives you a repeatable way to measure your own prompt, region, and streaming mode.

  • Use it when: Use this route when interaction time or first-token time is a release gate.
  • Limit: Do not infer latency from listed price, context, or an individual request.
Open this decision route

Video generation

Compare current video modalities, capabilities, and transparent per-unit USD service prices before testing your own prompt and duration.

  • Use it when: Use this route for story, motion, reference-image, and audio requirements.
  • Limit: The catalog is not a visual-quality or task-success ranking.
Open this decision route

Value

Use listed input price as a cost filter, then test accepted-result cost with your workload, retries, output length, and operating constraints.

  • Use it when: Use this route when budget is a hard constraint and a direct price comparison is useful.
  • Limit: The lowest listed input price is not a quality, latency, or total-cost claim.
Open this decision route

Model facts by task

Shortlist with facts, then validate your workload.

Models can appear in more than one group. Within these compact tables, rows are ordered by larger context window first, then lower listed input price.

Coding & agents

LLMs whose descriptions or capabilities reference coding, agents, tools, engineering, planning, or office work. This is catalog evidence, not a measured quality ranking.

ModelVendorContextInput $/1M
mimo-v2.5Xiaomi1M$0.14
glm-5.3-flashZhipu AI1M$0.15
qwen3.8-flashAlibaba1M$0.15
deepseek-flashDeepSeek1M$0.3
LongCat-2.0Meituan1M$0.3
mimo-v2.5-proXiaomi1M$0.435
hy4-previewTencent1M$0.8889
MiniMax-M3MiniMax1M$0.9
qwen3.6-flashAlibaba1M$1
step-5-previewStepFun1M$1
qwen3.7-plusAlibaba1M$1.2
deepseek-v4-proDeepSeek1M$1.32
glm-5.2Zhipu AI1M$1.4
glm-5.3Zhipu AI1M$1.4
qwen3.6-plusAlibaba1M$2
qwen3.8-maxAlibaba1M$2
qwen3.7-maxAlibaba1M$2.5
kimi-k3Moonshot1M$3
step-3.5-flashStepFun256K$0.1
step-3.5-flash-2603StepFun256K$0.1
hy3Tencent256K$0.1481
step-3.7-flashStepFun256K$0.2
doubao-seed-2-1-turbo-260628ByteDance256K$0.4444
doubao-seed-2-1-pro-260628ByteDance256K$0.8889
kimi-k2.6Moonshot256K$0.95
kimi-k2.7-codeMoonshot256K$0.95
kimi-k2.7-code-highspeedMoonshot256K$1.9
MiniMax-M2.7MiniMax200K$0.3
MiniMax-M2.7-highspeedMiniMax200K$0.6
glm-5Zhipu AI200K$1
glm-5-turboZhipu AI200K$1.2
glm-5.1Zhipu AI200K$1.4
stepaudio-2.5-chatStepFun$1.5

Best for Reasoning

LLMs tagged with Reasoning in the current data source.

ModelVendorContextInput $/1M
mimo-v2.5Xiaomi1M$0.14
glm-5.3-flashZhipu AI1M$0.15
qwen3.8-flashAlibaba1M$0.15
deepseek-flashDeepSeek1M$0.3
LongCat-2.0Meituan1M$0.3
mimo-v2.5-proXiaomi1M$0.435
hy4-previewTencent1M$0.8889
MiniMax-M3MiniMax1M$0.9
qwen3.6-flashAlibaba1M$1
step-5-previewStepFun1M$1
qwen3.7-plusAlibaba1M$1.2
deepseek-v4-proDeepSeek1M$1.32
glm-5.2Zhipu AI1M$1.4
glm-5.3Zhipu AI1M$1.4
qwen3.6-plusAlibaba1M$2
qwen3.8-maxAlibaba1M$2
qwen3.7-maxAlibaba1M$2.5
kimi-k3Moonshot1M$3
step-3.5-flashStepFun256K$0.1
step-3.5-flash-2603StepFun256K$0.1
hy3Tencent256K$0.1481
step-3.7-flashStepFun256K$0.2
doubao-seed-2-1-turbo-260628ByteDance256K$0.4444
doubao-seed-2-1-pro-260628ByteDance256K$0.8889
kimi-k2.6Moonshot256K$0.95
kimi-k2.7-codeMoonshot256K$0.95
kimi-k2.7-code-highspeedMoonshot256K$1.9
MiniMax-M2.7MiniMax200K$0.3
MiniMax-M2.7-highspeedMiniMax200K$0.6
glm-5Zhipu AI200K$1
glm-5-turboZhipu AI200K$1.2
glm-5.1Zhipu AI200K$1.4
stepaudio-2.5-chatStepFun$1.5

Long context

All LLMs with a 1M context window, ordered by lower listed input price.

ModelVendorContextInput $/1M
mimo-v2.5Xiaomi1M$0.14
glm-5.3-flashZhipu AI1M$0.15
qwen3.8-flashAlibaba1M$0.15
deepseek-flashDeepSeek1M$0.3
LongCat-2.0Meituan1M$0.3
mimo-v2.5-proXiaomi1M$0.435
hy4-previewTencent1M$0.8889
MiniMax-M3MiniMax1M$0.9
qwen3.6-flashAlibaba1M$1
step-5-previewStepFun1M$1
qwen3.7-plusAlibaba1M$1.2
deepseek-v4-proDeepSeek1M$1.32
glm-5.2Zhipu AI1M$1.4
glm-5.3Zhipu AI1M$1.4
qwen3.6-plusAlibaba1M$2
qwen3.8-maxAlibaba1M$2
qwen3.7-maxAlibaba1M$2.5
kimi-k3Moonshot1M$3

Multimodal

LLMs with vision=true. Text-only models are not listed in this group.

ModelVendorContextInput $/1M
mimo-v2.5Xiaomi1M$0.14
glm-5.3-flashZhipu AI1M$0.15
qwen3.8-flashAlibaba1M$0.15
deepseek-flashDeepSeek1M$0.3
mimo-v2.5-proXiaomi1M$0.435
MiniMax-M3MiniMax1M$0.9
qwen3.6-flashAlibaba1M$1
step-5-previewStepFun1M$1
qwen3.7-plusAlibaba1M$1.2
qwen3.6-plusAlibaba1M$2
qwen3.8-maxAlibaba1M$2
kimi-k3Moonshot1M$3
step-3.7-flashStepFun256K$0.2
doubao-seed-2-1-turbo-260628ByteDance256K$0.4444
doubao-seed-2-1-pro-260628ByteDance256K$0.8889
kimi-k2.6Moonshot256K$0.95
kimi-k2.7-codeMoonshot256K$0.95
kimi-k2.7-code-highspeedMoonshot256K$1.9

From shortlist to implementation

Choose a model, verify its setup path, then register.

Model cards narrow the catalog; the integration guides provide the next configuration step. Guide links are shown only for tools with a dated verification record, so a catalog capability tag is never presented as a compatibility guarantee.

Continue with a verified setup guide

Open the tool guide that matches your workflow, then start with a small test before routing production traffic.

Low-latency validation

Measure the workflow you will actually ship.

We do not publish a cross-region latency ranking because the current catalog does not contain a repeatable multi-region latency dataset. Keep the model, prompt, request shape, region, concurrency, and streaming mode fixed when you test.

A repeatable latency check

Record time to first token and total completion time across multiple runs. Compare only runs with the same prompt and output limit, then keep the raw measurements with the test date and region.

Limit: a single request, a provider specification, or a listed price cannot support a latency claim. Until a documented probe publishes sample size and region, this hub will not label any model as the lowest-latency option.

Longest context window

All LLMs ordered by context.

Sorted by context window descending, then listed input price ascending. Output pricing is shown for token-billed models.

RankModelVendorContextInput $/1MOutput $/1M
1mimo-v2.5Xiaomi1M$0.14$0.28
2glm-5.3-flashZhipu AI1M$0.15$0.5
3qwen3.8-flashAlibaba1M$0.15$0.47
4deepseek-flashDeepSeek1M$0.3$1.2
5LongCat-2.0Meituan1M$0.3$1.2
6mimo-v2.5-proXiaomi1M$0.435$0.87
7hy4-previewTencent1M$0.8889$2.6667
8MiniMax-M3MiniMax1M$0.9$3.6
9qwen3.6-flashAlibaba1M$1$4
10step-5-previewStepFun1M$1$2.7
11qwen3.7-plusAlibaba1M$1.2$4.8
12deepseek-v4-proDeepSeek1M$1.32$3.96
13glm-5.2Zhipu AI1M$1.4$4.4
14glm-5.3Zhipu AI1M$1.4$4.4
15qwen3.6-plusAlibaba1M$2$6
16qwen3.8-maxAlibaba1M$2$6
17qwen3.7-maxAlibaba1M$2.5$7.5
18kimi-k3Moonshot1M$3$15
19step-3.5-flashStepFun256K$0.1$0.3
20step-3.5-flash-2603StepFun256K$0.1$0.3
21hy3Tencent256K$0.1481$0.5926
22step-3.7-flashStepFun256K$0.2$1.15
23doubao-seed-2-1-turbo-260628ByteDance256K$0.4444$2.2222
24doubao-seed-2-1-pro-260628ByteDance256K$0.8889$4.4444
25kimi-k2.6Moonshot256K$0.95$4
26kimi-k2.7-codeMoonshot256K$0.95$4
27kimi-k2.7-code-highspeedMoonshot256K$1.9$8
28MiniMax-M2.7MiniMax200K$0.3$1.2
29MiniMax-M2.7-highspeedMiniMax200K$0.6$2.4
30glm-5Zhipu AI200K$1$3.2
31glm-5-turboZhipu AI200K$1.2$4
32glm-5.1Zhipu AI200K$1.4$4.4
33stepaudio-2.5-chatStepFun$1.5$3.5

Value filter

LLMs ordered by listed input price.

Sorted by listed input USD per 1M tokens, ascending. Use this as a cost filter, not a quality, latency, or accepted-result-cost ranking.

RankModelVendorContextInput $/1MOutput $/1M
1step-3.5-flashStepFun256K$0.1$0.3
2step-3.5-flash-2603StepFun256K$0.1$0.3
3mimo-v2.5Xiaomi1M$0.14$0.28
4hy3Tencent256K$0.1481$0.5926
5glm-5.3-flashZhipu AI1M$0.15$0.5
6qwen3.8-flashAlibaba1M$0.15$0.47
7step-3.7-flashStepFun256K$0.2$1.15
8deepseek-flashDeepSeek1M$0.3$1.2
9LongCat-2.0Meituan1M$0.3$1.2
10MiniMax-M2.7MiniMax200K$0.3$1.2
11mimo-v2.5-proXiaomi1M$0.435$0.87
12doubao-seed-2-1-turbo-260628ByteDance256K$0.4444$2.2222
13MiniMax-M2.7-highspeedMiniMax200K$0.6$2.4
14hy4-previewTencent1M$0.8889$2.6667
15doubao-seed-2-1-pro-260628ByteDance256K$0.8889$4.4444
16MiniMax-M3MiniMax1M$0.9$3.6
17kimi-k2.6Moonshot256K$0.95$4
18kimi-k2.7-codeMoonshot256K$0.95$4
19qwen3.6-flashAlibaba1M$1$4
20step-5-previewStepFun1M$1$2.7
21glm-5Zhipu AI200K$1$3.2
22qwen3.7-plusAlibaba1M$1.2$4.8
23glm-5-turboZhipu AI200K$1.2$4
24deepseek-v4-proDeepSeek1M$1.32$3.96
25glm-5.2Zhipu AI1M$1.4$4.4
26glm-5.3Zhipu AI1M$1.4$4.4
27glm-5.1Zhipu AI200K$1.4$4.4
28stepaudio-2.5-chatStepFun$1.5$3.5
29kimi-k2.7-code-highspeedMoonshot256K$1.9$8
30qwen3.6-plusAlibaba1M$2$6
31qwen3.8-maxAlibaba1M$2$6
32qwen3.7-maxAlibaba1M$2.5$7.5
33kimi-k3Moonshot1M$3$15

Video generation models

Kling, Seedance, Hailuo, Wan, and HappyHorse.

Video models do not have LLM context windows in the data source. They are billed in USD by the unit shown below. Test the prompt, duration, resolution, reference-control needs, and failure mode before choosing a production route.

ModelVendorCapabilitiesBilling
kling-3.0-turboKuaishouVideo, Text-to-Video, Image-to-Video, Audio, Fast$0.5556 per generation
MiniMax-H3MiniMaxVideo, Text-to-Video, Image-to-Video, Reference-to-Video, Audio, Open Weights$0.08 per second
doubao-seedance-2-5-260628ByteDanceVideo, Text-to-Video, Image-to-VideoPer generation
MiniMax-Hailuo-2.3MiniMaxVideo, Text-to-Video, Image-to-Video$0.2778 per generation
doubao-seedance-2-0-mini-260615ByteDanceVideo, Text-to-Video, FastPer generation
kling-v3KuaishouVideo, Text-to-Video, Image-to-Video, Audio$0.4799 per 5-second 720p silent generation
kling-v3-omniKuaishouVideo, Reference-to-Video, Editing$0.4167 per generation
MiniMax-Hailuo-02MiniMaxVideo, Text-to-Video, Image-to-Video$0.2778 per generation
doubao-seedance-2-0-260128ByteDanceVideo, Text-to-Video, Image-to-VideoPer generation
MiniMax-Hailuo-2.3-FastMiniMaxVideo, Image-to-Video, Fast$0.1875 per generation
doubao-seedance-2-0-fast-260128ByteDanceVideo, Text-to-Video, Image-to-Video, FastPer generation
happyhorse-1.1-t2vAlibabaVideo, Text-to-Video$0.075 per second
happyhorse-1.1-r2vAlibabaVideo, Reference-to-Video$0.075 per second
happyhorse-1.1-i2vAlibabaVideo, Image-to-Video$0.075 per second
wan2.7-t2vAlibabaVideo, Text-to-Video$0.1 per second
wan2.7-r2vAlibabaVideo, Reference-to-Video$0.35 per second
wan2.7-videoeditAlibabaVideo, Video-Edit$0.2 per second
wan2.7-i2vAlibabaVideo, Image-to-Video, Audio$0.1 per second

Image generation models

Seedream and Wan image models.

Image models do not have LLM context windows in the data source. They are billed per image in USD.

ModelVendorCapabilitiesBilling
doubao-seedream-5-0-260128ByteDanceImage, Text-to-Image, Image-to-Image, 4K$0.0306 per image
wan2.7-imageAlibabaImage, Text-to-Image, Image Editing, Multi-Reference$0.0278 per image
wan2.7-image-proAlibabaImage, Text-to-Image, Image Editing, Multi-Reference, 4K$0.0694 per image

How should I choose a Chinese AI model?

Start with the route that matches your constraint, then test candidates with your prompts and acceptance criteria. This hub publishes current catalog facts, not a universal quality leaderboard. For independent quality benchmarks, use a third-party source such as Artificial Analysis.

Which models have the longest context?

The 1M-context LLMs in the current data are mimo-v2.5, glm-5.3-flash, qwen3.8-flash, deepseek-flash, LongCat-2.0, mimo-v2.5-pro, hy4-preview, MiniMax-M3, qwen3.6-flash, step-5-preview, qwen3.7-plus, deepseek-v4-pro, glm-5.2, glm-5.3, qwen3.6-plus, qwen3.8-max, qwen3.7-max, kimi-k3.

Which models are cheapest?

Among token-billed LLMs, step-3.5-flash and step-3.5-flash-2603 share the lowest listed input price at $0.1 per 1M input tokens. The next listed input price is mimo-v2.5 at $0.14 per 1M input tokens.

Which model has the lowest latency?

ChinaAPI does not publish a cross-region latency ranking. Measure the same prompt, region, concurrency, and streaming mode before using latency as a production decision.

Start testing

Use the decision routes as a shortlist, then test your own workload.

Catalog facts narrow the field. Final model choice should still come from your prompts, data, latency needs, acceptance criteria, and total cost per accepted result.

Try ChinaAPI

Start with free credit or review live pricing before routing production traffic.