# Chinese AI API Pricing Landscape — July 2026

> What Chinese LLM, image, and video APIs cost in July 2026 — the sub-$0.20 tier, the 1M-context club, per-second vs per-generation video math, and what to pick by workload.

- Written by: ChinaAPI Research
- Published: 2026-07-06 (last updated 2026-07-10)
- Canonical: https://chinaapi.ai/insights/chinese-ai-api-pricing-2026-07/
- Try the models: https://dash.chinaapi.ai/register?lang=en&utm_source=chinaapi&utm_medium=md&utm_campaign=pricing-landscape-2026-07

Chinese model vendors have spent the last quarter undercutting each other, and the result is the cheapest frontier-adjacent API market in the world. This analysis currently covers **68 models live on our gateway**: 30 LLMs and 38 image/video models. Every price below is generated from [our public pricing dataset](/data/model-pricing/) rather than copied into this article by hand. Dataset revision: 2026-08-10; pricing fingerprint: `c59b3d55d9fea19c`.

One honest caveat up front: **price is not quality.** This article tells you what things cost; our benchmark series (task completion rates per model, measured through this same gateway) is the companion that tells you what they're worth.

## The sub-$0.20 input tier

6 models charge less than $0.20 per million input tokens. The important comparison inside this tier is output cost: two models with the same cheap input rate can diverge sharply once they start generating long answers.

| Model | Vendor | Input /1M | Output /1M | Context |
|---|---|---:|---:|---:|
| step-3.5-flash | StepFun | $0.1 | $0.3 | 256K |
| step-3.5-flash-2603 | StepFun | $0.1 | $0.3 | 256K |
| deepseek-v4-flash | DeepSeek | $0.14 | $0.28 | 1M |
| mimo-v2.5 | Xiaomi | $0.14 | $0.28 | 1M |
| hy3 | Tencent | $0.1538 | $0.6154 | 256K |
| qwen3.6-flash | Alibaba | $0.1846 | $1.1077 | 1M |

Two things stand out. First, `deepseek-v4-flash` and Xiaomi's `mimo-v2.5` currently have the same published input and output rates. Second, watch **output** prices, not just input: `qwen3.6-flash` sits in the same low-input tier but its output rate is $1.1077, versus $0.28 for `deepseek-v4-flash`. That gap dominates generation-heavy workloads.

## The 1M-context club

13 of the 30 LLMs in the current catalog publish a 1M-token context window: `qwen3.6-flash`, `qwen3.6-plus`, `qwen3.7-max`, `qwen3.7-plus`, `qwen3.8-max`, `deepseek-v4-flash`, `deepseek-v4-pro`, `LongCat-2.0`, `MiniMax-M3`, `kimi-k3`, `mimo-v2.5`, `mimo-v2.5-pro`, `glm-5.2`. A year ago 1M context was a premium feature; in the Chinese market it is becoming table stakes.

The cheapest published input rate among those million-context models is represented by `deepseek-v4-flash` at $0.14 per million input tokens. The most interesting long-context entry is Meituan's `LongCat-2.0`: an open MoE reasoning model at $0.3 input / $1.2 output with the full 1M window — flagship-class specs at flash-tier pricing.

## Flagship pricing by vendor

Taking each vendor's highest published **output** price in the current text catalog as a consistent proxy—not a quality ranking—produces this comparison:

| Vendor | Highest-priced text model | Input /1M | Output /1M | Context |
|---|---|---:|---:|---:|
| Moonshot | kimi-k3 | $2.7397 | $13.6986 | 1M |
| StepFun | step-1o-audio | $3.85 | $9.24 | — |
| Alibaba | qwen3.7-max | $2.5 | $7.5 | 1M |
| ByteDance | doubao-seed-2-1-pro-260628 | $0.9231 | $4.6154 | 256K |
| Zhipu AI | glm-5-turbo | $1.2 | $4.4571 | 200K |
| MiniMax | MiniMax-M2.7-highspeed | $0.5753 | $2.3014 | 200K |
| Meituan | LongCat-2.0 | $0.3 | $1.2 | 1M |
| DeepSeek | deepseek-v4-pro | $0.435 | $0.87 | 1M |
| Xiaomi | mimo-v2.5-pro | $0.435 | $0.87 | 1M |
| Tencent | hy3 | $0.1538 | $0.6154 | 256K |

Moonshot's high-speed coding line currently carries the highest output rate in this table at $7.3973 per million tokens. Tencent's `hy3` is at the other end at $0.6154. That spread is a pricing claim by the vendors, not evidence of the same gap in quality; only task-completion benchmarks can answer whether the premium holds.

Tencent's `hy3` deserves a flag: a major-vendor model at $0.1538 input / $0.6154 output per million tokens. Access friction outside China is exactly the gap this gateway exists to close.



## Video: per-second vs per-clip, with real math

Video pricing comes in two shapes, and comparing them requires arithmetic vendors don't do for you:

| Model | Vendor | Billing | A 5-second representative clip |
|---|---|---:|---:|
| happyhorse-1.1-t2v | Alibaba | $0.075 per second | $0.375 |
| wan2.7-t2v | Alibaba | $0.1 per second | $0.5 |
| MiniMax-Hailuo-2.3-Fast | MiniMax | $0.1875 per generation | $0.1875 |
| MiniMax-Hailuo-2.3 | MiniMax | $0.2778 per generation | $0.2778 |
| kling-v3 | Kuaishou | $0.4799 per 5-second 720p silent generation | $0.4799 |
| kling-3.0-turbo | Kuaishou | $0.5556 per generation | $0.5556 |

For short drafts, per-generation `MiniMax-Hailuo-2.3-Fast` is $0.1875 in the representative catalog configuration. Per-second billing (HappyHorse, Wan) can win when you need very short clips or precise duration control. ByteDance's Seedance line has configuration-specific pricing, so the public dataset intentionally marks rows without one representative unit price and points to live pricing instead.

## What to pick, by workload

Our stance, from pricing and specs alone (quality benchmarks follow separately):

- **Bulk extraction, classification, RAG pre-processing**: `deepseek-v4-flash` or `mimo-v2.5` — both currently publish $0.14/M input with a 1M window.
- **Long-document agents on a budget**: `LongCat-2.0` — 1M context and reasoning-tuned at $0.3 input / $1.2 output.
- **Multimodal agent loops (image and video input)**: `MiniMax-M3` — per the vendors' own published specs it is currently the only 1M-context natively multimodal model in this catalog.
- **Low-cost video drafts among representative per-generation rows**: `MiniMax-Hailuo-2.3-Fast`, currently $0.1875 per generation.
- **Premium coding**: Moonshot's `kimi-k2.7-code` line is the price outlier. Benchmark it on your own repository before committing spend.

## What this article doesn't tell you

- **Quality.** A cheap model that fails your task twice can cost more than a premium model that succeeds once. Our task-completion benchmark is the companion piece.
- **Stability under load.** Chinese vendor APIs vary in rate limits and peak-hour behavior; we route around some of this at the gateway layer, but haven't published numbers yet.
- **Tomorrow's prices.** Vendors reprice aggressively (usually downward, occasionally promo-then-revert). We now snapshot every price change, so future editions of this article will show trend lines, not just a snapshot.

All 68 models run behind one OpenAI-compatible endpoint with a single key. Download the [JSON](/data/model-pricing.json) or [CSV](/data/model-pricing.csv), read the [methodology](/data/model-pricing/), and use the $2 free credit to verify a published price with a real request.

---
<!-- GENERATED BY scripts/gen_insights.py; locale: en; Markdown twin -->
Markdown twin of https://chinaapi.ai/insights/chinese-ai-api-pricing-2026-07/. Full site reference: https://chinaapi.ai/llms-full.txt
