Zhipu / GLM API model

Zhipu: glm-5.3-flash

Zhipu GLM-5.3-Flash (320B-A18B) — the GLM-5 family's first native multimodal model for visual coding and agent workflows. It accepts text, image, video, and file input, with a 1M context window, up to 128K output, always-on thinking, function calling, structured output, streaming, and context caching.

Specs

Pricing and API details.

Transparent USD pricing. No mainland-China account or phone required. Token-model rates are synchronized from the gateway and may follow providers' official China list prices where configured. Live pricing is authoritative.

Context window

1M

Modalities

Text, image and video input

Input $/1M

$0.15

Output $/1M

$0.5

Endpoint

OpenAI-compatible /v1/chat/completions

About glm-5.3-flash

Zhipu GLM-5.3-Flash (320B-A18B) — the GLM-5 family's first native multimodal model for visual coding and agent workflows. It accepts text, image, video, and file input, with a 1M context window, up to 128K output, always-on thinking, function calling, structured output, streaming, and context caching.

Pricing

Transparent USD pricing.

glm-5.3-flash

$0.15/M input, $0.5/M output tokens. Paid in USD. Token-model rates are synchronized from the gateway and may follow providers' official China list prices where configured. Live pricing is authoritative. View live pricing Automatic cache hits are billed at $0.0300 per 1M input tokens; the standard input price above is the cache-miss rate.

Quickstart

Copy-paste API request.

curl -X POST https://api.chinaapi.ai/v1/chat/completions \
  -H "Authorization: Bearer $CHINAAPI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "glm-5.3-flash",
    "messages": [
      {"role": "user", "content": "Summarize the main tradeoffs for this API workload."}
    ]
  }'

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["CHINAAPI_API_KEY"],
    base_url="https://api.chinaapi.ai/v1",
)

response = client.chat.completions.create(
    model="glm-5.3-flash",
    messages=[
        {"role": "user", "content": "Summarize the main tradeoffs for this API workload."}
    ],
)
print(response.choices[0].message.content)

Internal links

Related model pages and guides.

How do I use glm-5.3-flash?

Use glm-5.3-flash through ChinaAPI with an API key and the OpenAI-compatible endpoint shown above.

How much does glm-5.3-flash cost?

glm-5.3-flash is $0.15/M input, $0.5/M output tokens, paid in USD. The live pricing page is authoritative. Token-model rates are synchronized from the gateway and may follow providers' official China list prices where configured. Live pricing is authoritative.

Does glm-5.3-flash need a Chinese account?

No. ChinaAPI provides access without a mainland-China account or phone number and includes a $2 free trial. Check live pricing for the displayed USD rate.