Zhipu / GLM API model

Zhipu: glm-5.3-flash

Zhipu GLM-5.3-Flash (320B-A18B) — the GLM-5 family's first native multimodal model for visual coding and agent workflows. It accepts text, image, video, and file input, with a 1M context window, up to 128K output, always-on thinking, function calling, structured output, streaming, and context caching.

Specs

Pricing and API details.

Transparent USD pricing. No mainland-China account or phone required. Standard model rates follow providers' international list prices at 100%; some models have a separate Paid rate after the first top-up.

Context window

1M

Modalities

Text, image and video input

Input $/1M

$0.15

Output $/1M

$0.5

Endpoint

OpenAI-compatible /v1/chat/completions

About glm-5.3-flash

Zhipu GLM-5.3-Flash (320B-A18B) — the GLM-5 family's first native multimodal model for visual coding and agent workflows. It accepts text, image, video, and file input, with a 1M context window, up to 128K output, always-on thinking, function calling, structured output, streaming, and context caching.

Pricing

Transparent USD pricing.

glm-5.3-flash

$0.15/M input, $0.5/M output tokens. Paid in USD. Standard model rates follow providers' international list prices at 100%; some models have a separate Paid rate after the first top-up. View live pricing Automatic cache hits are billed at $0.0300 per 1M input tokens; the standard input price above is the cache-miss rate.

Quickstart

Copy-paste API request.

curl -X POST https://api.chinaapi.ai/v1/chat/completions \
  -H "Authorization: Bearer $CHINAAPI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "glm-5.3-flash",
    "messages": [
      {"role": "user", "content": "Summarize the main tradeoffs for this API workload."}
    ]
  }'

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["CHINAAPI_API_KEY"],
    base_url="https://api.chinaapi.ai/v1",
)

response = client.chat.completions.create(
    model="glm-5.3-flash",
    messages=[
        {"role": "user", "content": "Summarize the main tradeoffs for this API workload."}
    ],
)
print(response.choices[0].message.content)

Internal links

Related model pages and guides.

How do I use glm-5.3-flash?

Use glm-5.3-flash through ChinaAPI with an API key and the OpenAI-compatible endpoint shown above.

How much does glm-5.3-flash cost?

glm-5.3-flash is $0.15/M input, $0.5/M output tokens, paid in USD. The live pricing page is authoritative. After the first top-up, the Paid price is $0.135/M input, $0.45/M output tokens (−10%). Standard model rates follow providers' international list prices at 100%; some models have a separate Paid rate after the first top-up.

Does glm-5.3-flash need a Chinese account?

No. ChinaAPI provides access without a mainland-China account or phone number and includes a $2 free trial. Check live pricing for the displayed USD rate.