# Zhipu: glm-5.3-flash — API on ChinaAPI

> Zhipu GLM-5.3-Flash (320B-A18B) — the GLM-5 family's first native multimodal model for visual coding and agent workflows. It accepts text, image, video, and file input, with a 1M context window, up to 128K output, always-on thinking, function calling, structured output, streaming, and context caching.

- Canonical page: https://chinaapi.ai/models/glm-5.3-flash/
- Model ID: `glm-5.3-flash`
- Context window: 1M
- Modalities: Text, image and video input
- Pricing: $0.15/M input, $0.5/M output tokens (USD)
- Endpoint: OpenAI-compatible `POST https://api.chinaapi.ai/v1/chat/completions`
- Price note: Automatic cache hits are billed at $0.0300 per 1M input tokens; the standard input price above is the cache-miss rate.
- Pricing basis: Token-model rates are synchronized from the gateway and may follow providers' official China list prices where configured. Live pricing is authoritative.
- Access: no mainland-China account or phone number required; $2 free trial credit
- Sign up: https://dash.chinaapi.ai/register?lang=en&utm_source=chinaapi&utm_medium=model-md&utm_campaign=glm-5.3-flash

## Quickstart

```
curl -X POST https://api.chinaapi.ai/v1/chat/completions \
  -H "Authorization: Bearer $CHINAAPI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "glm-5.3-flash",
    "messages": [
      {"role": "user", "content": "Summarize the main tradeoffs for this API workload."}
    ]
  }'

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["CHINAAPI_API_KEY"],
    base_url="https://api.chinaapi.ai/v1",
)

response = client.chat.completions.create(
    model="glm-5.3-flash",
    messages=[
        {"role": "user", "content": "Summarize the main tradeoffs for this API workload."}
    ],
)
print(response.choices[0].message.content)
```

## FAQ

**How do I use glm-5.3-flash?**
Use glm-5.3-flash through ChinaAPI with an API key and the OpenAI-compatible endpoint shown on this page.

**How much does glm-5.3-flash cost?**
glm-5.3-flash is $0.15/M input, $0.5/M output tokens, paid in USD. The live pricing page is authoritative. Token-model rates are synchronized from the gateway and may follow providers' official China list prices where configured. Live pricing is authoritative.

**Does glm-5.3-flash need a Chinese account?**
No. ChinaAPI provides access without a mainland-China account or phone number and includes a $2 free trial. Check live pricing for the displayed USD rate.

---
Markdown twin of https://chinaapi.ai/models/glm-5.3-flash/ — generated by scripts/gen_model_pages.py. Full catalog: https://chinaapi.ai/llms-full.txt
