MModelspectraIndependent AI Model Intelligence
Model profile · Updated 2026-09-08
Model Profile · Zhipu AI · CN

GLM-5.3-Flash

Best forUltra-fast, Ultra-low-cost, High concurrency
LicensingOpen weights
Updated2026-09-08
01

Snapshot

Overall
65
rank #18 of 22
Coding
60
rank #20 of 22
Multimodal
63
rank #21 of 22
Input / Output
$0.07 / $0.25
USD per 1M tokens
Context
128K
usable recall 80
Speed
~90 tok/s
TTFT 0.2s

Aggregated from public sources and independently weighted; methodology on the Terms page. Scores within 3 points are treated as statistically tied.

02

Specs & pricing

Every tracked field. List API prices in USD per 1M tokens; verify current vendor pricing before purchase.

VendorZhipu AI (CN)
Released / current2026.08
LicensingOpen weights
Context window128K
Max output64K
Effective-context score80
Input / output price$0.07 / $0.25 per 1M tokens
Cache discountNo published cache discount
Free tierChatGLM free
Speed / TTFT~90 tok/s / 0.2s
Function calling70
Refusal rate~6%
English / Chinese65 / 85
Modalitiestext
Fine-tuningYes
Private deploymentYes (enterprise)
SOC2 / no-trainno / yes
03

Strengths & watch-outs

Strengths

  • Among the cheapest
  • Fast
  • Open source

Watch-outs

  • Limited capability
  • Text-only
  • Small context
04

Where it ranks among 22 models

Overall 65/100 (#18), coding 60 (#20), multimodal 63 (#21). See the full boards and side-by-side compare on the leaderboard.
05

Pricing in practice

Illustrative monthly bill for 100M input + 30M output tokens: $15 at list. Your mix and cache-hit ratio change this.

Illustrative model using public list rates before any enterprise agreement; cache discount: No published cache discount.

06

Frequently asked questions

How much does GLM-5.3-Flash cost?
List API pricing is $0.07 input and $0.25 output per 1M tokens. No published cache discount. ChatGLM free
What is GLM-5.3-Flash's context window, and is it usable end to end?
It advertises a 128K window with 64K max output; its effective-context score is 80/100, which is the better predictor of whether long-document details are actually retained.
Is GLM-5.3-Flash good for coding?
Its coding aggregate is 60/100, rank #20 of 22. It is adequate for scripts and assisted completion rather than the most demanding SWE-style work.
Can I self-host GLM-5.3-Flash?
Yes — it ships open weights and supports private deployment, so you can run it on your own infrastructure for data control; budget for GPU and operations.
How strong is it in Chinese and on multimodal inputs?
Chinese score is 85 versus English 65; multimodal aggregate is 63/100 (rank #21), accepting text.