GLM-5.3 benchmarks, pricing and how it compares with Flash and rivals
GLM-5.3 scored 42% on Artificial Analysis's Terminal-Bench 4.0, against 33% for GLM-5.3-Flash. Z.ai charges roughly nine times Flash's per-token price for it. That makes the flagship worth testing on coding tasks Flash fails, while Flash remains the cheaper choice for work it handles.
Z.ai announced GLM-5.3 on August 14, 2026. It keeps GLM-5.2's base model and API list price, with further post-training aimed at long engineering tasks. The weights are on Hugging Face. Existing users need to update requests that disable thinking before switching models.
What changed from GLM-5.2?
Z.ai says it expanded its reinforcement-learning environments to include longer engineering tasks with executable checks. On its private Z.ai Code Bench, it reports 34.5% task completion at max effort using roughly 75K output tokens per task. GLM-5.2 reached 23.4% using 96K. Z.ai controls this test and its task mix.
Z.ai's public benchmarks point the same way. It reports 28.3 versus 4.6 on Terminal-Bench 3.0, 66.9 versus 46.2 on DeepSWE v1.1, and 28.5 versus 23.8 on Agents' Last Exam. For Terminal-Bench 3.0, it used Claude Code 2.1.207, max reasoning effort, 400K context, three rollouts per task, and limits of 600 turns and ten hours. Terminal-Bench 2.1 and 4.0 use different evaluations.
Requests that disable thinking fail on GLM-5.3. It accepts reasoning_effort values of low, high and max, with max as the default. If an application sends thinking.type: "disabled" to GLM-5.2, Z.ai says to change the request to "thinking": {"type": "enabled"} and "reasoning_effort": "low" before changing the model ID to "glm-5.3". Flash also accepts only enabled thinking, so the same disabled-thinking request needs revision before a switch to either model. Z.ai recommends max for complex coding tasks. Keep the effort level fixed when comparing them.
What do independent GLM-5.3 benchmarks show?
Artificial Analysis's Intelligence Index v4.3.2 gives GLM-5.3 at max effort a score of 45. The index combines ten evaluations across coding, agent work, reasoning and other tasks. On its direct comparison with Flash, the flagship scores 42% versus 33% on Terminal-Bench 4.0 and 59% versus 52% on SciCode. It is tied at 80% on the long-context AA-LCR v1.1 test and trails Flash on GDP.pdf, 11% to 15%. The composite lead, 45 to 42, hides those different outcomes.
In its September 25 snapshot, Artificial Analysis measured 63 output tokens per second for GLM-5.3 through Z.ai's API and 43 for Flash. The flagship was faster in this comparison, though its 63 tokens per second were below the site's 72-token median for similarly sized open-weight models. On the same comparison page, it took roughly one-third less weighted decode time per Intelligence Index task than Flash. This measure excludes the wait for the first token and other overhead. Output use per task was similar, at 71K versus 69K tokens, so verbosity does not explain the time gap. These live API measurements can change.
Z.ai's cyber evaluation reports GLM-5.3 at 84.5 on CyberGym versus 77.2 for 5.2. That is less than one point ahead of Fable 5 with fallback (83.8) and GPT-5.6 Sol (83.6) in the same table. Z.ai used a single run across 1,507 CyberGym tasks. On ExploitBench, it reports average coverage of 54.4 versus 24.4 for 5.2 across 41 tasks and three revisions with a 300-round limit.
GLM-5.3 vs GLM-5.3-Flash: what does the extra spend buy?
Z.ai says the flagship reuses GLM-5.2's base and accepts text. Flash has a newly trained base and about 320B total and 18B active parameters. Z.ai's API guide lists text, image and video input for Flash.
Z.ai's September 25 direct API prices put the models in different cost brackets:
| Z.ai direct API | GLM-5.3 | GLM-5.3-Flash |
|---|---|---|
| Input / 1M tokens | $1.40 | $0.15 |
| Cached input / 1M tokens | $0.26 | $0.03 |
| Output / 1M tokens | $4.40 | $0.50 |
| Artificial Analysis Index v4.3.2 | 45 | 42 |
| Artificial Analysis Terminal-Bench 4.0 | 42% | 33% |
| Artificial Analysis cost per index task | $2.01 | $0.25 |
The last three rows come from Artificial Analysis. Its cost per task uses the evaluator's prompts, token mix and index weights. A team's cost per accepted coding change also includes retries and review.
Z.ai lists GLM-5.3-FlashX between Flash and the flagship at $0.37 per million input tokens and $1.25 per million output tokens. It claims an inference speed of 200 tokens per second for that faster route. That is a Z.ai speed claim, not the Artificial Analysis measurement above. The guide says FlashX is not yet included in the Coding Plan.
For a simple uncached request with 100K input and 10K output tokens, the direct API list price works out to about $0.184 for the flagship and $0.020 for Flash. Reused prefixes lower both bills at their cached-input rates. The ratio also changes when one model needs more turns or output tokens to finish. Our Flash article examines its provider pricing, cache behavior and independent production reports in more detail.
Flash is the cheaper starting point for bounded tasks and for workflows that need image or video input. The flagship has the better case where missed steps, long tool sequences or code review failures are costly, particularly given the Terminal-Bench 4.0 difference. Run the flagship first on coding tasks where Flash fails, then compare accepted results per dollar and time spent reviewing them at a stated effort level.
GLM-5.3 vs Kimi K3, DeepSeek V4 Pro and Qwen3.8 Max
Z.ai's benchmark table compares the flagship with Kimi K3, DeepSeek V4 Pro-0813 and Qwen3.8-Max. The lead changes by task:
| Z.ai-reported evaluation | GLM-5.3 | Kimi K3 | DeepSeek V4 Pro-0813 | Qwen3.8-Max |
|---|---|---|---|---|
| Terminal-Bench 2.1 | 88.2 | 88.3 | 87.9 | 86.6 |
| DeepSWE v1.1 | 66.9 | 67.5 | 62.7 | 56.6 |
| NL2Repo | 58.0 | 58.0 | 61.1 | 55.9 |
| AutomationBench v1.0.6 | 48.2 | 46.7 | 43.2 | 39.8 |
GLM-5.3 vs Kimi K3 is particularly close on the first three rows. DeepSeek leads this group on NL2Repo, while GLM leads on AutomationBench. Z.ai's Terminal-Bench 3.0 row has no DeepSeek or Qwen result.
Artificial Analysis scores GLM-5.3 and Kimi K3 45 and 44 on its full index. The flagship leads 42% to 13% on Terminal-Bench 4.0, and SciCode is tied at 59%. Kimi leads five of the ten component evaluations, including long-context AA-LCR (89% to 80%) and GDP.pdf (22% to 11%). Their measured costs per index task nearly match at $2.01 and $2.00, respectively. Kimi's output rate is $15 per million tokens against GLM-5.3's $4.40. At the measured 48K and 71K output tokens per task, output accounts for about $0.72 of Kimi's $2.00 and $0.31 of GLM-5.3's $2.01. The comparison table does not break down the remaining charges by input, cache or other token types.
The GLM-5.3 vs DeepSeek V4 Pro-0813 comparison puts GLM-5.3 ahead 45 to 36 on the full index, 42% to 14% on Terminal-Bench 4.0, and 59% to 51% on SciCode. DeepSeek costs $0.67 per index task against $2.01 for GLM-5.3. In the September 25 max-effort snapshot, its API generated 84 output tokens per second against GLM-5.3's 63 and took about one-third less weighted decode time per task. Artificial Analysis's Terminal-Bench 4.0 is a different evaluation from Z.ai's Terminal-Bench 2.1 table above.
The GLM-5.3 vs Qwen3.8 Max (0902) comparison from Artificial Analysis ties the models at 45 on the overall index. Qwen leads four of the ten components: AA-Briefcase, GDPval-AA, Humanity's Last Exam and GDP.pdf. GLM-5.3 leads five, including Terminal-Bench 4.0 (42% to 39%) and SciCode (59% to 52%), while AA-LCR is tied. Qwen's weights are closed, and its measured cost is $5.41 per index task against GLM-5.3's $2.01. Qwen generated 108K output tokens per task at $6 per million output tokens, versus 71K at $4.40 for GLM-5.3. Those output charges work out to about $0.65 and $0.31 per task, so they account for only part of the $3.40 total gap. Z.ai's table above names Qwen3.8-Max without a revision, so its rows cannot be identified as results for 0902.
Against proprietary models, Z.ai reports GLM-5.3 above Opus 4.8 on Terminal-Bench 3.0 (28.3 against 21.1) but below Fable 5 with fallback (33.7) and GPT-5.6 Sol (34.6). On DeepSWE v1.1 it trails both Fable 5 with fallback and Sol.
GLM-5.3 specifications, pricing and API access
The public checkpoint configuration, Z.ai's API guide and direct API pricing give these reference values:
| Specification | GLM-5.3 |
|---|---|
| Developer | Z.ai |
| Announced | August 14, 2026 |
| Z.ai API model ID | glm-5.3 |
| Public checkpoint | zai-org/GLM-5.3 |
| Total / active parameters | Roughly 750B / 40B |
| Layers | 78 model layers, plus one multi-token prediction (MTP) layer |
| Architecture | Mixture of experts (glm_moe_dsa in the config) |
| Experts | 256 routed, 8 selected per token, plus 1 shared |
| Vocabulary | 154,880 tokens |
| Configured context | 1,048,576 positions |
| Maximum API output | 128K tokens |
| Input / output | Text / text |
| Thinking | Always enabled, with reasoning_effort of low, high or max (default) |
| Public weights | FP8 E4M3, 128 × 128 blocks |
| License | GLM-5.3 License |
| Input / 1M tokens (Z.ai direct API) | $1.40 |
| Cached input / 1M tokens (Z.ai direct API) | $0.26 |
| Output / 1M tokens (Z.ai direct API) | $4.40 |
A provider may impose a lower context limit than the checkpoint's 1,048,576 positions. For local deployment, the usable limit depends on the memory available for the KV cache.
Artificial Analysis lists 753B total and 40B active parameters, while the current vLLM recipe uses about 743B and 39B. The table rounds both figures. GLM-5.3 pricing matches GLM-5.2's direct API rates. Switching from 5.2 has no per-token list-price penalty, though output use and completion rate still determine the bill for finished work.
The GLM-5.3 developer guide lists OpenAI-compatible Chat Completions and Responses endpoints and an Anthropic Messages endpoint. It warns that accounts with a previous Coding Plan subscription, including an expired one, may currently have access only through the Chat Completions-compatible protocol. The same page lists /api/coding/paas/v4 as that protocol's base URL but uses /api/paas/v4/chat/completions in its examples. Coding Plan usage follows points and off-peak rules. Confirm the working endpoint and applicable billing route for your account before changing production requests.
GLM-5.3 hardware requirements and license
The public FP8 checkpoint is about 756 GB across 141 safetensors shards. The vLLM deployment recipe specifies eight H200 or H20 GPUs for a single-node FP8 setup and eight B200s with FP8 KV cache for its full one-million-token context example. Its AMD MI300X/MI355X example sets a 524,288-token maximum. BF16 weights are in zai-org/GLM-5.3-BF16 and need multiple nodes in that recipe. Inferact publishes a separate, approximately 465 GB Inferact/GLM-5.3-NVFP4 quantization for Blackwell GPUs. Active parameters describe compute per token, while a server still needs space for all weights and the KV cache.
The recipe currently labels the model "vLLM 0.29.0+" while its prerequisites and install command use 0.28.0. Confirm the supported build and available GPU memory before running it in production.
The weights are open under the GLM-5.3 License. It permits use, modification and distribution, including commercial use. If a licensee or any affiliate operates a model-as-a-service business and their aggregate revenue exceeds $10 billion over any consecutive 12 months, the licensee must pass Z.ai's security review before using GLM-5.3 or a derivative for any commercial purpose. Flash's weights carry an MIT license.
GLM-5.3 is the first model to test on terminal-heavy agent work that Flash fails: it leads Flash on Artificial Analysis's coding evaluations and costs about the same per index task as Kimi K3. Kimi wins five of the ten index components. Qwen3.8 Max (0902) matches GLM-5.3's overall index score at about 2.7 times the cost per task. DeepSeek V4 Pro costs about one-third as much per index task despite its lower coding scores. For cost-sensitive work, start with Flash or DeepSeek. The open question is cost per accepted change on your own tasks, measured with provider, effort level, tokens and review time recorded.
Member discussion