GLM-5.3 Scores 60 on Artificial Analysis Intelligence Index, Matching Kimi K3

Independent AI evaluation platform Artificial Analysis has published its benchmark results for Z.ai's GLM-5.3, awarding the reasoning model a score of 60 on its Intelligence Index v4.1.1. The result places GLM-5.3 level with Moonshot AI's Kimi K3 and three points behind frontier leader Claude Opus 5 (63). The evaluation tested GLM-5.3 at its maximum reasoning effort configuration across a nine-part battery that measures agentic tool execution, terminal coding, graduate-level scientific problem-

2 min
GLM-5.3 Scores 60 on Artificial Analysis Intelligence Index, Matching Kimi K3

Independent AI evaluation platform Artificial Analysis has published its benchmark results for Z.ai's GLM-5.3, awarding the reasoning model a score of 60 on its Intelligence Index v4.1.1. The result places GLM-5.3 level with Moonshot AI's Kimi K3 and three points behind frontier leader Claude Opus 5 (63).

The evaluation tested GLM-5.3 at its maximum reasoning effort configuration across a nine-part battery that measures agentic tool execution, terminal coding, graduate-level scientific problem-solving, physics, factual reliability, and long-context comprehension. Across a comparison class of 181 models with a median score of 35, GLM-5.3 ranks eighth overall.

GLM-5.3 Evaluation on Artificial Analysis Intelligence Index

Post-Training Gains on Base Weights

Z.ai released GLM-5.3 on August 14, 2026, without initiating a new base pretraining run. The 753-billion parameter model uses the identical base weights as its predecessor, GLM-5.2. All measured performance improvements stem from scaled reinforcement learning applied across long-horizon task environments over a four-week post-training cycle.

Vendor-reported data highlighted significant jumps on specialized coding benchmarks, with Terminal-Bench 3.0 climbing from 4.6 to 28.3 and DeepSWE v1.1 rising from 46.2 to 66.9. Artificial Analysis's independent test run confirms that these post-training gains translate into broader general reasoning and tool-use capabilities under standardized third-party evaluation.

Inference Cost and Token Economics

While GLM-5.3 matches Kimi K3's score of 60, the two models differ significantly in API pricing and token consumption patterns:

  • Token Pricing: Z.ai lists GLM-5.3 at $1.40 per million input tokens and $4.40 per million output tokens. In comparison, Moonshot prices Kimi K3 at $3.00 per million input and $15.00 per million output tokens.
  • Cost Per Evaluation Task: Across the Intelligence Index benchmark, GLM-5.3 averaged $0.68 per task, compared to $0.84 for Kimi K3 and $2.34 for Claude Opus 5. The full evaluation run across the suite cost $1,238.50 on Z.ai's API.
  • Generation Verbosity: GLM-5.3 exhibited high output token volume during reasoning steps, generating 170 million output tokens across the entire evaluation battery against a class median of 72 million tokens.

Deployment and Planned Open Weights

GLM-5.3 is currently accessible through Z.ai's hosted API and GLM Coding Plan subscriptions. The model enforces active reasoning traces and supports three selectable effort tiers, rejecting API calls that attempt to disable reasoning mode.

Z.ai has scheduled an open-weights release two weeks after launch following final safety evaluations, which would establish GLM-5.3 alongside Kimi K3 as the highest-scoring open-weight architectures on the Artificial Analysis leaderboard.

Sources

Written by

More to read

  • Anthropic Demonstrates Autonomous De Novo Protein Design and Chemical Analysis with Claude

    Anthropic Demonstrates Autonomous De Novo Protein Design and Chemical Analysis with Claude Anthropic has published experimental results demonstrating Claude's ability to autonomously design de novo protein binders with physical wet-lab validation and automate complex analytical chemistry workflows. The findings show frontier LLMs acting as autonomous agents across computational biology and molecular characterization pipelines. In the primary experiment, Anthropic evaluated Claude Mythos Previe

    1 min
  • Cerebras Unveils CS-4 Rack-Scale System Powered by Three WSE-3 Turbo Chips and Nexus Architecture

    Cerebras Unveils CS-4 Rack-Scale System Powered by Three WSE-3 Turbo Chips and Nexus Architecture Cerebras Systems has announced the CS-4, a rack-scale AI accelerator system designed around three of its next-generation Wafer Scale Engine 3 Turbo (WSE-3 Turbo) chips and a modular hardware architecture dubbed Nexus. Cerebras confirmed that initial customer shipments for the CS-4 are scheduled to begin in the current quarter. The new system marks a structural shift from Cerebras's single-wafer CS

    1 min
  • AI FinOps: Cutting LLM Inference Costs by 30-60% Through Model Tiering, Caching, and GPU Optimization

    AI FinOps: Cutting LLM Inference Costs by 30-60% Through Model Tiering, Caching, and GPU Optimization Inference costs have become the second-largest line item in enterprise AI budgets, trailing only talent spend according to RapidData's State of Enterprise AI 2026. This shift represents a fundamental inversion from the 2021-2023 era when training dominated AI expenditure. The compounding nature of serving costs—accumulating every hour as long as users hit the API—means that even modest producti

    1 min