GLM-5.2 Release: New Leading Open Weights Model on Artificial Analysis Index

Z ai's GLM-5.2 is now the leading open weights model on the Artificial Analysis Intelligence Index v4.1, scoring 51. This release places GLM-5.2 on the Pareto frontier of intelligence versus cost per task, offering the lowest cost per task among models at its specific intelligence level.

Benchmark Performance and Intelligence

GLM-5.2 outperforms other leading open weights models on the Intelligence Index v4.1, surpassing MiniMax-M3 (44), DeepSeek V4 Pro max (44), and Kimi K2.6 (43).

Scientific Reasoning and Agentic Performance

The model shows significant gains over its predecessor, GLM-5.1, particularly in scientific reasoning and agentic tasks:

  • GDPval-AA v2: GLM-5.2 scores 1524, leading MiniMax-M3 (1418) and DeepSeek V4 Pro max (1328). This score is effectively level with the proprietary GPT-5.5 (xhigh reasoning) at 1514.
  • Scientific Reasoning: Notable improvements include CritPt (+16 points to 21%), HLE (+12 points to 40%), and SciCode (+7 points to 50%).
  • Other Gains: AA-LCR increased by 9 points to 71%, tau3 banking by 15 points to 27%, and TerminalBench v2.1 by 16 points to 78%. GPQA Diamond improved by 3 points to 89%.
  • Omniscience Index: GLM-5.2 scored 4 (up from 2 for GLM-5.1), driven by higher accuracy (25.1% vs 24.2%) and a lower hallucination rate (28.1% vs 29.4%).

Technical Specifications and Availability

GLM-5.2 maintains the same architectural size as GLM-5.1 but introduces a significantly expanded context window.

  • Model Size: 744B total parameters with 40B active parameters (MoE).
  • Context Window: 1M tokens, a substantial increase from the 200K tokens available in GLM-5.1.
  • License: MIT License.
  • Pricing: $1.4 per 1M input tokens, $4.4 per 1M output tokens, and $0.26 per 1M cache hit tokens.
  • Availability: Accessible via Z ai's first-party API and third-party providers including DeepInfra, Novita, Nebius, Parasail, Siliconflow, GMI Cloud, Baseten, and Fireworks.

Efficiency and Resource Utilization

While highly intelligent, GLM-5.2 is less token-efficient than some of its peers, utilizing more output tokens to reach its conclusions.

  • Token Usage: The model uses 43k output tokens per Intelligence Index task (37k of which is reasoning), compared to 26k for GLM-5.1 and 24k for MiniMax-M3.
  • Cost per Task: GLM-5.2 costs approximately $0.46 per task. While higher than DeepSeek V4 Pro max ($0.05), it remains on the Pareto frontier for its intelligence level.

Community Insights and Practical Observations

Technical discussions from the community highlight a divide between benchmark success and real-world deployment experiences.

Reasoning and Speed

Some users have noted that the model's high reasoning capability comes at the cost of speed. One user reported that for a math evaluator library task, the model spent over 15 minutes reasoning and consumed 45k tokens before producing the first file, noting that "GPT 5.5 is extremely reasoning efficient" by comparison.

Modality Gaps

A recurring critique is the lack of vision capabilities. Unlike Gemma 4, Qwen 3.6, or the OpenAI/Anthropic suites, GLM-5.2 is text-input only. This is cited as a significant gap for tasks like web design where screenshot-to-code workflows are standard.

Reliability and Infrastructure

User reports indicate mixed results regarding stability and capacity:

  • Stability: Some users find GLM-5.2 more stable than newer proprietary models, comparing it to the consistency of Opus 4.6.
  • Capacity: Multiple users reported significant infrastructure struggles, including timeouts, 429 rate limits, and slow speeds on the official Z ai API.
  • KV Caching: Conversely, some users praised the implicit cache hit rate of the official API, claiming it exceeds 95%.

Comparative Positioning

Community members have debated the model's actual standing relative to the "frontier":

"Open models are on about a 4-7 month lag right now... if this keeps up, you might see an open-weights model doing claude fable 5 level work before the new year."

While some view it as a "massive win for the rest of the world" due to its pricing and open weights, others argue that for absolute frontier performance, proprietary models like GPT-5.5 and Claude Fable 5 still maintain a significant lead in efficiency and abstract reasoning.

Sources