The newly released GLM 5.2 ARC-AGI benchmark results have set a historic record for open-weight models. Developed by Z.ai, this 753-billion parameter model achieved an impressive 77.0% on ARC-AGI-1 and 22.8% on ARC-AGI-2. These scores represent the highest ever recorded for an open-weight system on these challenging reasoning tests.
This achievement highlights how open-source artificial intelligence is rapidly closing the capability gap with proprietary models. According to a report by Office Chai, the model performs comparably to low-reasoning configurations of GPT-5.4 and GPT-5.5.
What the GLM 5.2 ARC-AGI Scores Mean
The Abstraction and Reasoning Corpus for Artificial General Intelligence (ARC-AGI) measures general fluid intelligence. It forces models to solve entirely novel visual grid-based puzzles with minimal prior training. GLM 5.2 achieved its ARC-AGI-1 score at just $0.19 per task, while the more difficult ARC-AGI-2 cost $0.25 per task.
Architectural Innovations Behind the Efficiency
To achieve this level of efficiency on modern computers, Z.ai introduced several key architectural upgrades. The model utilizes an optimization called IndexShare, which shares a single attention index across multiple sparse layers. This reduces per-token compute by roughly three times at long context lengths. Additionally, an upgraded Multi-Token Prediction layer boosts speculative decoding by up to 20% during inference.
The Remaining Gap to Frontier Systems
Despite these record-breaking GLM 5.2 ARC-AGI results, a significant gap remains when compared to the absolute frontier. OpenAI’s GPT-5.5 currently leads the ARC-AGI-2 benchmark with an 85% score, followed closely by GPT-5.4 Pro at 83.3%. GLM 5.2’s score of 22.8% shows that proprietary systems still hold a massive advantage in complex, novel reasoning tasks.
Market Impact and Chinese AI Competitiveness
The release has had an immediate effect on the global tech economy. Shares of Knowledge Atlas, the publicly listed entity connected to Z.ai, have roughly doubled since the model’s debut. Furthermore, VentureBeat reports that GLM 5.2 beats GPT-5.5 on multiple long-horizon coding benchmarks for a fraction of the cost, signaling a rapid acceleration in Chinese AI development.




