AI Benchmarks

Daily AI Roundup: Autonomous Robotics, Cloud Agents, and Frontier Compute
Today’s AI roundup covers Unitree humanoid models, OpenAI training on 100,000 GPUs, Rustuna runtime tools, and Tencent retrieval benchmarks.

Daily AI Roundup: New Local Models, Hardware Accelerators, and Web Standards
Today’s AI roundup highlights Perplexity’s on-device model, Cerebras CS-4 hardware, new agent connectors, and OpenAI’s WebMCP open web standard.

NVIDIA System Scores 100% on ARC-AGI-3 Benchmark
NVIDIA’s CUDA-optimized system achieves a perfect score on the ARC-AGI-3 benchmark, solving all 25 public games and 183 levels.

GLM 5.2 Achieves Record-Breaking Open-Weight Scores on ARC-AGI Benchmarks
GLM 5.2 has set a new record for open-weight models on the ARC-AGI benchmarks, delivering high-level reasoning at a fraction of the cost of proprietary systems.

GLM 5.2 Model Claims Top Spot on Global Benchmarks Following US Export Restrictions
The GLM 5.2 model has claimed the top spot on the Bridgebench database, outperforming competitors shortly after US export bans took effect.

AI Benchmark Flaws Exposed: How Simple Exploits Can Fake Near-Perfect Scores
Researchers behind the Terminator-1 coding agent say seven design flaws in major AI benchmarks allow simple exploits to fake near-perfect scores without solving actual tasks.

Xiaomi MiMo-V2-Pro: A 1-Trillion Parameter Model Challenging Western AI Leaders on Cost and Performance
Xiaomi released MiMo-V2-Pro, a 1-trillion parameter AI model, on March 18, 2026, ranking 10th globally on Artificial Analysis’s Intelligence Index at roughly one-seventh the cost of comparable Western models.

MiniMax M2.7 Model Launches on Agent Platform and API with Self-Improvement Capabilities
MiniMax has released its M2.7 model via the MiniMax Agent and API Platform, featuring reinforcement learning-based self-improvement and a 97% skill adherence rate across 40+ complex tasks.

Mistral Small 4 Unifies Reasoning, Multimodal, and Coding in One Open-Source Model
Mistral AI has released Mistral Small 4, an open-source model combining reasoning, multimodal, and coding capabilities under the Apache 2.0 license. The model delivers a 40% reduction in completion time and three times more throughput than its predecessor.
