SWE-bench

AI Benchmark Flaws Exposed: How Simple Exploits Can Fake Near-Perfect Scores
Researchers behind the Terminator-1 coding agent say seven design flaws in major AI benchmarks allow simple exploits to fake near-perfect scores without solving actual tasks.
