New paper questions reliability of common reasoning benchmark
The authors argue a widely cited benchmark may be more contaminated by training data than previously assumed.
A new preprint questions the reliability of a widely cited reasoning benchmark, arguing that contamination from training data may inflate scores more than previously assumed across several major model families.
The core argument
The authors compared performance on the original benchmark against a newly constructed variant with no public exposure, finding a meaningful score drop across every model tested, though the size of the drop varied significantly.
- The gap was largest for smaller, more heavily fine tuned models
- Larger frontier models showed a smaller but still notable gap
- The authors call for benchmark rotation as a standard practice