Miles and Pallets
Ports & hubs

New paper questions reliability of common reasoning benchmark

The authors argue a widely cited benchmark may be more contaminated by training data than previously assumed.

New paper questions reliability of common reasoning benchmark

A new preprint questions the reliability of a widely cited reasoning benchmark, arguing that contamination from training data may inflate scores more than previously assumed across several major model families.

The core argument

The authors compared performance on the original benchmark against a newly constructed variant with no public exposure, finding a meaningful score drop across every model tested, though the size of the drop varied significantly.

  • The gap was largest for smaller, more heavily fine tuned models
  • Larger frontier models showed a smaller but still notable gap
  • The authors call for benchmark rotation as a standard practice

Related coverage

More from Ports & hubs